Jev's Four Application Patterns: Constraining LLM Generation into Structured Decisions

The article analyzes four community-driven patterns for using Jev — semantic filtering, action selection, control-plane routing, and evidence gating — each narrowing open-ended LLM generation into a constrained judgment, with concrete examples, benchmarks, and production readiness checks.

Architect
Architect
Architect
Jev's Four Application Patterns: Constraining LLM Generation into Structured Decisions

Jev, a TypeSafe "System One" model, returns only constrained answers (Choice, Noul, Score) with probabilities given a state and named questions. The author maps half a month of community demos into four distinct integration points along a request's data flow.

What Jev Actually Returns

TypeSafe's API accepts a state payload and one or more named questions. Jev does not generate explanations; it returns only the constrained answer and a probability distribution. Three question types are provided: Choice – pick one item from caller-supplied candidates, returning the candidate distribution. Noul – judge the probability that a proposition holds. Score – rate on an ordered scale.

Multiple independent questions can be answered in parallel against the same state; dependent questions require sequential calls orchestrated by code. Jev only answers; application code decides what the answer triggers.

Pattern 1: Semantic Filtering – Turning Natural Language into Callable Conditions

Jev Semantic Grep feeds text line-by-line to the model, asking "does this line express meaning X?" and filters by Noul probability. Unlike literal grep, it matches propositions (e.g., a line describing a wolf wearing a grandmother's hat matches "disguise" even without the word). Uses include semantic search, log monitoring, email/ticket triage, pre-execution checks (command safety, PR merge eligibility, prompt-injection detection), and post-execution trace filtering.

This acts as a "probabilistic if" : deterministic code handles timestamps, status codes, numeric thresholds; Jev handles semantic conditions hard to express in regex. Code then combines results with AND/OR/NOT and threshold policies.

Production concerns: at 10 GB of logs, per-line remote calls are slow and costly. Practical pipelines pre-filter by time window, fields, and keywords before sending complete events to Jev. Data privacy: logs containing accounts, tokens, request params, and customer data must pass through field allow-lists, redaction, and tenant isolation before hitting the TypeSafe API — local grep and remote semantic judgment operate under different trust assumptions.

The article notes Newsjack as a similar decomposition (search, relevance, fact-check, routing) but its README does not claim Jev usage or provide Jev-attributable benchmarks; it illustrates task decomposition, not Jev performance.

Pattern 2: Action Selection – Choosing the Next Step from Legal Candidates

Jev Browser and Computer Use demos exemplify this. The environment supplies executable actions; Jev selects one; execution result feeds the next round.

Jev Browser collects clickable, inputtable, selectable DOM elements plus scroll, back, done actions. Each main loop answers three questions around the same page state: Choice for action, two Noul for goal-achieved and stuck detection. Code checks stop conditions before execution; on repeated invalid actions it can fall back to the second-best candidate from the probability distribution.

A published Wikipedia trajectory required 3 Jev calls, estimated cost $0.0022. Traces log suggested action, actual action, probabilities, page errors, stop reasons. done and goal_achieved are independent judgments; both must agree for a reliable stop signal.

Explicit capability boundaries: Jev only selects — it does not generate search terms or form content (a separate small model handles string input); passwords are supplied by code, not the model. Each step caps at 240 page elements (truncated beyond); Shadow DOM, iframes, hover menus are unsupported in v0.1.

End-to-end effectiveness depends on the whole harness: how the page becomes candidates, whether required elements enter the list, post-execution verification, failure recovery, budget exhaustion handling. Jev handles selection; Harness handles perception, execution, budget, verification, recovery.

Game demos (Doom, Wikiracing, 50 parallel Subway Surfers, Mario, Slay the Spire, pixel-perfect color picking) and real-world tasks (Browser Use, flight search, 3D driving sim, MuJoCo rocket landing, robot control, ReAct loops, on-chain orders) reuse the same state-candidate-select-execute loop. Many are still demo videos or second-hand descriptions; the author refrains from quoting their speed/cost numbers. The key takeaway: any program that can enumerate state and legal actions can plug Jev into the selection node.

As stakes rise (real devices, money, production systems), mis-execution cost grows. Web mis-clicks are recoverable; robot or on-chain actions may be irreversible. Candidate sets must be pre-filtered by permissions; high-impact actions need approval, rate limits, sandboxes, or human confirmation. The model chooses only from "allowed actions" — high probability does not grant new permissions.

Pattern 3: Control-Plane Routing – Picking Skills, Tools, or Models

When an agent loads many skills/tools/models, stuffing all descriptions into the primary model wastes context and repeats the same routing decision each turn. A routing layer reads the request first and decides which capability to load.

A public Jev Agent Skill Router experiment used 24 synthetic skills and 72 synthetic requests. Jev routed 68 correctly (94.4%); a lexical baseline scored 51 (70.8%). Jev's edge appeared on requests with clear semantic intent but no keyword/name match.

The author cautions: data is synthetic, routing prompts were refined after early full runs, the same batch was reused during development, and results include 3 unnecessary human reviews and 1 client-side validation failure. The 94.4% signals the direction is worth further testing, not production accuracy.

More valuable than the single number is the process: skill catalog is first split into deterministic batches; Jev retains candidates from each batch; a merge round follows. Final output can be route, no_skill, or review. When no skill fits or confidence is low, the flow pauses for review instead of forcing a guess.

A community TypeSafe Router project summarizes the boundary: "Jev selects. Your application authorizes and executes." It is marked as a community project, not an official SDK or production authorization system.

Related work includes tool routing, model routing, specialist agent dispatch, Jev Codex Router for task-difficulty sharding, and LLM pre-routing (LangChain's external decision framework and routing middleware). All decide who handles the task; tenant model access, budget sufficiency, tool production access remain enforced by deterministic policies.

Type safety only guarantees the return value falls in the predefined set, not that the model picks correctly. Akshay Pachaar notes models can still choose wrong among legal options. When probabilities are close, review is hit, or input distribution shifts, the system must have escalation paths to a general model or human. Thresholds must be measured on your own labeled data, not copied from a demo.

Pattern 4: Evidence Gating – Deciding What Enters the Context

Long-running agents accumulate logs, tool results, retrieval chunks, and history. Jev can pre-filter, passing only evidence likely to affect the current question to the general model.

The underlying call may be identical to semantic filtering; the difference is where the result goes. Semantic filtering decides which business objects proceed downstream; evidence gating directly rewrites the next-turn context. Missing key evidence cannot be recovered by stronger downstream reasoning.

Implementations focus on RAG result filtering/reranking, long-context compaction (e.g., fast-jev-compaction), old conversation pruning, and tool-result trimming. Each item is scored for relevance to the upcoming task; code keeps or discards.

jevskill published a small A/B test: six tasks × 3 runs. Raw input to answer model totaled 113,632 tokens; after Jev filtering, 831 tokens (99.3% reduction). Raw input answered 15/18 correctly; Jev-filtered and structured-filtered both answered 18/18. One Deployment JSON task compressed from 20,542 to 111 tokens, finding the anomaly all three times.

These numbers suggest semantic gating can cut cost and noise simultaneously, but each condition had only 3 runs; the author calls them directional. The author is more interested in the one failure mode.

An earlier test filtered YAML line-by-line; the task required comparing default vs. production values inside the same config block. Jev found prod: true but the field name was on the previous line — the model couldn't see the relationship, scoring 0/3. The fix: change the decision unit to "heading plus its full content block", turning the task to 3/3.

Relationships lost at the chunking stage cannot be recovered by the downstream model. In context compression, keeping a few extra chunks only costs tokens; dropping decisive evidence changes the answer.

For RAG filtering, log compaction, and task-completion verification, compression ratio alone is insufficient. The author would additionally measure: key-evidence recall rate, miss rates per input type, and fallback cost. Structured rules that precisely locate content should retain it first; cross-line, cross-field, cross-call relationships must be placed in the same minimal complete unit. High-risk tasks can keep the uncompressed context and re-run on low confidence.

Putting All Four Patterns in One Agent

A hypothetical incident-response agent receiving "payment success rate dropping" alert would chain the four patterns:

Semantic filtering on incoming logs/alerts to isolate relevant events.

Evidence gating to compress relevant logs, tool outputs, and history into the diagnostic context.

Control-plane routing to select the appropriate diagnostic skill (e.g., payment-gateway skill vs. database skill).

Action selection within the chosen skill to drive browser/CLI/API steps for mitigation.

Jev only interposes on these constrained judgments. Code guards permissions and state machines; Harness manages loops, budgets, recovery, stop conditions; tools execute side-effecting actions; the general model handles open-ended diagnosis; high-risk gray zones go to humans.

Not every pattern must be adopted. If log filtering works with SQL/rules, keep it. If only three skills exist, feeding descriptions to the primary model may be simpler. Jev integration is justified only when a constrained semantic judgment is frequent enough and demonstrably reduces downstream calls or context.

Constrained Choices Don't Always Need Jev

The 2048 game illustrates the boundary. LogJev's demo uses rules to eliminate illegal moves, then lets the model pick from up/down/left/right. The interface could be Choice, but legal moves are exactly computable by rules, and board search has mature algorithms. The problem needs position calculation, not natural-language semantic understanding.

Moreover, that demo uses LogJev — a wrapper around standard OpenAI-compatible top_logprobs packaged into Choice/Score/Noul style — not TypeSafe's official Jev. It demonstrates Decision Model programming style, not Jev's speed, cost, or accuracy.

Sean Goedecke questions whether Jev's gains come from constrained candidates, short outputs, and parallel inference — replicable by regular models via single-token selection. A thorough controlled experiment is still missing, but the architectural insight stands: decision nodes can be designed against interfaces and evals without binding to a specific model brand.

PrimeLine's pre-registration comparison is closer to an evaluation methodology (freeze task, metrics, pass criteria, then compare solutions) than a fifth application pattern. It helps decide whether to adopt Jev, but isn't a usage pattern itself.

When facing a new scenario, the author asks four questions: Can the answer set be enumerated upfront? Is semantic understanding irreplaceable? Is the call volume high enough? Can the system recover from a mistake? Only when all four have clear answers does Jev enter A/B testing.

In the author's view, Jev belongs in the agent control plane, handling fast, constrained semantic judgments ; code and infrastructure retain authorization, execution, and rollback. A working demo is step one. Pre-production testing should focus on three things: Are candidates complete? Does the decision unit preserve evidence relationships? Can the system recover from a misjudgment?

References

TypeSafe AI, Introducing System One Models & Jev (https://typesafe.ai/blog/introducing-system-one-models-and-jev)

TypeSafe AI, Agent Skill (https://docs.typesafe.ai/agent-skill)

jkudish, Jev Browser (https://github.com/jkudish/jev-browser)

Junji Uehara, Jev's Killer App: Semantic Grep (https://zenn.dev/uehaj/articles/jev-semgrep-grep-by-meaning)

GodsBoy, Jev Agent Skill Router (https://github.com/GodsBoy/jev-agent-skill-router)

TypeSafeAI Community, Jev Tool & Model Router (https://github.com/TypeSafeAI/typesafe-router)

lazniak, jevskill (https://github.com/lazniak/jevskill)

elvisun, Newsjack (https://github.com/elvisun/newsjack)

DumoeDss, LogJev (https://github.com/DumoeDss/logjev)

Akshay Pachaar, Jev Clearly Explained (https://x.com/akshay_pachaar/status/2101037514945597645)

Sean Goedecke, Jev Means Structured Output Is Interesting Again (https://www.seangoedecke.com/jev-means-structured-output-is-interesting-again/)

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LLM agentscontext compressionsemantic filteringskill routingJevTypeSafeaction selectionconstrained generation
Architect
Written by

Architect

Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.