Jev: Separating Semantic Judgment from Generation in LLM Applications

TypeSafe's Jev introduces a decision model that returns typed probability outputs (Noul, Choice, Score) instead of natural language, moving semantic judgment from runtime generation to a designable, testable 'decision surface' developed with LLM assistance but executed by a specialized model, enabling calibrated thresholds, parallel evaluation, and deterministic code control flow.

Architect
Architect
Architect
Jev: Separating Semantic Judgment from Generation in LLM Applications

Core Problem: Semantic Judgment Coupled with Text Generation

In typical LLM agents, small repeated judgments — whether to call a tool, escalate to human, select a model, or retain retrieved context — are handled by prompting a general-purpose LLM to read state, generate an explanation, and parse a boolean, label, or parameters from the text. This couples "understanding semantics" with "generating text," incurring full generation latency and cost for a single true, and requiring retry logic for format errors and schema drift.

Jev: A System One Model for Decisions

Released by TypeSafe in September 2026, Jev is called a System One Model : a decision model invoked directly by software. The caller supplies a state string and a set of narrow questions with predefined answer spaces. Jev returns structured probability distributions — not natural language — for three primitives: Noul: probability a statement holds (e.g., "user explicitly requests unsubscribe") Choice: distribution over a finite candidate set (e.g., "ticket category: billing, shipping, technical") Score: ordered ranking (e.g., "risk level low→medium→high")

Multiple independent questions are evaluated in parallel in one request. Results include candidate distributions and confidence scores; application code then applies thresholds to decide actions (auto-handle, escalate, call stronger model).

Illustrative Example: Marketing SMS Triage

state = """
Sent SMS: Member day 20% off, reply TD to unsubscribe
User reply: Stop sending me messages
"""

questions = {
  "unsubscribe": {
    "type": "noul",
    "instructions": "Does the user explicitly indicate they don't want more marketing SMS?"
  },
  "complaint": {
    "type": "noul",
    "instructions": "Is the user expressing dissatisfaction with the SMS or service?"
  },
}

answer = jev(state=state, questions=questions)

if answer["unsubscribe"]["noul"] >= 0.90:
  stop_marketing_message()
if answer["complaint"]["noul"] >= 0.80:
  create_customer_service_ticket()

The model no longer answers "how to handle this message." It answers two verifiable questions; actions are composed in code. A reply can be both unsubscribe and complaint without forced single-label classification. When single-choice is needed, Choice with an explicit other option is used; Score expresses degree, not Noul 's 0.5 as "medium."

Comparison with OpenAI Structured Outputs

OpenAI's Responses API with strict: true and JSON Schema (or Pydantic models) constrains output shape — e.g., a Triage class with unsubscribe: bool, complaint: bool. This solves "can the program read the output" but differs from Jev:

GPT remains a generation model; Schema constrains output shape, not the decision primitive.

No built-in Noul / Choice / Score primitives with calibrated probabilities. logprobs reflect token-level likelihood, not decision confidence.

Threshold calibration, fallback strategies, and runtime cost remain the application's burden; Jev places constrained judgment and probability results at the API center.

OpenAI's API can implement part of a Jev-style workflow (LLM outputs boolean/enum/tool params, code branches), but the decision surface definition, calibration, and evaluation are still external.

New Development Paradigm: Decision Surface as Explicit Artifact

Traditional approaches: (1) encode all logic as deterministic rules — brittle for semantic nuance; (2) delegate rules, understanding, and action to a runtime LLM — fast to integrate but each call re-interprets policy, with format drift and auditability issues.

Jev enables a third workflow:

Development time: LLM (or coding agent) drafts the decision surface from business rules and sample data — decomposing "handle unsubscribe" into factual questions, adding other options, writing candidate actions and initial thresholds — then human review.

Evaluation: Annotated dataset evaluation, threshold calibration, shadow runs.

Runtime: Jev executes the reviewed decision surface; deterministic code handles branching, permissions, rollback, audit.

Business rules + sample data
  ↓
LLM drafts questions, options, strategy code
  ↓
Human reviews decision surface
  ↓
Annotated eval, calibrate thresholds, shadow run
  ↓
Jev executes semantic judgment in production
  ↓
Deterministic code completes branching, permissions, rollback

The decision surface — state fields, question phrasing, candidate options, probability-to-action mapping — becomes a reviewable, testable, versioned engineering artifact instead of hiding inside a long prompt.

Performance Claims and Independent Verification

TypeSafe publishes 70–500 ms end-to-end latency, $0.042/M input tokens, output free, claiming ~193.6× speedup and ~444.6× cost reduction — but notes these come from self-designed workflows representing the high end of real gains; "accuracy" is mainly agreement with a reference model, not human-annotated ground truth.

Independent tests (xbill, 8 days) show speed multipliers from <1 to 12.1×, cost multipliers from <1 to 478×, varying by task, baseline model, and network. These multipliers cannot be cited task-agnostically. Total cost must account for reduced generation calls plus new calibration, fallback, and human-review overhead.

Probability ≠ Accuracy: Calibration and Threshold Engineering

Probability is not "accuracy." A 0.90 Noul does not mean 90% correctness on any task. TypeSafe emphasizes RLCD (Reinforcement Learning for Calibrated Decisions), but public reliability curves across tasks are not yet available. Probabilities can shift high or conservative when task or data distribution changes.

A small internal test (92 marketing SMS replies, 2 questions, 3 phrasing variants × 2 state formats × 3 runs = 1,656 calls, $0.037) showed median latency ~716 ms (max 3.97 s incl. network). Sarcastic reply "Send another one?" was not stably recognized as refusal across 6 combinations (refusal probability 0.04–0.29). This illustrates that phrasing, context, and thresholds jointly change actions; a single probability cannot be blindly promoted to production.

Recommended adoption sequence:

Offline evaluation on historical samples.

Shadow run: record probabilities only, no action changes.

Measure real error rates, coverage, and human-review cost per probability bucket.

Gradually route high-confidence traffic to auto-handling.

Continuous sampling of auto-approved cases (high confidence can still be wrong).

Re-evaluate per language, domain terminology, cultural context — English thresholds do not transfer directly to Chinese production data.

Type safety constrains output shape, not semantic correctness. Defining options "billing, technical, sales" prevents a fourth department, but the model can still pick the wrong valid option. Prompt injection in state is not neutralized by enum outputs.

Three Primary Agent Integration Points

Routing: Classify request as simple query, routine modification, or complex architecture issue; select model by cost/capability. Routing result and probability logged for misroute analysis and upgrade cost tracking.

Guardrails: Before tool execution, judge whether command modifies files, deletes data, touches production, or bypasses permissions. High-risk → confirmation/sandbox; low-risk → proceed. Permissions and sandboxing remain system responsibilities; Jev provides only semantic signal.

Verification: When agent claims "task complete," judge whether tests actually pass, output satisfies constraints, or same step repeated. Hard assertions stay in code; Jev covers semantic checks rules cannot express.

LangChain's public examples already embed Jev in agent middleware for model routing and pre-tool risk judgment. Jev does not replace the agent loop; it makes implicit loop judgments explicit nodes.

Boundary with Rule Engines / DMN

Similarities: decision surface human-defined, branching and responsibility boundaries reviewable. Difference: Jev executes semantic judgments requiring natural language or unstructured state understanding. Rule engines excel at "amount > X," "date expired," "field empty" — deterministic, cheap, provable. Jev fills semantic gaps ("does this text express refusal?" "is this ticket closer to shipping or billing?") but does not eliminate rule maintenance, sample annotation, or regression testing.

In a real call chain: code guards hard boundaries, Jev provides semantic signals, general LLM handles open-ended reasoning, human catches gray zones and accountability. More models demand clearer system boundaries.

Engineering Trade-offs: What Replaces the Saved Calls

Runtime judgment code shrinks, but engineering work shifts to: reviewing question phrasing, candidate order, other presence, state context selection (longer state → more irrelevant content and injection risk); calibrating thresholds; re-evaluating on model/question changes; iterating decision surface on new production expressions (not just prompt tweaks).

Observability must capture: state summary, question definition version, model version, probability distribution, final action, fallback reason. Without this, a misroute leaves only "routing wrong" without knowing what the model saw or which threshold fired. Versioned question definitions and threshold changes require re-running regression suites.

End-to-end ROI must be calculated: if every request calls Jev then the original general model, and Jev doesn't reduce downstream work, the system just adds a network hop. Net benefit requires reducing expensive calls, early interception of dangerous actions, or batching low-value judgments.

Limits of the New Paradigm

Jev is not a general AI replacement. Its accuracy, calibration, cross-task generalization, and injection resistance need per-data validation. Complex causal analysis, long-horizon planning, and open-ended generation remain outside its scope.

The "new paradigm" is a change in LLM application development practice , not a model capability leap. Software integrating models often needs a judgment that code can handle, monitor, and assign clear ownership — not another paragraph of explanation. Once the decision surface is designed, reviewed, measured, and versioned, the LLM application gains a traceable intermediate control layer, no longer buried inside a single prompt call.

This reshapes agent design order: first identify which actions are deterministic rules, which need semantic judgment, which need deep reasoning; then assign code, specialized decision model, general model, or human path. Control flow is no longer monopolized by the model; decision surfaces and execution boundaries become co-equal architectural objects.

Ultimate test: in real data and real accountability, is the decision surface easier to understand, measure, and roll back than the original prompt? If not, Jev is just another call layer; if yes, judgment logic truly enters versioning, testing, and audit pipelines.

References

TypeSafe AI, Introducing System One Models & Jev (2026-09-15)

TypeSafe AI, API Reference

LangChain, Building a Harness with Jev

Diogo Almeida, Latent Space podcast Jev: System One models for Prod, not God (2026-09-21)

Sean Goedecke, Jev means structured output is interesting again (2026-09-16)

Akshay Pachaar, Jev Clearly Explained (2026-09-19)

阿蔺 A-Lin, Jev 入门指南与四个开源项目 (2026-09-22)

doronp/jevc, GitHub

xbill, Jev After Eight Days of Independent Tests (2026-09-24)

OpenAI, Structured model outputs

OpenAI, Function calling

OpenAI API Reference, Chat Completions: Create

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Structured OutputAgent EngineeringJevSystem One ModelTypeSafeSemantic JudgmentDecision SurfaceLLM Application Architecture
Architect
Written by

Architect

Professional architect sharing high‑quality architecture insights. Topics include high‑availability, high‑performance, high‑stability architectures, big data, machine learning, Java, system and distributed architecture, AI, and practical large‑scale architecture case studies. Open to ideas‑driven architects who enjoy sharing and learning.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.