Jev Decision Model: Extracting Structured Judgments from LLMs for Code-Driven Routing
This article details the Jev decision model from TypeSafe, which provides typed answers (Choice, Score, Noul) with probabilities instead of text, enabling code-driven routing, verification, and retrieval reranking with lower latency and cost than LLMs, while explaining its four architectural patterns, confidence gating, atomic decomposition, and limitations.
What Is the Jev Decision Model?
TypeSafe's official documentation positions Jev (System One) as a model for building AI-powered software , not for agents. It does not generate code or choose actions. Instead, given a state (text, ticket, records, JSON) and a typed question , it returns a typed answer plus a probability distribution (and confidence for Choice and Score). Your code then branches, sorts, or routes based on that structured output.
state + typed question → typed answer + probabilities (+ confidence) → code executes strategyThree Output Primitives
Choice : pick one from given options → returns choice + probabilities + confidence Score : rate against an ordered rubric → returns score + probabilities + confidence Noul : whether a proposition holds → returns noul (0–1 probability), no confidence
Multiple primitives can be combined in a single request. The official customer-support example asks one Choice (which team to route to), one Score (customer dissatisfaction level), and several Noul questions (requests credentials, sender/domain mismatch, refund request).
Why “No Free Text” Matters
LLMs output text for humans. When code needs a judgment, you must generate formatted text then parse it back — a mismatch. Jev returns structure natively: a typed value with a full probability distribution, and confidence for Choice/Score. It is “type-safe by construction.”
Input → generate text → parse text → decide (standard LLM path)
State + typed question → typed answer + probabilities → code executes (Jev path)The division is not about intelligence but about whether code can directly act on the output. The cookbook draws a clear line: blank lines and explicit markers ( -, 1., #) are read by code; arithmetic and calendar math stay in code — “the model only reads what the text says, never does calendar math.”
Three Architectures: Who Controls the Flow
Traditional software : code controls flow; use where rules are clear.
LLM agents : model decides each step; usable only with human oversight; each loop risks drift.
AI-powered software (TypeSafe’s target): code controls flow ; model appears only where “programmable common sense” or unstructured-data interpretation is needed.
Therefore, “adding a Jev decision layer inside an agent” is misleading: agents assume the model decides the next step, while System One assumes do not give the model decision authority . For agent builders, official use cases show Universal Verification and Harness Engineering : cheap narrow questions verify agent outputs (tool calls, schema compliance, citation support, retrieval manipulation) — Jev sits beside the agent, not driving it.
Why Jev Exists: Latency, Cost, Output Shape
Slow : dozens of loops per task, tens of thousands of tickets daily, each waiting for text generation.
Expensive : per-output-token billing for long text when only a low/medium/high judgment is needed.
Wrong shape : you need an if -ready value; the model gives a sentence to parse.
Official positioning data (source noted):
Latency : “most queries complete in ~100 ms”; real-time apps at 150 ms level.
Cost : model jev-1.13.0 priced at $42 / Btok ($0.042 / Mtok) , input tokens only, output tokens free; rate limit 250k tokens/s, 1,200 req/min; context 64k tokens (state + all questions).
Target : “>100× intelligence-to-speed-and-cost ratio.”
Cost advantage comes from post-training path: RLHF (chat), RLVR (reasoning, slower/costlier), RLCD (calibration-oriented RL) — Jev uses RLCD to output calibrated judgments and probabilities, not text.
Background: official AI primer states RLHF co-invented by Diogo Almeida (RLHF used for InstructGPT/ChatGPT), later TypeSafe co-founder/CEO; DCVC public materials note his OpenAI experience. Cited as “official docs say” / “investor materials say,” not independently verified.
Atomic Decomposition: The Core Concept
Official How-to-build page emphasizes: a broad question hides multiple judgments in one answer; atomic questions split them so you can inspect, tune, and combine in code.
Example — instead of asking “Is this email spam?” (one opaque number), ask independent narrow questions:
Does the email ask for passwords or login credentials?
Does it claim an unexpected prize, payment, or reward?
Does it create urgency to act quickly?
Does the sender display name organization contradict the email domain?
Does the link text mask or distort the true URL?Broad question → unauditable number. Atomic questions → separate signals you can observe, threshold, replace individually, then combine with weighted code logic. The “verify tool calls” example uses nine atomic questions (tool choice, city match, schema compliance, result ID match, date match, unit match…).
Key: decomposition does not add latency. Questions on the same state run in parallel, independently — no hidden cross-context (no context rot), “decomposition does not require more round trips.” Cookbook comparison: a GDPR wiki page (53,777 chars) + 13 questions in one request cost $0.000497, took 0.27s; 13 separate requests cost $0.006090, took 2.71s — 12.2× cheaper, 10.0× faster, answers unchanged (5 repeats, most probability std-dev = 0.0; noise comes from question difficulty, not batching).
Providing State
State = material for judgment (string, JSON object, array). Three official principles:
Only include context relevant to current questions.
Structure where possible (nested JSON, reference paths in questions with backticks, e.g., `support.tickets[0].message`).
Don’t mix questions into state — state holds content/facts, questions defined separately.
Hard limits: Jev only accepts text (images/audio/video must be converted first); context 64k tokens, state + single longest question ≤ 32k.
Four Official Patterns
Speculative Fan-Out : ask many questions (including speculative) in one request; code decides which to use. Benefits: cost, speed.
Confidence-Gated Routing : use confidence as a second decision axis. Benefits: reliability, safety.
Composite Scoring : combine multiple dimensions into one score while keeping individual judgments. Benefits: cost, reliability, speed.
Intent Routing : classify intent first, then route to best handler. Benefits: cost, speed.
Intent Routing example: incoming support message → one request asks Choice (intent: order status / product inquiry / return / complaint) and Score (complexity). Code decides : intent.confidence < 0.5 → human; order status → pure code query; product/return → specialist LLMs; complaint → check complexity score + its confidence, else human. Remember: classification by model, routing by code.
Confidence: Not Probability, Not Accuracy
Choice and Score return full probability distributions; confidence compresses the distribution’s “shape” into 0–1 (1.0 = all mass on one option, lower = flatter). It’s a convenience statistic; you can compute your own from probabilities. Noul has no confidence — it is already a 0–1 probability.
Official starter pattern: three tiers — high confidence → auto-execute; medium → cautious proceed (human confirm, flag review, gather info); low → don’t act (escalate, clarify, fallback). Threshold is not a single number; different actions in the same system get risk-based tiers. Example: confidence < 0.5 always human; above that, reversible actions (balance check) execute directly, high-risk (approve transfer) requires confidence > 0.9.
Honest boundary: calibration is a cross-sample property, does not guarantee single-answer correctness. Before production, plot confidence vs. accuracy on your data to set thresholds.
Cookbook experiment: 60 10-K filings, 75-option Choice for industry group. Split at confidence ≥ 0.9 — high-confidence half: 27/30 correct; low-confidence half forced specific answer: 12/30 correct, but allowed to answer coarser parent category → useful answers rose to 48/60. At low confidence, “answer coarser” beats “force specific,” and that logic lives in code, no second model call.
Five Real-World Use Cases (Official Cookbook Numbers)
All numbers below from official cookbook with conditions (model version, dataset, sample size, repeats) noted; they are results in official experimental environments, not universal guarantees.
Classify then route : one Choice + one Score for intent + complexity; code routes to “pure code query / specialist LLM / human,” so expensive resources only used where needed.
Moderation & trust & safety : self-consistency experiment on a borderline post with 8 Choice moderation questions, repeated 15 times: label consistency 90.8% ; add code rule — if max probability < 0.60 output “uncertain” — consistency rises to 99.2% , while 74.2% of answers still auto-decide (latency 114ms, cost ~ $0.000046 per run). Author stresses: this experiment does not measure accuracy. Another guardrails case: same scores ( jailbreak=0.74, severity=0.51), strict policy → block, lenient → review. Official line: “TypeSafe provides evaluation, your application owns the decision.”
Verify other AI — cheap narrow judgments check expensive generative models:
Check citations : RFC 7519 (58,365 chars, 45 sections) tested with 8 citations (4 deliberately corrupted). Result: 4 verified, 1 fabricated, 1 contradicted, 2 unsupported. The citation not in source at all needs no model — string match in code suffices.
Filter RAG passages : 12 candidate passages, vector similarity clustered at 0.584–0.455, couldn’t separate “corrects premise” from “attempts hijack.” Switched to 4 narrow Noul (on-topic?, usable for direct answer?, conflicts with premise?, attempts manipulation?) plus explicit code thresholds (e.g., injection ≤ 0.70). Injection passage had relevance pass but injection score 0.99 → dropped. Author reminder: this is not a security boundary.
Verify tool calls : the nine atomic questions above act as a health check on agent execution traces.
Retrieval & reranking : on CLERC legal retrieval dataset, BM25 recalls 30 candidates (100% contain correct answer, but correct answer at rank 1 only 5%). Rerank with one Noul per query-candidate pair → Top-1 from 5% to 18%, Top-10 from 38% to 62% (1,200 calls cost $0.0645). Reason: generative models are extra time/cost for “just need a number” tasks, and repeated calls may give inconsistent scores.
Batch extraction & features :
Cascade : small model extracts, then Jev verifies field-by-field with Noul (“is this field fabricated?”, “is it off-topic?”). Only on hit upgrade to reasoning model. Example: small model passed JSON schema but content was hallucinated; Jev gave hallucinated 0.95, upgrade corrected field to empty string — “turn a confident false field into an honest empty field,” paying reasoning-model cost only for that field.
Feature extraction : autoresearch experiment on 2,000 wine-tasting notes predicting score (80–98). Held-out 800 RMSE: mean baseline 3.09 , bag-of-words 2.47 , single 10-level Score 2.15 , 18 atomic questions → 1.87 , after 5 iterations 1.77 . Same text, multiple narrow questions yield more usable signals than one overall score; these numeric features feed a classic ML model, which does the final prediction.
Integration Checklist: Six Steps
Ask if code can decide directly. Arithmetic, calendar, field existence, explicit markers, structured data reads — all stay in code. Model only for “questions code cannot answer from text alone.”
Decompose broad judgments into atomic questions. One attribute per question; include criteria/boundaries in question (official allows instructions / criteria as structured objects defining coverage, non-coverage, examples).
Ask multiple questions in one request. Same state, parallel evaluation, no interference; decomposition adds no round trips.
Combine in code. Deterministic rules or weighted sums; official also supports feeding probabilities as features to downstream classic models.
Gate by confidence, thresholds per risk. High-risk irreversible actions demand higher confidence; low confidence always escalate to human first.
Log and calibrate. Record model version, answer, probabilities, final outcome. For long-running experiments, pin exact version (not jev-latest).
When Not to Use Jev
Open-ended generation tasks : copywriting, explanations, code generation, long-document summarization — still generative model territory. Official cookbooks keep LLMs for generation: “TypeSafe scores paragraphs and tags routes, answer still written by LLM”; cascade just “spend big money only when needed.”
Rules, arithmetic, calendar solve it reliably : don’t bring a model.
Irreversible error nodes : must keep human confirmation — same scores under different policies lead to different actions; decision logic must stay in your code.
Very low call volume or no evaluation data : accumulate data and evaluation method first, then discuss thresholds.
Non-English content requires caution : official states English is primary training language; other languages (including CJK) “work but accuracy not equivalent” — must self-test on your content and factor confidence into routing.
Hard limits : text input only; 64k context; not a security boundary (official emphasizes this in injection-filtering cookbook).
Conclusion: Copy the Decomposition Mindset, Not Just an API
Treating Jev as “cheaper LLM” yields only cost optimization. Treating it as an interface contract gives you a writing style: first ask “is this judgment a finite choice?” then “can code decide it?” — only the remainder goes to the model, and even then as a sufficiently narrow question.
Four portable questions: Which nodes in my system are just finite choices? Which judgments can be split into independent narrow questions asked at once? Which nodes must have human sign-off? Which errors must never be automated? Answer those clearly, and the model becomes a replaceable decision backend.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
