Jev Engineering: Moving Judgment Tasks Out of LLMs for Faster AI Agents
This article explores Jev (System One Model) as a specialized probability judgment layer for AI agents, detailing its three primitives (Choice, Score, Noul), the fast-jev-compaction project for context pruning, and a demo implementation showing 91.5% context reduction while preserving critical debugging details like file paths and error codes.
What Is Jev: Turning Generation into Judgment
Jev is TypeSafe AI's System One Model. Unlike chat models that generate text, Jev deliberately abandons free-form generation and returns only typed probability decisions. Its input can be heterogeneous: user goals, page state, tool result summaries, business fields, candidate actions. Its output is narrow: which option, what score, or the probability that a proposition holds.
Jev exposes three primitives:
Choice (单选题): Output shape — Option + per-option probability + confidence. Typical use cases: Tool routing, action selection, classification.
Score (打分题): Output shape — Continuous score + probability distribution + confidence. Typical use cases: Risk scoring, quality scoring, relevance scoring.
Noul (判断题): Output shape — 0–1 probability. Typical use cases: Whether to retain, whether dangerous, whether goal achieved.
For example, given a test log, a traditional LLM might output an explanatory paragraph. Jev returns a structured probability:
{
"answers": {
"result_t3": {
"noul": 0.87
}
}
}This can directly drive a code branch:
if (keepResult >= 0.5) {
keepFullResult();
}The core shift: model output moves from "text" to "executable probability judgments." Unlike GPT/Claude structured output (which still generates tokens), Jev is trained from the start for closed-output-space probability judgments.
This enables a four-layer agent architecture:
Slow Thinking — Responsibility: Planning, explaining, generating, complex reasoning. Typical components: GPT / Claude / Gemini.
Fast Judgment — Responsibility: Routing, filtering, scoring, gating. Typical components: Jev / Jev-like models.
Deterministic — Responsibility: Permissions, state, side effects, rollback. Typical components: Ordinary code.
Fallback — Responsibility: High-risk or low-confidence handling. Typical components: Human / stronger model.
Jev is not a "cheaper GPT"; it suits high-frequency, closed-output, semantic-understanding micro-judgments inside an agent — the reflex layer, not the brain.
Why Agents Need a Fast Judgment Layer
Realistic coding-agent sessions accumulate massive context: user constraints, assistant reasoning, tool calls, tool results, error logs, file snippets. The biggest context bloat comes from tool results.
User request — Often long? No. Must keep full? Yes. Typical risk: Dropping loses constraints.
Assistant reply — Often long? Medium. Must keep full? Usually. Typical risk: Dropping loses plans/commitments.
read_file result — Often long? Yes. Must keep full? Depends. Typical risk: Full file may be stale or re-readable.
search_content result — Often long? Yes. Must keep full? Depends. Typical risk: Many matches are just exploration traces.
execute_command log — Often long? Yes. Must keep full? Depends. Typical risk: Failure stack matters; success logs may be useless.
list_dir result — Often long? Yes. Must keep full? Usually not. Typical risk: Directory trees consume space easily.
Traditional summarization loses precise details (file paths, error codes, line numbers, auth headers) that are critical for subsequent debugging actions. Using a general LLM for compression hits three mismatches:
Generation cost vs. judgment need: Many tasks only need a 0–1 probability (keep this log?), not a generated explanation.
Summaries break verifiability: Tool results are evidence; summarization rewrites exact tokens into approximate descriptions.
Confidence doesn't enter code branches: LLMs may say "I'm 80% sure" in text, but that's uncalibrated prose. Jev returns calibrated probabilities as API values so code can threshold actions.
Probability bands and system actions:
>= 0.8 — Auto-execute. Note: Low-risk, high-repeat scenarios.
0.5 – 0.8 — Conservative handling. Note: Truncate, keep summary, request more info.
< 0.5 — Drop or rollback. Note: Low-value content cleaned; high-risk escalates to human.
Probabilities must be calibrated with shadow-mode evaluation on own data before production use.
fast-jev-compaction: No Summarization, Only Retention Decisions
fast-jev-compactionis a GitHub project (Claude Code plugin / npm library) that applies Jev to a concrete problem: when a coding agent's history grows too large, which tool calls and results can be deleted? Its strategy is restrained:
User text untouched.
Assistant text untouched.
Tool calls and results processed in pairs.
Old tool results either kept whole, truncated, or deleted.
Fallback to traditional summary if Jev fails or compression gain is insufficient.
Instead of rewriting history into a summary, it splits compression into two judgment questions per tool-call pair:
keepCall: Does this tool call still matter?
keepResult: Does this tool result full text still need to be kept?The two probabilities yield three actions: keep both, keep call but truncate result, or drop both. This preserves exactness — critical tokens (paths, error codes) stay verbatim; low-value bulk (directory trees) gets truncated or dropped.
Demo Implementation Walkthrough
The demo ( src/compact.js) follows six steps:
Collect tool_use and tool_result from history.
Pair them by tool_use_id.
Pin first message and most recent N messages (protected from deletion).
Build a compact state for Jev.
For each non-pinned call, ask two Noul questions.
Rebuild message list based on keepCall / keepResult thresholds.
4.1 Tool Call Pairing
Assistant emits tool_use; user returns tool_result with matching toolUseId. Pairing prevents orphaned calls or results. The collectToolCalls function (in src/state.js) produces unified objects with fields: id (t1, t2…), toolUseId, tool, input, callIndex, resultIndex, resultChars, isError, pinned.
4.2 State Construction: Only What Jev Needs
Full tool results are not sent to Jev. Each result becomes a one-line note: ok, 4213 chars (omitted) or error, 830 chars (omitted). The historyEntries function (in src/state.js) builds entries like:
{
id: call.id,
tool: call.tool,
input: truncate(safeJson(call.input), inputChars),
result: call.resultNote,
}Mapping of original content to state:
Full file content → Not included; only tool name + input.
Full test log → Not included; only error, N chars.
Tool input params → Truncated but retained.
User text → Retained; long text may be abridged.
Recent messages → Prioritized for retention.
4.3 fitState: Graceful Degradation Under Token Budget
If the constructed state exceeds maxStateTokens, fitState applies a staged degradation chain (implemented in src/state.js):
full — Tool inputs up to 1000 chars. Loss level: Low.
inputs<=200 — Tool inputs reduced to 200 chars. Loss level: Low.
inputs<=60 — Tool inputs reduced to 60 chars. Loss level: Medium.
texts abridged — Long texts keep head/tail. Loss level: Medium.
old messages collapsed — Old text folded into ellipsis note. Loss level: High.
old calls compacted — Old tool calls compressed to one line. Loss level: High.
old messages left out — Delete old pure-text messages. Loss level: Very High.
old calls merged — Merge consecutive old tool calls. Loss level: Very High.
Order: weaken detail first, then drop content; protect recent messages first; preserve task continuity over compression ratio. Telemetry fields ( stateTokens, stateStage, requests) make the process observable.
4.4 Two Nouls → Three Compression Actions
Jev principle: one question, one judgment. Instead of a compound question, split into two Noul s:
Does knowing this tool call (with its inputs) still matter for the rest of the task?
Does the full text of this tool result still need to be retained?
Decision logic ( decideCall in src/compact.js):
keep — Condition: keepResult >= threshold. Handling: Retain call and result fully.
drop_result — Condition: keepResult < threshold AND keepCall >= threshold. Handling: Keep call, truncate result.
drop_call — Condition: Both below threshold. Handling: Delete call and result.
4.5 Batching and Fallback
When many calls exist, batchCalls splits questions across requests, each carrying the full state (simple, same context per judgment). Fallback strategies:
Jev request fails → Fall back to built-in summary or skip compression.
Malformed response → Discard this compression round.
Insufficient compression gain → Keep original transcript.
State cannot fit → Increase budget or switch to traditional compression.
Low confidence → Conservatively retain; avoid aggressive deletion.
The demo abstracts the model behind an asker interface: real Jev via TYPESAFE_API_KEY, or a local heuristic mock for offline testing. Model swappable; decision protocol stable.
Demo Results
Sample session: fixing an order-export timeout. Four tool calls:
toolu_001 — Tool: list_dir. Result: Long directory tree. Relevance: Call useful, full text not.
toolu_002 — Tool: read_file. Result: orderExport.ts with getAuthToken. Relevance: Useful.
toolu_003 — Tool: execute_command. Result: Test failure with timeout + auth header. Relevance: High value.
toolu_004 — Tool: read_file. Result: Unrelated legacy file. Relevance: Irrelevant.
Run: npm test (11 pass) then npm start. Output:
id tool action keepCall keepResult
-- --------------- ----------- -------- ----------
t1 list_dir drop_result 0.54 0.00
t2 read_file keep 0.59 0.69
t3 execute_command keep 0.75 1.00
t4 read_file drop_call 0.47 0.00Compression metrics:
Characters: Before 10029, After 852, Change: -9177 (91.5% reduction).
Tool calls: Before 4, After 3 visible. Change: 1 irrelevant call deleted.
Fully retained: 2 (Key file + failure log kept).
Truncated results: 1 (Directory tree head only).
Matches expectations: directory tree truncated (re-runnable), orderExport.ts kept (auth token relevance), failure log kept (timeout + auth header), legacy file dropped.
Selection Guide: Jev vs. Structured Output vs. Traditional Classifiers
LLM + JSON Schema — Pros: Expressive, explains complex rationale. Cons: Still token-by-token generation; high cost/latency. Best fit: Low-frequency complex judgments.
Tool Calling — Pros: Integrates with tool system. Cons: Decision still relies on generative model. Best fit: Main-loop agent actions.
Traditional Classifier — Pros: Cheap, local deployment. Cons: Rigid labels, weak semantic generalization. Best fit: Stable labels, high-volume labeled data.
Jev — Pros: Closed output, usable probabilities, high-frequency ready. Cons: No generation, weak complex reasoning. Best fit: High-frequency local judgments.
Quick checklist:
Can output options be enumerated upfront? → Suitable for Jev.
Need natural-language explanation? → Prefer LLM.
High-frequency calls? → Suitable for Jev.
Errors tolerable with fallback? → Suitable for Jev.
Task stable long-term with ample labels? → Consider traditional classifier.
Requires cross-file multi-step reasoning? → Not for Jev alone.
Two common misreadings:
"No hallucination" → actually type-safe: it won't emit outside the schema (e.g., invent maybe_keep), but can still pick the wrong valid option.
"Probabilities are trustworthy" → calibration is a goal; public RLCD details and third-party calibration data are still limited. Engineering practice: shadow mode → measure real accuracy per probability bucket → set auto-execute thresholds → conservative handling for high-risk.
Risk-level gating:
L1 Read-only — Action type: Search, read, classify, rerank. Strategy: Can auto-execute.
L2 Reversible — Action type: Truncate context, create temp records. Strategy: Low confidence → conservative retain.
L3 Irreversible — Action type: Delete data, charge, revoke permissions. Strategy: Jev must not decide alone.
From Demo to Engineering Pattern
Community projects show Jev's real niche as an internal agent component:
Browser automation — Example project: browser-use/jev-ultrafast. Core idea: Number page elements; Jev picks next action + target element.
Agent framework — Example project: vercel/eve. Core idea: Use Jev for evaluation and routing inside framework.
Context pruning — Example project: fast-jev-compaction. Core idea: Judge whether tool calls/results should stay in context.
Code review — Example project: jev-review. Core idea: Decompose review into yes/no judgments and risk scores.
Local replication — Example projects: SemIf, kev, laya. Core idea: Replicate Jev-like interface with open small models.
Database extension — Example project: pg-jev. Core idea: Write semantic judgment conditions in SQL.
Game simulation — Example projects: Doom, Mario, Snake. Core idea: Per-tick or per-turn action selection.
Game demos spread fast, but fast-jev-compaction is more instructive: context pruning is universal for agents, and it cleanly demonstrates Jev's engineering role — not writing code, not running the agent, just judging "should this tool history stay?"
The article's central point: Jev's value lies not in replacing LLMs, but in extracting high-frequency, closed, fallback-able judgments out of the generation pipeline. This yields three shifts:
Efficiency — No longer spin up a full generation chain for every micro-judgment.
Capability — Routing, filtering, scoring, gating become stable unified interfaces.
Paradigm — Agent moves from "all LLM thinking" to "generation / judgment / execution" layered collaboration.
Conclusion
Jev isn't mysterious; its pieces (classification, constrained output, logit scoring, probability calibration, discriminative route) existed before. Its contribution is a clean developer interface that forces rethinking agent decomposition: what needs generation, what is just judgment, what belongs to deterministic code. fast-jev-compaction is a perfect entry point — it doesn't try to make Jev a generalist, but places it in a narrow, real spot: judging tool history retention. This pattern applies to all high-frequency, closed, fallback-able internal agent decisions.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Tencent Technical Engineering
Official account of Tencent Technology. A platform for publishing and analyzing Tencent's technological innovations and cutting-edge developments.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
