Jev: The 40-200x Faster 'System 1' Model Reshaping AI Agent Architecture

Jev is a new structured judgment model from TypeSafe AI that delivers Choice, Score, and Noul outputs at 70-500ms latency and 40-400x cost savings versus generative LLMs, enabling high-frequency decision primitives for browser automation, game AI, trading, agent supervision, and large-scale classification with confidence-driven routing between Jev and LLMs.

Fighter's World
Fighter's World
Fighter's World
Jev: The 40-200x Faster 'System 1' Model Reshaping AI Agent Architecture

Jev, released by TypeSafe AI (founded by former OpenAI researcher Diogo Almeida), is a "System One" model that does not generate text but performs structured judgments: Choice (selection), Score (rating), and Noul (yes/no probability). It operates at 70–500 ms latency, with input cost around $0.042 per million tokens and free output, claiming 40–200x speed and 40–400x cost advantages over mainstream generative LLMs.

Use Cases

1. Browser/Computer Use Automation

Browser Use Ultrafast : Jev selects the next DOM action (click, input) each step; a small LLM is called only when text is needed. The flagship demo searches Google Flights from Zurich to London in ~7.1 seconds at $0.0039.

Voice-controlled browser : Speech → transcription → Jev decides action → browser executes; single decision ~$0.0002, ~300 ms.

Computer Use acceleration : Stagehand combines accessibility trees, Cua restricted action menus, Mac desktop OCR + Jev clicks, parallel browser adversarial testing, Amazon product filtering.

2. Email/Ticket Classification & Routing

Classify 1,700 emails simultaneously for category, priority, spam probability, and reply need; total cost ~$0.18.

Customer service: per email judge department, urgency, refund risk, human escalation; high confidence enables full automation.

3. Real-time Game Decisions

Official/community Doom demo: ~10 Hz calls, ~$7/hour.

Super Mario: extracts object JSON directly from emulator RAM, no screenshots needed.

Millisecond decisions: Tetris, Wikiracing, 1v1 FPS (~9 Hz for move/aim/shoot), Subway Surfers parallel. Games serve as the most intuitive "millisecond decision" showcase.

4. Trading/Quant Decisions (High Risk)

jev-trader : Jev judges buy/sell every block (~300 ms) on Monad chain Kuru order book; model latency ~81 ms.

Liquidity pool toxicity detection, mean reversion (mostly advisory mode).

5. Agent Supervision, Code Review, Security

Code review (jev-review) : staged risk judgment on Git diffs + local dashboard.

Agent supervision (Foreman, pi-warden) : check task completion/stuck/deviation, dangerous tool calls.

Context compaction : Jev decides keep/delete to avoid information loss from traditional summarization.

Security : prompt injection detection (self-reported 96.5% accuracy), tool-call firewall, content moderation.

6. Large-scale Data Classification & Filtering

News/content triage: ~1 second, $0.001 per day.

Other: children's snack scoring (3,000 items ~28 s, $0.11), PostgreSQL natural-language WHERE filtering, social media post filtering (clickbait/ads), paper title/abstract screening.

Other Promising Demos

Model/tool routing: decide which LLM or tool to call next.

Drones: camera state → Jev suggests action (safety guaranteed by code).

Real-time interaction: pixel-level drawing, NPC simulation.

Common Patterns Across Use Cases

The community treats Jev as a fast decision primitive inside software, collaborating with LLMs: Jev handles high-frequency narrow judgments, LLMs handle generation/planning. Cost and speed advantages make "run judgment on all data" feasible.

1. Core is High-frequency, Narrow-scope Structured Judgment

Scenarios rarely ask the model to write, code, or chat; they repeatedly answer "which one", "score", "yes/no". Jev acts as a software-internal fast decision primitive (Choice/Score/Noul) outputting calibrated probabilities consumed directly by code.

2. Extreme Latency and Cost Sensitivity

Browser ops, games, trading, real-time agent supervision need millisecond response (70–500 ms).

Email classification, massive data filtering, content moderation need massive parallel judgments; per-call cost must be ~$0.0001–0.001.

Existing LLMs are either too slow or too expensive; Jev fills this gap.

3. "Program-led + Jev Local Judgment" Architecture

Typical flow:

Program prepares current state (page elements, email, game state, code diff, market data).

Program asks Jev a set of predefined typed questions.

Jev returns choice/score/probability.

Program decides next action (execute, route, escalate, terminate) based on result and confidence.

Jev never drives global logic; it only performs atomic judgments. Complex orchestration, execution, retry, safety boundaries remain in code.

4. Frequent Collaboration with Traditional LLMs

Hybrid architecture:

Jev handles high-frequency micro-decisions (button selection, routing, filtering, scoring).

Traditional LLM invoked only for text generation, complex planning, or explanation.

This drastically cuts cost and latency while retaining generative capability.

5. Confidence as Key Control Switch

High confidence → auto-execute.

Low confidence → escalate to stronger model or human review.

Balances automation degree and risk control.

6. Natural Fit for Batch + Parallel

Multiple questions evaluated in parallel per call; output is free. Community uses Jev for:

Map-reduce style judgment on massive data.

Every step decision in agent loops.

Real-time high-frequency control loops.

Implications for Agent Applications

Treat Jev as a lightweight decision layer inside the Harness (execution framework), not as a replacement for the main LLM. Jev handles high-frequency, narrow, structured judgments; the main model handles planning and generation.

1. Harness Engineering (Most Watched Direction)

LangChain founder Harrison Chase noted: "Jev isn't meant for text generation, it's for simpler constrained output. Turns out, that can actually be very useful when building a harness!" Community uses Jev around the harness, not inside the agent core.

Typical usages:

Model Routing : At each turn start, Jev's Choice judges "cheap fast model vs strong model". LangChain's experimental ModelRouterMiddleware selects model by task complexity, significantly reducing cost. For multi-model agent products, intent decomposition and model routing can improve margins.

Skill/Tool Selection : Large skill libraries → Jev scores/selects all available skills, loads only relevant ones, drastically shrinking context. Hermes integrates Jev for skill selection, memory filtering, compaction selection.

Context Compaction : Keep/delete judgments (Noul) on tool-call history or memory blocks, avoiding traditional summary information loss. Caution: simple per-line filtering may break model's context understanding; needs thorough validation.

Stuck Detection / Completion Verification : Judge whether agent repeats same strategy, task truly completed, output supported by evidence.

Risk Guardrails : Pre-tool-execution Noul checks "dangerous/irreversible/deviates from intent"; high risk blocks or escalates. Projects: AutoModeMiddleware, AgentGhost, pi-warden.

Practice tip : Jev only judges; code owns thresholds and side effects. Low confidence auto-escalates to stronger model or human.

2. Intent Classification

Ideal for Jev's Choice/Noul.

Customer service: incoming request → Jev asks "account query? product issue? policy interpretation? high-risk/ambiguous?" → route accordingly. Previous attempts at formalized planning/routing now have a viable tool.

Parallel questions: urgency, human needed, spam, etc.

Semantic routing: Hono + Jev middleware defines routing rules in natural language instead of hard-coded paths.

If reliability holds, benefit is "millisecond + ultra-low cost", orders of magnitude cheaper than generative LLM classification.

3. Routing

Includes task routing, model routing, sub-agent routing, exception routing.

Task/request routing: dispatch to different handlers/workflows/humans.

Multi-agent routing: parallel coding agents → decide which agent takes subtask, or confidence-based majority-vote permission control.

Cascade routing: simple request → cheap model; complex/low confidence → strong model or human.

Common community pattern: "Cascade" – Jev does fast triage, then decides whether to call expensive model.

4. Tool Strategy Selection

Pre-tool-call: Choice from closed tool set picks best tool, or judges "need tool?".

During execution: pre-execution risk check – irreversible? matches user intent?

Post-execution: judge result sufficiency, retry need, continue decision.

AgentGhost, various MCP wrappers package Jev as "semantic firewall before tool call", outputting ALLOW/ASK/DENY.

Simple Principles for Leveraging Jev

Only atomic, bounded judgments : design questions as predefined options; don't let Jev generate text or do complex reasoning.

Ask multiple questions in parallel : single call yields intent, risk, complexity, tool-need, etc.; cost barely increases.

Confidence-driven strategy : high confidence auto-execute, low confidence escalate – key to harness reliability.

Code retains ultimate control : Jev returns probabilities and choices; program decides thresholds and actions.

Explicit division with main LLM : Jev = fast-reacting System 1; generative main model = deep-thinking System 2.

Start exploration at harness boundary layers : entry intent classification, per-step tool guards, model routing, completion checks – likely highest ROI.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agentsbrowser automationmodel routingagent harnessJevTypeSafe AIstructured judgmentconfidence thresholds
Fighter's World
Written by

Fighter's World

Live in the future, then build what's missing

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.