Jev: The 'Non-Chatting' AI That's 193x Faster & 444x Cheaper Than LLMs for Agent Judgment Tasks

This article dissects Jev, a specialized judgment model from TypeSafe AI that replaces LLM-based classification with single-forward-pass inference, achieving 70-500ms latency and $0.042/M input tokens (output free). It covers Jev's three output primitives (Noul, Choice, Score), benchmarks against DeepSeek and Claude, and real-world use cases in browser automation, game AI, agent guardrails, and model routing.

Tech Freedom Circle
Tech Freedom Circle
Tech Freedom Circle
Jev: The 'Non-Chatting' AI That's 193x Faster & 444x Cheaper Than LLMs for Agent Judgment Tasks

What Is Jev?

Jev is a System One Model — a judgment-only AI that does not generate text, chat, or write code. Given a state and predefined questions, it returns structured decisions: a choice with confidence, a probability (Noul), or a weighted score (Score). Its creator, Diogo Almeida (ex-OpenAI, co-inventor of RLHF), founded TypeSafe AI, which raised a $40M seed round led by DCVC. Jev launched September 15, 2026, opened fully on September 21 with a $5 credit (~120M tokens).

Why a "Non-Chatting" AI?

Agent workflows require dozens of micro-judgments per task (which button to click, which dropdown option, whether a command is destructive). Using a general LLM for each judgment introduces three problems:

Slow: LLMs generate tokens autoregressively. A classification that needs only "A or B" forces the model to emit 100+ tokens of reasoning. Benchmark: DeepSeek V4.1 Flash averages 11.3s for a customer-service classification; Jev averages 1.0s (11x faster).

Expensive: LLMs charge per input and output token. Jev charges $0.042/M input tokens, output free. Apple-earnings test: baseline LLM $0.0006–0.001 per judgment vs. Jev $0.000177 (3.5–5.6x cheaper). LangChain eval: 100 agent evaluations with Claude Sonnet 4.6 cost $28.17; with Jev $0.34 (83x cheaper).

Unpredictable output: LLMs wrap JSON in markdown, add explanatory text, or deviate from schemas. Jev returns only the requested primitive (option index, probability, or score), eliminating parsing logic.

Core Architecture: Single Forward Pass + Parallel Evaluation + Typed Output

Jev uses a bidirectional encoder (not autoregressive decoder). All questions for a given state are evaluated in one forward pass, producing logits for each primitive type simultaneously. This yields 70–500ms end-to-end latency (most calls ~100ms), comparable to a mobile-app tap response.

Three Output Primitives

Noul — Yes/No judgment. Returns a 0–1 probability. Example: "Will it rain tomorrow?" → 80%.

Choice — Pick one from N options. Returns the selected option plus a probability distribution over all options. Example: Menu ordering → most likely dish.

Score — Rate on a scale. Returns a weighted score plus per-level probabilities. Example: Restaurant rating → 4.3 stars.

Integration Example

import typesafe

result = typesafe.ask(
    state="User says: I was charged twice, want a refund",
    questions=[
        {"type": "choice", "question": "Which department?",
         "options": ["Pre-sales", "After-sales", "Technical", "Billing"]},
        {"type": "score", "question": "How angry is the customer?",
         "scale": [1,2,3,4,5,6,7,8,9,10]},
        {"type": "noul", "question": "Explicit refund request?"}
    ]
)
{
    "Which department?": {"choice": "Billing", "confidence": 0.97},
    "How angry is the customer?": {"score": 8.2, "confidence": 0.91},
    "Explicit refund request?": {"probability": 0.97}
}

Community Projects (Days After Launch)

Browser automation (Jev Ultrafast): Books a flight in ~7s. Jev decides click/input/select/wait each step; a small LLM fills text fields. Code loop queries Jev for next action and target element.

Game AI (Minecraft/Doom bots): Reads structured game state (distance, health), queries Jev every few tens of ms for action (flee, attack, hide, gather). Doom bot: ~10 queries/sec, estimated $7/hr.

Agent guardrails: Pre-execution safety check for shell commands. Jev classifies command as read-only, reversible, or destructive; low confidence triggers human review. Example: ambiguous rm -rf scored 0.33 confidence for irreversible → flagged.

Model routing: Jev rates query complexity (simple/medium/complex); simple → small model, complex → large model. Cuts cost by avoiding heavy models for trivial queries.

Why This Moment?

The industry is shifting from "make models smarter" (parameter scaling) to "specialized division of labor." Agents need thousands of fast, cheap judgments per workflow. Jev fills the judgment layer, letting large models focus on generation and reasoning. TypeSafe's $40M seed signals investor belief in "decision models" as a new category — analogous to phone makers shifting from raw performance to battery life.

Key Metrics at a Glance

Latency: 70–500ms (typical ~100ms)

Pricing: $0.042/M input tokens, output free

Free tier: $5 credit ≈ 120M tokens

Adoption: 38M+ announcement views, 140k developers in 3 days, ~500 open-source projects, 13% of Vercel paid teams active within 24h (2x GPT-5.6 day-one, 6x+ Claude Fable 5.1)

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

agent architecturemodel routingJevSystem One ModelTypeSafe AInon-autoregressive inferenceAI performance optimizationjudgment models
Tech Freedom Circle
Written by

Tech Freedom Circle

Crazy Maker Circle (Tech Freedom Architecture Circle): a community of tech enthusiasts, experts, and high‑performance fans. Many top‑level masters, architects, and hobbyists have achieved tech freedom; another wave of go‑getters are hustling hard toward tech freedom.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.