Reverse-Engineering Jev: 10K API Calls Expose Closed-Source Model Architecture
An independent researcher reverse-engineered TypeSafe's closed-source Jev classification model using 10,000 API calls, revealing its architecture uses a classification head with shared prefix inference, listwise scoring, and exceptional calibration (ECE 0.0313), challenging assumptions about black-box security.
Background: Jev, a Closed-Source Classification Model
In September 2026, TypeSafe released Jev, a "System One decision model" designed for structured tasks like classification, routing, and risk control. Unlike chat models, Jev does not generate text; given a state text and a set of questions with allowed answers, it outputs a probability distribution over the options (e.g., payment 91%, account 6%, other 3%). TypeSafe published no weights, no technical paper, only a promotional blog and API documentation.
Fundamental Design Flaw in Autoregressive Models for Decision Making
The article first explains why standard autoregressive LLMs (ChatGPT, Claude) are ill-suited for calibrated decision making. When asked for a confidence score, they produce a textual token like "90%" based on language patterns, not a mathematically derived probability. The probability 0.9 is never computed internally; the model merely predicts the next token that looks like a confidence value. Feeding such text into downstream risk systems amounts to "garbage in, garbage out."
Jev takes a different approach: after the forward pass, a linear layer plus softmax (a classification head) directly emits probabilities from the final hidden vector h via z = Wh + b, then softmax. No token-by-token decoding occurs.
Black-Box Probing Methodology: 10,000 API Calls
Researcher archerhume systematically probed Jev over two weeks, varying inputs and observing three signals:
Latency curves – how response time changes with input length and question count.
Token counts – the output_tokens field in API responses.
Refusal rules – how the model behaves when options are reordered or added.
This black-box probing technique, previously used to guess GPT-3's tokenizer and layer count, was applied systematically to a newly released commercial model, yielding a full architecture reconstruction report.
Evidence That Jev Does Not Decode Autoregressively
The output_tokens field appeared to report generated token count, but experiments contradicted this:
A question with 200 options returned output_tokens=1911, yet latency was identical to a 2-option question (tens of milliseconds). Autoregressive generation of 1911 tokens would take seconds.
Changing an option's probability from 0.0 to 0.01 (adding a character) left output_tokens unchanged.
Conclusion: output_tokens is a billing metric computed from the serialized response size, not a decode step count. Jev finishes inference after the prefill phase; no autoregressive decoding occurs.
Shared Prefix Inference (Shared-State Architecture)
Jev's API accepts a shared state (context) and a list of questions. The researcher tested whether the state is re-encoded per question or once:
Fixed state, varied question count from 1 to 1,500. Latency rose from 86 ms to 610 ms – answering 1,500 independent questions in ~0.6 s.
If each question re-processed the state, latency would be minutes.
Token limits: 65,536 total tokens per request, 32,768 per branch. This matches a shared-prefix design: a 23,000-token state plus 5,000 questions fits in one request.
This matches shared-prefix inference (also called shared-prefix attention), where the KV cache for the common prefix is computed once and reused across branches. Academic precedents: Hydragen (2024) restructures shared-prefix attention into matrix multiplications; DeFT (2024) optimizes tree-structured inference with FlashAttention. Jev's latency curve shape aligns with these theoretical expectations, indicating production deployment of recent research.
Listwise Scoring: Options Influence Each Other
Jev's probabilities are not independent per option (pointwise). Adding an irrelevant option changes relative odds among existing options:
Base case: four options (bank, provider, customer, unknown). Log-odds ratio customer vs. unknown = +0.38.
Add fifth irrelevant option ("bad weather"). Log-odds drops to +0.11 (average change -0.28 across 10 random groups).
If scores were independent, the ratio would be invariant to other options. The shift proves options attend to each other via the attention mechanism – a listwise scoring paradigm common in learning-to-rank.
Position sensitivity also matters: placing a reference card at the end of the option list yielded 16/16 correct; in the middle 11/16; at start 12/16. Moving reference info into the shared state achieved 48/48 correct. This is a critical deployment risk: option order changes may require threshold recalibration.
Calibration: The Critical Metric for Decision Systems
For decision-making, calibration (reliability of predicted probabilities) matters more than accuracy. Expected Calibration Error (ECE) measures the gap between predicted confidence and empirical accuracy.
Tested on 1,200 MMLU questions: Jev's ECE = 0.0313.
Typical open-source models show ECE 0.1–0.3.
TypeSafe calls their training method RLCD (Reinforcement Learning for Calibrated Decisions). Using a proper scoring rule (e.g., Brier score, log loss) as the loss function mathematically guarantees that minimizing expected loss yields calibrated probabilities. However, calibration can degrade due to finite data, model capacity, or distribution shift (Guo et al., 2017 showed modern neural nets can be miscalibrated despite high accuracy). Jev's calibration varies by task: 3-digit multiplication (86.7% accuracy, avg max prob 0.83) vs. two-step word problems (32% accuracy, avg max prob 0.30) – the latter shows appropriate uncertainty but overall probability distribution remains slightly optimistic.
The confidence Field Is Not a Learned Confidence
Jev's API returns a confidence field. TypeSafe's public Python adapter reveals it is a deterministic transformation of the output distribution: c = (p_max - 1/K) / (1 - 1/K) where p_max is the highest probability and K the number of options. This measures how peaked the distribution is relative to uniform, not the probability of being correct. A sharp distribution can be confidently wrong; a flat one may honestly reflect ignorance. The field is useful for internal thresholding but must not be presented to users as a correctness guarantee.
Speculation: Sparse Mixture-of-Experts (MoE) Backbone
Based on latency and performance, archerhume hypothesizes Jev uses a sparse MoE architecture:
~160 ms for 30,000-token input. A dense 70B model on 8×H100 would need ≥1 s.
MMLU-Pro accuracy 84.6% suggests a large model, not a small one.
Speed + knowledge capacity ≈ MoE with ~10B active parameters.
MoE's usual bottleneck – memory bandwidth during autoregressive decoding – disappears because Jev only does prefill. This remains unproven without a technical report or more extensive probing (e.g., 100k calls).
Limitations: Jev Is Not a Chat Model
Jev cannot converse, generate code, summarize long documents, or perform multi-step reasoning. It is a pure judgment engine. It does not compete with GPT-4/Claude; it serves a different niche (classification, routing, risk control) where it can be 10× cheaper and more accurate. Engineering evaluation should prioritize latency, throughput, and calibration over benchmark scores.
Three Takeaways for Practitioners
Ask for calibration proof. When a vendor claims "AI decision making," demand calibration tests (e.g., 1,200 samples) in the contract. Don't accept marketing talk.
Closed source ≠ black box ≠ security. TypeSafe's secrecy was pierced by 10k API calls. Real moats are data, domain scenarios, and iteration speed – things black-box probing cannot extract.
Measure non-star metrics. Jev's 84.6% MMLU-Pro is less relevant than its 65k token batch capacity, 0.0313 ECE, and sub-second latency. Quantify latency, throughput, and calibration for any AI tool.
Open Question: The Tokenizer
Archerhume tested 445 tokenizer probes against 192 public tokenizers; none fully matched Jev. Closest was Qwen (348/445 matches) but digit splitting and merge rules differed. OpenAI's o200k vocabulary direction matched but not exactly. TypeSafe likely uses a custom or modified tokenizer. This mystery may persist until a technical report or deeper probing emerges.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
