Reproducing Jev's Structured Decision Engine for $0.19: 400 Lines of C++ in llama.cpp

The article details reproducing TypeSafe AI's Jev prototype—a single-forward-pass structured decision engine—using Baidu Qianfan Token Plan, implementing 400 lines of C++ in llama.cpp to achieve 70ms parallel decisions with calibrated probabilities, costing only 1.395 yuan (313.9 credits), and shares prompt engineering practices for AI-assisted development.

Baidu Geek Talk
Baidu Geek Talk
Baidu Geek Talk
Reproducing Jev's Structured Decision Engine for $0.19: 400 Lines of C++ in llama.cpp

Background: Jev and Structured Decision Engine

In September 2026, TypeSafe AI released "System One Model"—Jev. Unlike traditional LLMs that generate text autoregressively, Jev outputs structured decisions with calibrated probabilities via a single forward pass, supporting parallel sampling with ~70ms latency. This avoids the "generate→parse→validate" loop, offering a new approach for classification, routing, and moderation.

Reproduction Using Baidu Qianfan Token Plan

The author used Opencode (based on DeepSeek-V4-Flash) within Baidu Qianfan Token Plan personal tier to reproduce Jev's core mechanism in llama.cpp. The entire process—concept understanding, source code exploration, compilation/debugging, and article writing—was autonomously performed by AI. Total cost: 313.9 credits (≈1.395 CNY / $0.19).

Why Token Plan Suits Exploratory AI Development

AI development involves iterative trial-and-error: reading dozens of files (many irrelevant), multiple compilation rounds, and article revisions. Under pay-per-call pricing (e.g., GPT-4o), such exploration could cost >20 CNY. Token Plan's prepaid 200 CNY for 45,000 credits eliminates cost anxiety, enabling unrestricted experimentation.

Implementation Details

Core Mechanism Transformation

Traditional LLM mode: For multi-question scenarios, requires "token-by-token generation → text extraction → regex parsing", slow and prone to format collapse.

Jev decision mechanism: Breaks autoregressive generation; packs 5 business questions into one batch, applies Softmax on the last token's logits for each sequence, completing all inference and probability calibration in a single llama_decode() call without generating any text. Response time reduced to ~70ms.

The AI located the relevant code in llama.cpp and wrote ~400 lines of C++, fixing bugs such as multi-sequence logits position index offsets.

Standardized Input/Output

Input: Standard JSON configuration containing scene state (e.g., "customer refund, damaged packaging, 3-year member") and list of decision questions.

Output: Structured JSON returning all decisions (e.g., escalate / refund) with confidence probabilities.

Experimental Results

Tested with Qwen3-0.6B-Q8_0.gguf on a customer-service scenario (5 questions). One inference, zero string generation. Each decision included probability scores. Lowering temperature from 0.8 to 0.3 increased deny confidence from 62.5% to 89.4%—lower temperature sharpens distribution, making model appear more confident.

Critical distinction: Confidence ≠ accuracy. Confidence reflects model's certainty about its own output; lower temperature only sharpens distribution. No labeled test set existed to compute accuracy. For questions with clear signals (churn risk, priority), decisions matched intuition at reasonable temperatures. For subjective judgments (customer sentiment), confidence was low (55.9%), which itself is a useful signal for downstream routing to human review.

Gaps from Original Jev

Jev employs specialized model architecture and RLCD training for structured decisions. This reproduction simulates the mechanism on an existing LLM. However, the core loop—single forward pass, parallel decisions, probability output—is functional. Validating feasibility cost only 1.395 CNY.

Prompt Engineering Practices for AI-Assisted Development

The interaction with the AI agent followed a structured three-phase workflow:

Confirm foundations: Provide all materials (blog HTML, GGUF model path) and explicit three-step goals (organize paths, modify llama.cpp for Jev). Key: give both concept source and model path; let agent explore autonomously.

Acceptance testing: After agent claims completion, define test criteria (e.g., specific customer scenario with emotion classification). This exposed issues: on 0.6B model, label-token probabilities were near-random (~0.25 each). Agent then implemented permutation debiasing and continuation scoring.

Result correction: Human reviews and corrects technical conclusions. Example corrections: "Zero hallucination and high-confidence Jev cannot be guaranteed; Jev only ensures type safety." "Lower temperature raises confidence but not accuracy; without test set, accuracy cannot be computed."

Core principle: Agent executes, human judges. Agent excels at reading source files, writing C++, compiling/debugging, but technical accuracy requires human oversight—especially for concepts like "zero hallucination" and "calibrated probabilities" that agent may copy verbatim from source without engineering precision.

Conclusion

Jev's insight: many AI applications need structured decisions, not text generation. Traditional LLMs take a detour via text generation; Jev extracts decision probabilities directly from logits, skipping string generation/parsing overhead. Jev guarantees type safety (answers within predefined options) but not correctness. Confidence ≠ accuracy; it measures model's self-certainty. Real value: software can route based on confidence—low confidence → human review, high confidence → auto-process.

Baidu Qianfan Token Plan's value: reduces exploratory development trial-and-error cost to near zero. 1.395 CNY validates a cutting-edge AI concept—no need for large API budgets or per-call hesitation. This enables letting AI explore freely.

AI-agent interaction workflow
AI-agent interaction workflow
Test results: 5 questions, confidence scores at different temperatures
Test results: 5 questions, confidence scores at different temperatures
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI-assisted developmentPrompt EngineeringC++llama.cppJevBaidu Qianfan Token Plansingle forward passstructured decision
Baidu Geek Talk
Written by

Baidu Geek Talk

Follow us to discover more Baidu tech insights.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.