Rizzo Flow: Local Open-Source Implementation of Jev's Token-Free Decision Model

Rizzo Flow replicates Jev's System One Model interface locally using Spark-X2.5 and LoRA, outputting calibrated probabilities for typed decisions without token generation, achieving 50ms latency on consumer GPUs while acknowledging accuracy gaps versus the proprietary Jev service.

Geek Labs
Geek Labs
Geek Labs
Rizzo Flow: Local Open-Source Implementation of Jev's Token-Free Decision Model

Background: Jev System One Model

TypeSafe AI released Jev (System One Model) on September 15, 2026. Unlike conventional LLMs that generate text, Jev takes unstructured state (text, JSON, logs) and returns typed decisions with calibrated probabilities — yes/no, choice from options, score, or numeric value. Key differences from regular LLMs:

Output form: Pre-defined structure, no parsing/validation needed, claimed zero type errors.

Sampling: Single forward pass yields full probability distribution; TypeSafe claims two orders of magnitude faster than token-by-token generation.

Cost: Output tokens free; input $42 per billion tokens.

Confidence: Calibrated probabilities via RLCD (Reinforcement Learning for Calibrated Decisions), avoiding overconfidence typical of LLMs asked "how sure are you?".

Jev is a closed hosted service, requiring data egress and cloud pricing.

Rizzo Flow: Open Local Replication

Released September 21, 2026 by Rizzo AI Academy (maintainer Simone Rizzo), rizzo-flow is an Apache-2.0 Python project (810 stars as of 2026-10-04) that replicates Jev's /v1/systemone API endpoint and request/response schema. It uses the open-source Spark-X2.5 model (4B/1.7B, Apache-2.0) with a LoRA adapter trained on public data. The project explicitly states it is not affiliated with TypeSafe, does not copy Jev's proprietary architecture or RLCD training, and does not claim quality parity with Jev or SemIf. Probabilities are uncalibrated by default.

GitHub:

github.com/Rizzo-AI-Academy/rizzo-flow

How It Avoids Token Generation

The approach uses four steps:

Every question becomes multiple choice. Candidate answers map to single uppercase letters (A–Z). Tokenizer check ensures each letter is exactly one token, limiting to 26 options per question (Jev supports 255).

State processed once. State placed at prompt start, prefilled into KV cache. Subsequent questions branch from this shared cache — KV cells are shared, not copied.

Read only answer-letter logits. For each question, read one position and the allowed letters' probabilities; no sampling, no other tokens read.

Pure Python converts logits to structured output. Softmax, temperature, expected score, abstention policy, then schema-validated JSON.

Result: no decode loop, no output parsing, no JSON repair, construction-time type safety. Trade-off: "zero generated tokens" ≠ zero latency. Prefill and question processing still compute. On RTX 5060 Ti 8-bit: ~50 ms per short decision; 21 questions over 2000-token state ~1 second.

Snake Demo: Real-Time Decisions

Built-in /snake page: each move is a POST /v1/decisions call. State describes board; a choice question lists legal moves; model's letter probabilities render on candidate squares. No text generated. On M4 Pro, 10×10 board: three games recorded at ~150 ms/step (real-time, no speedup). Best game: 22 food, 208 steps, self-collision. Critical detail: representation matters. With sensor data (next cell content, distance to food, reachable empty cells — computed by game, model only chooses), it plays well. With raw ASCII board only, 4B model fails (0 score, death at 26 and 13 steps) — it cannot read grid spatially.

Quality Benchmarks (typed-decisions, RTX 5060 Ti, Q8_0)

Spark-X2.5-4B base: Accuracy 0.574, KL Div 2.899, Brier 0.480, ECE 0.349

Rizzo Flow 4B (fine-tuned): Accuracy 0.648, KL Div 0.452, Brier 0.205, ECE 0.112

TypeSafe Jev 1.13.0 (from dataset card): Accuracy 0.727, KL Div 1.442, Brier 0.148, ECE –

Fine-tuning adds 7.4 percentage points accuracy; probability shape improves more dramatically (KL 6× lower, ECE 0.349→0.112). Yet accuracy still trails Jev's 0.727. ECE (Expected Calibration Error) improvement is notable: base model gave 0.9999 confidence even when wrong; fine-tuned learns to spread probability when evidence is split. Most visible on agent<em>trace</em>observability workflow: accuracy 0.366→0.502. Originally almost every trace returned "severe, needs human review" at 0.9999.

Training data: 28,321 questions from three public datasets (tasksource/procedural-typed-decisions, ZefanCai/Open-Jev, Praveenrajus/jev-bench). Evaluation-set contamination removed; SemIf fixtures zero contamination. One epoch, AdamW, lr 5e-5, 1,770 steps; 4B on RTX PRO 6000 took 41 minutes. LoRA adapter and full logs on Hugging Face.

Hardware Support (Tested by Maintainers)

Windows/Linux + NVIDIA (cuda): Tested on Windows 10 + RTX 5060 Ti — all project numbers from here. Linux untested.

Windows/Linux + AMD/Intel (vulkan): Community reports cover AMD Radeon 780M and Intel Iris Xe; maintainers have not independently reproduced.

Mac Apple Silicon (metal): Community reports cover M3 Pro; maintainers have not independently reproduced.

No GPU (vulkan CPU fallback): Only reported on Intel laptop with iGPU; headless no-GPU host untested.

Project states: "Tested by us on one machine" — rare candor in open source.

Deployment

git clone https://github.com/Rizzo-AI-Academy/rizzo-flow && cd rizzo-flow uv sync --locked uv run rizzo download uv run rizzo serve

Open 127.0.0.1:8017/playground. Model 4.4 GB; llama.cpp runtime auto-selected per platform (CUDA ~570 MB, Vulkan 30 MB, Mac 11 MB).

When to Use / When to Avoid

Suitable: Embedding a "fuzzy if" in systems — ticket classification, content scoring, log severity judgment. Tasks where traditional code lacks precision but full LLM is too costly/slow. Rizzo Flow runs locally, outputs directly usable, probabilities calibratable on your data.

Unsuitable: Any need for text generation or world knowledge. VRAM not a blocker: 4B 8-bit uses 5.6 GB, fits consumer cards.

Before starting: Probabilities uncalibrated by default; status: ok ≠ correct answer. Calibrate with your labeled data via rizzo calibrate (temperature scaling). Calibration file binds to runtime, backend, and weight fingerprint — CUDA calibration unusable on Vulkan. Option limit: 26 per question (vs Jev's 255); exceed by splitting into two levels.

Known gaps: Underuses abstention — confident wrong answers when evidence missing (6/36 errors vs SemIf's 1/36). Residual position bias, no permutation debiasing. Single model instance, serializes concurrent requests. No hardening for public exposure.

Project Maturity & Engineering Habits

13 days old, 810 stars, 53 forks, 21 open issues. Clear positioning: Jev proved a "no chat, just decisions" model direction; Rizzo Flow shows the interface pattern works with open models on consumer hardware — albeit less accurate. Contributors page lists concrete contributions: first AMD GPU run, first Mac run, NaN input returning 500 vs 422, Jev consistency test revealing three discrepancies. Ends with "Want to be next?" listing needed test reports. The thoroughness of data, limitations, hardware coverage, and acknowledgments in a 13-day project reflects strong engineering discipline.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LoRAllama.cpplocal inferenceconsumer GPUJevprobability calibrationSystem One ModelRizzo FlowSpark-X2.5typed decisions
Geek Labs
Written by

Geek Labs

Daily shares of interesting GitHub open-source projects. AI tools, automation gems, technical tutorials, open-source inspiration.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.