CLM-8B: System One Decision Model 9x Faster Than Jev
CLM-8B is a contrastive language model that maps states and actions to a shared vector space for fast decision-making, matching Jev's zero-shot performance on benchmarks like WikiRacing and Super Mario while being 9x faster, and serves as a verifier setting new SOTA on DeepSWE and Terminal-Bench with 4-5x speedup using decoupled encoders and vector caching.
CLM (Contrastive Language Model) is a decision model trained with contrastive learning. It maps states and actions into the same vector space and scores actions by dot product. Given a current state and a set of candidate actions, CLM scores each action and picks the highest. This simple mechanism yields strong results.
The term "System One" comes from Kahneman's fast and slow thinking. Most agent decisions rely on step-by-step LLM reasoning — slow but reliable. CLM mimics the intuitive "glance and know" reaction, and the authors explicitly call it a System One model.
Zero-Shot Benchmark Results
Zero-shot tests cover computer use, games, and tool calling. On T-Rex, BFCL v4, WikiRacing, and Super Mario, CLM-8B matches Jev's performance while inference is much faster. The advantage grows with more candidate actions. In WikiRacing, where the action space is large, CLM-8B's latency is about one-ninth of Jev's.
Verifier Role
CLM also works as a verifier: sample several candidate answers, then use CLM to pick the best. On 38 held-out DeepSWE tasks, a lightly fine-tuned CLM reaches 81.6%; on 30 held-out Terminal-Bench 2.1 tasks, it reaches 87.6%, both new state-of-the-art. By contrast, Jev cannot even achieve pass@1 on these long-horizon tasks, making it ineffective as a verifier. CLM verification is 4.1–5.7x faster than Jev.
Core Design: Decoupled Encoders and Vector Caching
The key design is decoupling state and action encoders. Because they are independent, action embeddings can be cached and reused. In real scenarios the state changes constantly while candidate actions often stay fixed. The server maintains a vector cache similar to vLLM's KV cache reservation. On a cache hit, the encoder forward pass is skipped and only a dot product is computed. In a test with 20 repeatedly visited rooms and 50 candidate actions, server p50 latency dropped from 2.0 ms to 0.7 ms.
Three-Stage Training
Training proceeds in three stages:
Pre-training on 60 million Nemotron DQA question-answer pairs.
Intermediate training on 30 million synthetic hard negatives.
Post-training on 1 million agent trajectories.
The intermediate stage's value is clear: adding hard negatives from the start yields only 62.4% top-1 accuracy; pre-training first then refining reaches 69.2%. Hard negatives are a refinement tool, not a substitute for pre-training. The post-training stage mixes in 40% pre-training data to prevent performance regression.
Playground and API
Code and models are open-sourced on GitHub and Hugging Face. The repository includes a playground where you can write a state, add a question, and see the answer distribution. Deployment is straightforward: a Qwen3-8B encoder service plus a clm-serve process. The API supports three task types: Noul (binary probability), Choice (classification), Score (scoring), plus a pure ranking interface for candidate ranking or best-of-N selection.
Scaling Laws
CLM's test loss follows a power-law decay with training compute, model size, and data volume. The optimal head size and training token count show an almost linear relationship: roughly 310 tokens per parameter. This provides a predictive basis for future scaling.
Conclusion
CLM does not replace LLMs; it is a lightweight decision module suited for scenarios requiring fast response. While the field piles on inference compute, contrastive learning has produced an "intuition" model with competitive results. The project is fully open-sourced on GitHub (https://github.com/Contrastive-LM/CLM) and Hugging Face (https://huggingface.co/Contrastive-LM) with documentation, weights, and data available.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
AI Engineering
Focused on cutting‑edge product and technology information and practical experience sharing in the AI field (large models, MLOps/LLMOps, AI application development, AI infrastructure).
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
