Tagged articles

FP4 quantization

9 articles · Page 1 of 1
Architect
Architect
Sep 14, 2026 · Artificial Intelligence

DeepSeek V4.1 Flash Architecture: The State Lifecycle Behind 1M Context

This article dissects DeepSeek V4.1 Flash's architecture for 1M-token context, explaining how CED reduces prefill compute, CSA2 shares global KV across layers, hierarchical sparse indexing narrows search, FP4 quantization shrinks storage, SWA uses bounded replay for recovery, and Engram adds a conditional memory path—revealing five cost dimensions of long-context deployment.

Bounded ReplayCEDCSA2
0 likes · 21 min read
DeepSeek V4.1 Flash Architecture: The State Lifecycle Behind 1M Context
Xiaomi Tech
Xiaomi Tech
Jun 9, 2026 · Artificial Intelligence

How Xiaomi’s MiMo‑V2.5‑Pro UltraSpeed Achieves 1000 TPS on a 1‑Trillion‑Parameter Model

Xiaomi’s MiMo‑V2.5‑Pro UltraSpeed mode breaks the 1000 tokens‑per‑second barrier for a 1‑trillion‑parameter model by combining FP4 expert‑only quantization, DFlash block‑masked speculative decoding, and TileRT’s ultra‑low‑latency GPU system, and the API is now available through a limited‑time trial.

AI inferenceDFlashFP4 quantization
0 likes · 13 min read
How Xiaomi’s MiMo‑V2.5‑Pro UltraSpeed Achieves 1000 TPS on a 1‑Trillion‑Parameter Model
Architect's Guide
Architect's Guide
May 29, 2026 · Artificial Intelligence

What Makes DeepSeek V4 Different? A Deep Technical Dive into Its Innovations

DeepSeek V4 introduces a suite of architectural breakthroughs—including mixed‑expert MoE, manifold‑constrained hyper‑connections, CSA/HCA hybrid attention, and FP4 quantization—that slash inference cost by up to tenfold while delivering million‑token context, competitive benchmarks, dual model variants, and a disruptive pricing strategy.

AI model benchmarkAgentic AIDeepSeek-V4
0 likes · 41 min read
What Makes DeepSeek V4 Different? A Deep Technical Dive into Its Innovations
Architect's Must-Have
Architect's Must-Have
Apr 28, 2026 · Artificial Intelligence

Why DeepSeek V4 Stands Apart: A Deep Dive into Its Architecture and Performance

DeepSeek V4 introduces a suite of architectural innovations—including mixed attention, manifold‑constrained hyper‑connections, the Muon optimizer, and FP4‑aware quantization—that together slash million‑token inference cost to a tenth of its predecessor while delivering benchmark results that rival top‑tier closed‑source models.

DeepSeek-V4FP4 quantizationMixture of Experts
0 likes · 44 min read
Why DeepSeek V4 Stands Apart: A Deep Dive into Its Architecture and Performance
DeepHub IMBA
DeepHub IMBA
Apr 27, 2026 · Artificial Intelligence

DeepSeek‑V4 Deep Dive: Engineering Million‑Token Context Efficiency

The article provides a thorough technical analysis of DeepSeek‑V4, detailing how mixed sparse attention (CSA + HCA), manifold‑constrained hyper‑connections, the Muon optimizer, FP4 quantization, and a suite of infrastructure tricks enable stable training and inference with up to one‑million token contexts while achieving state‑of‑the‑art benchmark results.

CSADeepSeek-V4FP4 quantization
0 likes · 22 min read
DeepSeek‑V4 Deep Dive: Engineering Million‑Token Context Efficiency
TechVision Expert Circle
TechVision Expert Circle
Apr 25, 2026 · Artificial Intelligence

DeepSeek V4 Arrives After 15 Months: Hybrid Attention Beats Expectations

DeepSeek's V4 preview, released 15 months after V3, introduces million‑token context, a hybrid attention architecture, and two model variants—Pro and Flash—offering significant compute savings, new training techniques, and performance gains that challenge top closed‑source LLMs.

DeepSeek-V4FP4 quantizationMillion-token Context
0 likes · 12 min read
DeepSeek V4 Arrives After 15 Months: Hybrid Attention Beats Expectations