Tagged articles

DFlash

8 articles · Page 1 of 1
DataFunTalk
DataFunTalk
Jun 29, 2026 · Artificial Intelligence

DSpark Explained: 10 Key Concepts You Need to Know

The DSpark system from DeepSeek combines batch decoding, speculative decoding, draft‑model tricks, Eagle‑MTP, DFlash parallelism, variable‑length scheduling and online confidence calibration to deliver up to 85% speedup and four‑fold throughput gains while maintaining generation quality.

Batch DecodingDFlashDSpark
0 likes · 12 min read
DSpark Explained: 10 Key Concepts You Need to Know
Xiaomi Tech
Xiaomi Tech
Jun 9, 2026 · Artificial Intelligence

What 1000 tokens/s Really Means: Inside Xiaomi MiMo’s UltraSpeed Breakthrough

The article explains how Xiaomi’s MiMo‑V2.5‑Pro‑UltraSpeed mode achieves a record‑breaking 1000 tokens per second inference speed, why such ultra‑fast performance matters for real‑time AI applications, and the FP4 quantization, DFlash decoding and TileRT inference technologies that make it possible without sacrificing model quality.

DFlashFP4InferenceSpeed
0 likes · 10 min read
What 1000 tokens/s Really Means: Inside Xiaomi MiMo’s UltraSpeed Breakthrough
Xiaomi Tech
Xiaomi Tech
Jun 9, 2026 · Artificial Intelligence

How Xiaomi’s MiMo‑V2.5‑Pro UltraSpeed Achieves 1000 TPS on a 1‑Trillion‑Parameter Model

Xiaomi’s MiMo‑V2.5‑Pro UltraSpeed mode breaks the 1000 tokens‑per‑second barrier for a 1‑trillion‑parameter model by combining FP4 expert‑only quantization, DFlash block‑masked speculative decoding, and TileRT’s ultra‑low‑latency GPU system, and the API is now available through a limited‑time trial.

AI inferenceDFlashFP4 Quantization
0 likes · 13 min read
How Xiaomi’s MiMo‑V2.5‑Pro UltraSpeed Achieves 1000 TPS on a 1‑Trillion‑Parameter Model
Old Zhang's AI Learning
Old Zhang's AI Learning
Apr 19, 2026 · Artificial Intelligence

Qwen3.6-35B: 4‑bit Quantization, DFlash Speedup, Claude Opus Distillation

The article reviews three optimization paths for the Qwen3.6‑35B model—four‑bit AWQ quantization variants, the DFlash speculative decoding accelerator, and a Claude Opus‑based distillation—detailing their implementation steps, benchmark results, and guidance on selecting the best version for different hardware and performance needs.

AIDFlashDistillation
0 likes · 11 min read
Qwen3.6-35B: 4‑bit Quantization, DFlash Speedup, Claude Opus Distillation
Old Zhang's AI Learning
Old Zhang's AI Learning
Apr 17, 2026 · Artificial Intelligence

How DFlash Achieves 8× Lossless Acceleration for Large‑Model Inference (Qwen3.5‑27B Example)

The article explains how DFlash’s block‑diffusion draft model and KV Injection boost speculative decoding speed by 5‑8× without sacrificing output quality, and how DDTree further raises the gain to over 8×, backed by benchmark results and integration guides for major inference frameworks.

DDTreeDFlashSpeculative Decoding
0 likes · 7 min read
How DFlash Achieves 8× Lossless Acceleration for Large‑Model Inference (Qwen3.5‑27B Example)
Old Zhang's AI Learning
Old Zhang's AI Learning
Apr 14, 2026 · Artificial Intelligence

Qwen3.5-27B-DFlash Delivers Up to 5× Faster Inference Without Quality Loss

The DFlash approach replaces speculative decoding’s autoregressive drafter with a block diffusion model and injects target‑model hidden features into every KV‑cache layer, achieving up to 5× speed‑up for Qwen3.5‑27B on single‑GPU and 1.5–1.9× on high‑concurrency workloads while preserving output quality.

DFlashInference AccelerationQwen3.5
0 likes · 12 min read
Qwen3.5-27B-DFlash Delivers Up to 5× Faster Inference Without Quality Loss