Tagged articles

TileRT

4 articles · Page 1 of 1
Xiaomi Tech
Xiaomi Tech
Jun 9, 2026 · Artificial Intelligence

What 1000 tokens/s Really Means: Inside Xiaomi MiMo’s UltraSpeed Breakthrough

The article explains how Xiaomi’s MiMo‑V2.5‑Pro‑UltraSpeed mode achieves a record‑breaking 1000 tokens per second inference speed, why such ultra‑fast performance matters for real‑time AI applications, and the FP4 quantization, DFlash decoding and TileRT inference technologies that make it possible without sacrificing model quality.

DFlashFP4InferenceSpeed
0 likes · 10 min read
What 1000 tokens/s Really Means: Inside Xiaomi MiMo’s UltraSpeed Breakthrough
Xiaomi Tech
Xiaomi Tech
Jun 9, 2026 · Artificial Intelligence

How Xiaomi’s MiMo‑V2.5‑Pro UltraSpeed Achieves 1000 TPS on a 1‑Trillion‑Parameter Model

Xiaomi’s MiMo‑V2.5‑Pro UltraSpeed mode breaks the 1000 tokens‑per‑second barrier for a 1‑trillion‑parameter model by combining FP4 expert‑only quantization, DFlash block‑masked speculative decoding, and TileRT’s ultra‑low‑latency GPU system, and the API is now available through a limited‑time trial.

AI inferenceDFlashFP4 Quantization
0 likes · 13 min read
How Xiaomi’s MiMo‑V2.5‑Pro UltraSpeed Achieves 1000 TPS on a 1‑Trillion‑Parameter Model
SuanNi
SuanNi
May 22, 2026 · Artificial Intelligence

How GLM‑5.1‑highspeed Achieves 7× Faster Inference to Become the World’s Fastest Flagship Model

On May 22, Zhipu launched the GLM‑5.1‑highspeed API, delivering 400 tokens per second—about 7× faster than the original model and twice as fast as Gemini 3.5 Flash—through a three‑layer optimization that rewrites the MoE inference path, introduces dynamic scheduling, and leverages TileRT’s AOT engine to cut latency while preserving full flagship capabilities.

GLM-5.1Real-time AITileRT
0 likes · 10 min read
How GLM‑5.1‑highspeed Achieves 7× Faster Inference to Become the World’s Fastest Flagship Model