Tagged articles

GR-Inference

1 articles · Page 1 of 1
DataFunTalk
DataFunTalk
Oct 2, 2026 · Artificial Intelligence

Xiaohongshu's GR-Inference: Custom Engine for 3.6x Faster Generative Retrieval

Xiaohongshu built a custom inference engine GR-Inference for generative search retrieval, addressing unique load characteristics — long context, short decode, large dynamic beam, and constrained generation — that break general frameworks, achieving 1.5–3.6x throughput over SGLang and improving recall and click-through rates.

GR-InferenceLLM servingSGLang
0 likes · 7 min read
Xiaohongshu's GR-Inference: Custom Engine for 3.6x Faster Generative Retrieval