Xiaohongshu's GR-Inference: Custom Engine for 3.6x Faster Generative Retrieval
Xiaohongshu built a custom inference engine GR-Inference for generative search retrieval, addressing unique load characteristics — long context, short decode, large dynamic beam, and constrained generation — that break general frameworks, achieving 1.5–3.6x throughput over SGLang and improving recall and click-through rates.
Load Profile: Why General Frameworks Fail
Xiaohongshu's search generative retrieval workload differs fundamentally from typical LLM serving: each request carries a long context of 1k–5k tokens but decodes only 3–5 steps, while maintaining a beam that dynamically expands from 20 to 300 to 900, and every generated path must satisfy a legal Semantic ID (SID) constraint. Traditional ANN dual-tower retrieval encodes query and document separately; Xiaohongshu instead uses offline hierarchical residual clustering to produce multi-level SIDs, then online an LLM generates SIDs level by level conditioned on query and context, mapping back to candidate notes. The inference goal is not a single answer but many valid candidates within a strict latency budget.
The combination of long context, short decode, large beam, and constrained generation invalidates three default assumptions of general inference frameworks:
State management: Managing each beam as an independent sequence duplicates the long context, wasting memory linearly with beam width.
Compute organization: Expanding decode attention per beam prevents reuse of the shared context, causing massive redundant reads.
Execution path: Relying on external loops to assemble dynamic beam and item-constrained generation pushes scheduling, synchronization, and maintenance costs onto engineers.
The key question is not whether a general framework can run, but whether it can run efficiently within the target SLA and resource budget.
GR-Inference: Request-Centric Design
Xiaohongshu collaborated with NVIDIA to build GR-Inference, a specialized inference engine centered on the Request-Centric principle: the request — not a single beam — is the basic unit of state ownership. The request state is split into three cooperating objects:
ContextKV: Request-level shared long context, stored once, read-only during decode.
BeamKV: Per-beam key-value caches that grow as beam expands from 20 to 900.
BeamPath: Indices tracking each beam's generation history.
This separation keeps ContextKV constant while BeamKV and BeamPath absorb the growth.
Sparse-Dense Cooperative Decode Attention
The decode attention kernel is split into three segments:
K1 (dense): Processes the long ContextKV using Tensor Cores.
K2 (sparse): Handles the irregular gather from BeamKV.
K3 (merge): Merges the two results via an online log-sum-exp reduction without materializing the full attention matrix.
GPU-Resident SID Trie for Constrained Generation
The SID trie is stored in CSR format and pinned in GPU memory. A fused kernel performs prefix traversal, legal-token gather, and local top-k selection, shrinking the candidate set from the full vocabulary to only the legal branches at the current trie node.
Benchmarks and Results
Under a fixed configuration — Qwen3-0.6B, BF16, prefill 100–1000 tokens, decode 3 steps, end-to-end latency ≤100 ms, fixed beam 900 — GR-Inference was compared against three SGLang beam-search implementations:
Standard beam search: GR-Inference achieves 1.5–2.53× higher throughput.
Item-constrained generation: GR-Inference achieves 3.24–3.63× higher throughput.
After deploying the new autoregressive SID retrieval channel in Xiaohongshu search:
Offline Recall@1000 improved by +5.7% .
Click-through rate increased by +0.03 percentage points .
Effective click-through rate increased by +0.2% .
Caveats and Boundaries
The author explicitly notes that the reported speedups hold only for the stated benchmark configuration and do not directly extrapolate to all context lengths or deployment setups. Moreover, the business metric gains are not solely attributable to the inference optimization; the engine provides serving efficiency and launch support, while the retrieval channel itself delivers the business effect.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunTalk
Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
