360 Zhihui Cloud Developer
Sep 18, 2026 · Artificial Intelligence
Two-Level AI Inference Gateway: Prefix-Aware Routing & vLLM Router Fixes
This article details a two-level AI inference gateway architecture using Higress with a custom WASM plugin for cross-cluster instance routing and a modified vllm-router for intra-instance replica selection, featuring prefix-cache affinity, multi-signal scoring, RAII-guarded in-flight counting, and a soft_k_choice strategy to balance load and avoid thundering herd.
AI inference gatewayHigressRAII guard
0 likes · 29 min read
