Tagged articles

LLM serving

10 articles · Page 1 of 1
DataFunTalk
DataFunTalk
Oct 2, 2026 · Artificial Intelligence

Xiaohongshu's GR-Inference: Custom Engine for 3.6x Faster Generative Retrieval

Xiaohongshu built a custom inference engine GR-Inference for generative search retrieval, addressing unique load characteristics — long context, short decode, large dynamic beam, and constrained generation — that break general frameworks, achieving 1.5–3.6x throughput over SGLang and improving recall and click-through rates.

Beam SearchGR-InferenceInference Optimization
0 likes · 7 min read
Xiaohongshu's GR-Inference: Custom Engine for 3.6x Faster Generative Retrieval
Cambridge Mofang Notes
Cambridge Mofang Notes
Sep 2, 2026 · Artificial Intelligence

Inference Frameworks vs Platforms: How LLMs Actually Run on Your Hardware

This article distinguishes between inference frameworks (llama.cpp, vLLM, SGLang) that execute model computations and inference platforms (Ollama, LM Studio, Xinference) that manage deployment, explaining their roles, interactions, and how to choose tools for local or server-side LLM inference.

AI inferenceLLM servingLM Studio
0 likes · 16 min read
Inference Frameworks vs Platforms: How LLMs Actually Run on Your Hardware
JD Retail Technology
JD Retail Technology
Aug 21, 2026 · Artificial Intelligence

Janus: Dual‑Timescale Scheduling for Production‑Scale Multi‑LLM Serving

Janus, a Service‑Engine co‑design system built on Oxygen xLLM, uses dual‑timescale scheduling, performance‑oracle‑driven placement, and a three‑state model lifecycle to handle bursty traffic, power‑law application hotness, and heterogeneous resource demands, achieving 0.97–1.0 SLO rates while cutting device usage by 27%.

AI InfrastructureLLM servingSOSP 2026
0 likes · 19 min read
Janus: Dual‑Timescale Scheduling for Production‑Scale Multi‑LLM Serving
AI Engineering
AI Engineering
Jul 4, 2026 · Backend Development

How SGLang Encoded Engineering Experience into Agents and Achieved Up to 2.75× Kernel Speedups

The SGLang team turned their benchmarking, profiling, CUDA kernel tuning, and production‑issue triage know‑how into reusable agent skills, merging three KDA‑Pilot PRs that delivered up to 2.75× kernel acceleration, a 71.4% throughput boost for Qwen3‑Next and a TTFT reduction from 456 ms to 168 ms, while outlining a repeatable workflow and practical rules for large‑scale performance engineering.

Agent AutomationCUDA optimizationLLM serving
0 likes · 16 min read
How SGLang Encoded Engineering Experience into Agents and Achieved Up to 2.75× Kernel Speedups
Data Party THU
Data Party THU
Nov 2, 2025 · Operations

How to Maximize vLLM Throughput: Batch Size, Quantization, and Monitoring Tips

This guide explains how to unleash vLLM’s full potential by optimizing batch size, leveraging 4‑bit quantization, tuning concurrency parameters, planning capacity with token‑per‑second metrics, and implementing robust monitoring to balance latency, cost, and scalability in production deployments.

LLM servingbatchingcapacity planning
0 likes · 10 min read
How to Maximize vLLM Throughput: Batch Size, Quantization, and Monitoring Tips
NewBeeNLP
NewBeeNLP
Jan 14, 2025 · R&D Management

How to Kickstart Your CS Research Journey and Find LLM Serving Ideas

The author shares a candid half‑year reflection on entering computer‑science research, outlining practical steps for discovering research ideas, navigating papers, focusing on LLM serving systems, and emphasizing collaboration to help newcomers succeed in academia.

LLM servingSystem Designacademic journey
0 likes · 9 min read
How to Kickstart Your CS Research Journey and Find LLM Serving Ideas
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Sep 17, 2024 · Artificial Intelligence

Boosting LLM Inference: How NanoFlow Doubles Throughput

The article introduces NanoFlow, a novel service framework that leverages intra‑device parallelism, operation‑based pipelining, and async scheduling to significantly improve large language model serving throughput, achieving up to 1.91× higher performance while integrating with Alibaba Cloud PAI.

Alibaba Cloud PAIGPU SchedulingLLM serving
0 likes · 7 min read
Boosting LLM Inference: How NanoFlow Doubles Throughput