Tencent Backend Veteran's Reset: Mastering LLM Inference Engineering from Zero

A Tencent senior backend engineer with 15 years experience shares his transition to LLM inference engineering, detailing learning paths for model architecture, inference optimization, and self-assessment exercises, while arguing that traditional infrastructure skills remain crucial for AI Infra and Agent Infra development.

Tencent Technical Engineering
Tencent Technical Engineering
Tencent Technical Engineering
Tencent Backend Veteran's Reset: Mastering LLM Inference Engineering from Zero

1. Preface

The author worked 15 years, joined Tencent in 2016, promoted to T12 in 2022 via search middleware work, as a senior search backend engineer. In early 2025, he switched to a large model inference engineering team, setting aside past seniority to start from zero. This article serves as self-reflection and encouragement for peers hesitating in the AI era.

2. Becoming a Zero-Level Novice in LLM Inference Engineering

From 2024 to 2025, the author shifted from senior search backend engineer to LLM inference engineering novice — literally a beginner with no senior title, no task assignment. First three months focused on peripheral tasks: performance evaluation, service deployment. Meanwhile, he deeply read SGLang source code.

2.1 Choosing the AI Era

Thirty years ago PC internet arrived when he was a child; fifteen years ago mobile internet rose when he just graduated chasing high-pay offers. Today AI surges while he approaches 40, at peak mental capacity. Staying in comfort zone would leave him unable to answer future generations: "When the AI revolution came, what were you doing?"

ChatGPT's late 2022 debut made Transformer-based LLMs sweep the world. The scaling law — larger parameters yield higher intelligence baseline — still holds. Best way to embrace this wave: choose the direction closest to LLMs, even enter core LLM domains.

How to enter? Two years ago search products were close to LLM applications in form, but internally still used traditional tech for recall/ranking; multi-turn dialogue, long context, Agent Runtime lacked clear product shapes. He chose LLM inference engineering because regardless of upper-layer product evolution, compute and inference acceleration infrastructure remain indispensable foundations.

2.2 Technical Dilemma of Managers

Managers shouldn't be mere information conduits or survive on information asymmetry. Management is a rigorous discipline requiring continuous learning. Yet facing a brand-new tech wave, managers struggle to focus 100% on learning, leading to dilemma: either technical learning stays superficial or management quality slips.

Traditional path often suggests abandoning technology: many technical managers become pure "career managers". Their core duty should be building pragmatic team culture, inspiring engineers to create more value. Reality shows some managers want management perks while keeping technical titles, causing anomalies: academics with many papers but admitting "details done by students" at Q&A industry leads named on big projects but unable to deeply participate, using them for promotion. Fortunately, evaluation is returning to essence — technology and rank must serve project contribution.

As a programmer with technical faith, he didn't want to prematurely become a career manager. But truly diving into LLM tech demands massive time and extreme focus. Organization won't allow a high-level manager to be a "beginner" in LLM direction. So he chose to shed management title and baggage, restart with pure mindset.

Looking back, for managers or senior engineers wanting to enter new fields, besides his zero-reset approach, there's a second path: AI Infra and general Infra are very close. Especially with Agent product roadmap clarifying, evolving from general Infra to AI Infra is a smooth natural progression. In an AI product team, AI Infra isn't just GPU compute acceleration; it still needs solid platform engineering, R&D efficiency, stability assurance — traditional infrastructure support. Starting from familiar entry points, thinking how to "AI-ify" general Infra, is equally viable.

2.3 Organizational Change as Catalyst

In 2024, his team merged into a new department. For months the group lacked mainline tasks; he led them to "find work", handling edge miscellany, making stability harder. Then org planned to merge team into new department's product mainline, requiring him to drop management title. He accepted — aligning with department's "task-first" philosophy and his belief in "engineer culture": What determines a person's value is never their title, but what they are doing.

He chose to enter from familiar technical direction, familiarized with product mainline R&D system and business form, and through a three-month independent code refactor, mapped the core pipeline end-to-end.

Internal struggles existed: moving from Dept A to B meant completely different R&D systems, identity switch, role gap. He considered finding another management slot, even interviewed. But each decision he reminded himself: In the AI era, you should stay in the team closest to AI. At that time, search seemed closest to AI — Transformer born at search giant, RAG natural extension of traditional search, most dialogue needs essentially information retrieval/processing. He valued long-term technical direction over management title.

By early 2025, after deep understanding of new department's infrastructure, he saw AI Infra team hiring info, had deep talks with multiple leads, and joined that team for LLM inference engineering.

3. LLM Inference Engineering Entry Guide

LLM inference engineering involves model structure, GPU hardware characteristics, distributed systems, scheduling — a complex systems engineering. Not merely DL model deployment; core contradiction is trade-off among compute, VRAM, communication: Prefill phase compute-intensive, Decode phase memory-bandwidth-intensive; KV Cache, Continuous Batch, Prefix Cache, TP are essentially resource swaps under hardware constraints to achieve higher throughput and lower latency at lowest cost.

Knowledge coverage broad; suggest top-down approach: first build systematic cognition, then break down individually.

3.1 Build Systematic Cognition of Model Structure

Mainstream LLMs all use Transformer architecture. First step: systematic cognition of model structure.

Familiarize Transformer model structure, know core components: Embedding, Attention, FFN, Normalize, Residual, LM Head; able to draw model structure diagram.

Understand DL tech system basics and key tech: perceptron, multi-layer neural networks, gradient descent, backpropagation, hyperparameters, CNN, RNN, Seq2Seq, Attention structure, etc.

Know key figures, companies, investments in DL history.

Can use AI to learn but usually not systematic enough; recommend books or systematic overview articles. Recommended books:

DL and NLP basics: Deep Learning from Scratch: Theory and Implementation with Python , Advanced Deep Learning: Natural Language Processing

DL human stories: The Deep Learning Revolution

Transformer structure is simplest among above three; recommend using AI or searching Zhihu articles (despite splash ads, Zhihu remains highest density of LLM knowledge/discussion in China).

Author previously summarized NLP evolution to Transformer: https://zhuanlan.zhihu.com/p/2027080361256005976 (From Word Vectors to LLMs: Brief Notes on NLP Tech Evolution).

3.2 Build Systematic Cognition of LLM Inference Engineering

When analyzing technical problems, first clarify two questions: what problem to solve, what evaluation system. Grasping core objectives and evaluation system gives essential judgment on each component: how it supports overall goal, how to quantify its pros/cons.

LLM inference engineering top-level goals align with traditional software engineering: functionality, performance, cost, stability. Evaluation system includes core metrics: TTFT, TPOT, QPS, TPS, Concurrency, and cost. Around these metrics, every system-level optimization can be abstracted as resource swap among compute, VRAM, communication. Thus three-layer model: top business goals, middle quantitative evaluation standards, bottom concrete implementation means.

On engineering system and org division, implementation means further split into three sub-directions:

Business Layer : interfaces inference system with upper business. Includes inference service SDK, inference platform, service deployment & monitoring, business-customized request traffic scheduling.

Scheduling Layer : runtime carrier of model inference, directly determines core performance; most inference acceleration optimizations land here. Typical tech: distributed parallelism, Prefix Cache, Overlap Scheduler, Continuous Batch, Chunked Prefill, Speculative Decoding, etc.

Operator Layer : implements high-performance basic compute units, focuses single-point compute optimization, e.g., QKV compute, Normalize, Linear compute.

Layered architecture's core value: reduce complexity of large-scale systems and large R&D teams, let each direction team focus on own responsibility boundaries. For smaller business scale, lower system complexity, such fine-grained split unnecessary; engineers can work across multiple layers.

Can also use AI to learn basic concepts. For experienced developers, forming systematic cognition is relatively simple; hard part is diving into details, e.g., how Prefix Cache, Overlap Scheduler are implemented. Suggest learning via source code; nano-vLLM, mini-sglang are good entry codebases. Author previously summarized readings: https://zhuanlan.zhihu.com/p/1989806890381746916 (Learning Key LLM Inference Functions via nano-vLLM) https://zhuanlan.zhihu.com/p/2009377131159897395 (Learning Key LLM Inference Functions via mini-sglang)

3.3 Understand Mainstream Model Structure Evolution Trends

In LLM era, traditional "algorithm" and "engineering" role boundaries blur. For inference engineering, only knowing system architecture is far from enough; deep understanding of model structure is also required — because new-gen model structure evolution is essentially algorithm innovation deeply combined with engineering optimization. From initial Full Attention + Dense FFN to today's sparse attention, linear attention, hybrid architecture, MoE, and Pre-Norm, RoPE, MLA, cross-layer KV Cache reuse, etc., reviewing this evolution reveals: models become smarter while evolving toward drastically lower compute complexity, better long-context adaptation, achieving higher throughput at lower inference cost. DeepSeek is typical representative: attention layer's MLA, HCA, CSA designs, and DSpark speculative decoding are cases of algorithm-engineering fusion.

Further thought reveals current LLM structure evolution is essentially an "experimental science" . Many innovations are conceivable; when innovation ideas become cheap, expensive part is verifying whether experiments can be done better and faster. At this point infra is obviously critical; understanding model structure, clarifying motivations behind evolution, helps clarify trends, enabling more complete, forward-looking infra.

Recommended reading:

https://magazine.sebastianraschka.com/p/the-big-llm-architecture-comparison

(The Big LLM Architecture Comparison) https://zhuanlan.zhihu.com/p/2035408576978629074 (From 305 GB to 7.4 GB: Panorama of LLM KVCache Architecture Evolution)

3.4 Deep Dive into One Optimization Direction

LLM inference engineering knowledge system extremely vast. After building macro systematic cognition, suggest combining concrete business pain points or personal interest, pick a subdivided topic for analysis learning. Currently high ROI and industry-focused sub-directions include:

Prefill/Decode (PD) Separation : decoupling deployment of Prefill and Decode phases significantly improves whole-machine throughput and resource utilization, especially critical for teams with limited compute budget, scarce GPU resources. Though open source has good implementations, adapting to optimal performance for own business scenarios remains challenging.

Distributed KV Cache : for business scenarios with multi-turn interaction, long context, can greatly improve Prefix Cache hit rate, significantly reduce redundant compute cost.

Speculative Decoding : exemplar of algorithm innovation deeply combined with engineering optimization (e.g., DSpark architecture); cleverly utilizes Decode phase's unsaturated compute to greatly boost system throughput.

Beyond these major directions, many interesting independent "small tech points" worth deep diving. E.g., author previously systematically analyzed Overlap Scheduler: https://zhuanlan.zhihu.com/p/1993845553814069701 (LLM Inference Acceleration: Deep Analysis of Overlap Scheduling and Performance Trade-off Art). When digging down bottom execution chain, often discover hidden reefs overlooked by others: under specific constraints, Overlap Scheduler may actually break Continuous Batch's batching efficiency. Such micro-mechanism research can be verified via demo code, yielding more solid learning accumulation.

4. LLM Inference Engineering Entry Self-Assessment Exercises

In today's widespread AI usage, if learning stays at "passive input" lacking "active output", easy to fall into false mastery of "understand when reading, clueless when doing". To help verify real understanding of inference underlying mechanisms, author designed three exercises based on past year's practical project experience.

4.1 Shared Prefix Attention Optimization Within Same Batch

Prefix sharing often triggers cross-request/cross-batch Prefix Cache. But if within a single batch multiple requests share common prefix — e.g., Request1 = [a, b, c, 1, 2, 3], Request2 = [a, b, c, 4, 5, 6] — how to avoid repeated Prefill compute and KV Cache duplicate writes for prefix [a, b, c]?

Such needs common in personalized recommendation. E.g., document scoring tasks input form user_info + doc_N: prefix user_info represents user attributes (identical within same batch), suffixes are different docs to evaluate.

Abstracted, essence is: attention compute optimization for shared prefixes within same batch. To fully solve and implement, need to integrate multi-dimensional bottom-layer knowledge:

Deep understanding of Causal Mask construction;

Code-level proficiency in KV Cache allocation and usage;

Understanding FlashAttention / FlashInfer mainstream Attention Kernel interface design;

Know GPU Thread Block parallel characteristics.

Can first try to conceive solutions independently, then read author's systematic analysis: https://zhuanlan.zhihu.com/p/2072091176719669236 (LLM Inference Acceleration: Comprehensive Analysis of Shared Prefix Optimization Technology). In actual business landing, same-batch shared prefix often just a single-point problem, still needs combining with traditional backend tech; those peripheral integration points (e.g., personalized model structure, business system embedding) often more difficult.

4.2 Beam Search Engineering Implementation

Beam Search tech hot in search/ad/recommendation last 1-2 years, closely related to Kuaishou's OneRec paper ( https://arxiv.org/html/2506.13695v3, OneRec Technical Report), many companies already landed it in internal search/recommendation projects.

Beam Search principle straightforward: each step after backbone and LM Head execution, instead of keeping only single best token, keep multiple candidate tokens as multiple sub-paths to continue derivation; each step selecting next token, merge all sub-paths' candidates, sort by cumulative logprob of entire sub-path, keep top Beam Width sub-paths, iterate until generation ends, then pick best sub-path.

Key technical points involved:

KV Cache reuse. Multiple sub-paths share common prefix, should share prefix, but they are multiple sub-requests expanded from single request; need to understand inference engine's KV Cache implementation for efficient reuse;

High-efficiency pruning. When selecting next batch of legal tokens, need to merge all paths' candidates for unified compute and screening. When Beam Width large, e.g., BW=1024, need to pick best 1024 from 1024 * Vocabulary Size candidates, compute huge;

Constrained decoding. Beam Search applied to recommendation often restricts generation to legal DocIDs, so each step needs vocabulary constraint; how to achieve highest performance;

Integration with inference framework. Relative to Greedy decoding, Beam Search is niche; how to avoid niche feature becoming burden on mainstream feature iteration.

Can try implementing on mini-sglang or similar small engines. Author implemented high-performance Beam Search on SGLang last year, achieved industry SOTA performance at that time (now stronger custom engines exist), wrote summary: https://zhuanlan.zhihu.com/p/2031882274229179444 (Beam Search in LLM Inference Engine: Engineering Challenges, Mainstream Implementations, and SGLang Deep Optimization). SGLang official developers also refactored/merged based on his branch, finally merged into SGLang v0.5.19.

4.3 Reading LLM-Direction Technical Articles

LLM-direction technical articles abundant: papers, paper interpretations, model structure analyses. Entry milestone is ability to comprehend these articles. Reading HCA/CSA, hybrid architecture, DSpark etc. technical papers/articles, can understand basic principles.

5. AI Infra: Natural Expansion of Traditional Infra Map

After mastering conventional LLM inference acceleration methods, switching to product perspective reveals LLM inference engineering not a brand-new field born from zero, but natural expansion of traditional Infra map. Two share origin: most AI Infra skill points built on general Infra foundation. Two also diverge: AI Infra, Agent Infra differ significantly from traditional backend systems in resource model and service paradigm. For an Infra engineer, learning AI Infra isn't starting from scratch, but natural extension of knowledge boundaries — traditional backend development's accumulated system capabilities remain highly valuable in AI era.

5.1 Traditional Infra Knowledge System Still Foundation of LLM Inference Services

Today's external LLM inference API services still built on traditional Infra capabilities: microservices, distributed systems, RPC, storage, observability, R&D efficiency systems all indispensable.

A high-performance, powerful LLM inference interface is core component of model product. But if team developers only master vLLM/SGLang-type inference engines, system stays at demo stage. When user scale grows, it inevitably hits classic large-scale distributed system problems, highly dependent on traditional Infra capabilities:

Traffic estimation, capacity planning, data center & cluster resource planning;

Risk isolation, stability assurance while maintaining iteration efficiency, building CI/CD system;

Service release, container scheduling, full-chain observability system construction;

Microservice vs monolith trade-offs, stateful vs stateless service boundary division.

5.2 AI Infra: Resource Constraints Differ, But System Thinking Largely Inherited from Traditional Infra

LLM inference engineering, though emerging, birthed Continuous Batch, Overlap Scheduler etc. proprietary skills, but its underlying engineering thinking largely traces back to traditional systems engineering. Take PagedAttention: essentially moves OS virtual memory page table idea to VRAM management — logically contiguous KV Cache cut into fixed-size page blocks, mapped via Block Table to physically discrete VRAM pages. Similar thought migrations abound: TP, CP model distributed parallel schemes, core goal consistent with traditional distributed systems: split compute, share load; CUDA Stream multi-stream mechanism, conceptually mirrors CPU multi-threading; distributed KV Cache sharding, lifecycle management schemes can heavily borrow mature experience from distributed storage systems. In short, GPU-related engineering not a completely disjoint knowledge system. Senior backend engineers with years of CPU-side depth can fully transfer accumulated systems thinking to GPU inference engineering.

5.3 Traditional Backend Development Value Reassessment & Agent Infra New Battlefield

5.3.1 Traditional Engineering Experience More Important in AI Era

Current AI Infra hiring prefers AI Native talent with training/inference framework, kernel operator development experience, but author believes traditional backend engineers' transition value underestimated. AI specialists may onboard faster, but traditional engineers with solid system design foundation, aided by AI tools, can catch up quickly and leverage experience advantages.

AI popularization replaces basic coding repetitive work, making code review, architecture design core matters. Traditional engineers' long-accumulated code taste, architecture thinking, complex system control ability become more critical in AI era, effectively avoiding AI-generated code's architectural chaos, redundancy, hidden risks.

Thus, in current AI Infra talent scarcity, letting passionate backend engineers expand knowledge boundaries into AI Infra is good choice for team management and talent cultivation.

5.3.2 Layered Technical Direction Demand: Fewer at Bottom, More at Top

AI Infra industry has clear layered characteristic: lower layer stronger commonality, less demand; upper layer greater business differentiation, more manpower demand, overall "bottom standardized, top differentiated" pattern. For LLM inference engineering, each layer demand characteristics:

Operator Layer : model structure, hardware ecology will converge; new operator, new structure adaptation similar to Linux kernel iteration — core value extremely high, industry-wide reusable, but no need for large manpower to repeatedly develop; only few top teams needed to continuously iterate adapting new models, new hardware; overall manpower demand minimal.

Inference Engine Scheduling Layer : general scheduling capabilities already maturely covered by open source ecosystem; vast majority of businesses need no deep customization, only few business-bound scenarios have slight customization needs; overall manpower demand limited.

Business Layer : custom adaptation for own business scenarios, pipeline optimization, stability polishing — differentiated work where manpower demand is larger.

Business layer's differentiated work is precisely where traditional backend engineers can transfer system design foundation and leverage experience advantages.

5.3.3 Agent Infra: Traditional Backend Developers' New Battlefield

As AI bottom foundation, operators, inference engines gradually standardize and open-source, bottom differentiation competition space continuously narrows, industry innovation center fully shifts upward, Agent Infra becomes traditional backend developers' most suitable new battlefield.

Agent is junction of general intelligence and vertical business; different scenarios' Agent usage forms, runtime logic, engineering demands differ vastly: e.g., research scenario, OpenAI can schedule 10k+ Agents to solve Navier-Stokes equations, while ordinary scenarios only need single Agent for lightweight spreadsheet organization, but massive personal scenarios form larger scale. Highly differentiated scenarios give Agent products and Agent Infra huge engineering landing and optimization space.

From service paradigm view, Agent era completely reconstructs traditional backend service model. Traditional API service is "convenience store model": short requests, stateless, ultra-low latency, call-and-go, immediate resource release. Agent service is "hotel model": core characteristics — multi-turn interaction, long duration, stateful, long lifecycle; needs memory storage, tool calls, external resource interaction to complete tasks; session persists, state changes frequently.

Comparison of traditional API service (convenience store) vs Agent service (hotel) paradigms
Comparison of traditional API service (convenience store) vs Agent service (hotel) paradigms

This paradigm shift brings entirely new engineering challenges: long-session scheduling, state persistence, long/short-term memory management, tool sandbox isolation, Agent observability & effectiveness evaluation — all belong to core domain of traditional distributed systems engineering, where traditional backend experience can shine.

6. Conclusion

This article's knowledge, experience, industry analysis come from personal experience and observation, representing personal views only. As a 15-year "internet veteran", in this AI wave, there's both excitement for new era transformation and shock at domestic internet industry's age discrimination; time slips away in hesitation — not choosing is the worst choice. Though he feels he could work another 10-20 years until brain stops, era rolls forward; past experience may flourish or be washed away by history's torrent. Each tech revolution's reset merely puts you back at starting line 15 years ago — believe in yourself, keep running.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Backend EngineeringSpeculative DecodingLLM InferenceBeam SearchCareer TransitionAI InfrastructureKV CachePagedAttention
Tencent Technical Engineering
Written by

Tencent Technical Engineering

Official account of Tencent Technology. A platform for publishing and analyzing Tencent's technological innovations and cutting-edge developments.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.