Multi-GPU LLM Inference: Step-by-Step Deployment Guide for Memory-Constrained Environments
This comprehensive tutorial covers multi-GPU inference for large language models, detailing VRAM estimation formulas, parallel strategies (tensor, pipeline, data), GPU interconnects (NVLink, PCIe), framework comparisons (vLLM, DeepSpeed, Transformers), quantization techniques (GPTQ, AWQ, bitsandbytes), performance benchmarking, and production deployment with Docker and Kubernetes.
Problem Background
Large model parameter counts (GPT-3 175B, LLaMA 70B, Qwen 72B) exceed single GPU VRAM (24GB/48GB/80GB), making multi-GPU inference necessary. Target roles: algorithm engineers, ops engineers, MLOps engineers, R&D engineers.
Multi-GPU Inference Characteristics
Model parameters distributed across GPUs
Cross-GPU communication required
VRAM and compute scale linearly
Inference latency may increase
Configuration complexity rises
Applicable Scenarios
Single GPU VRAM insufficient for full model
Need higher inference throughput
Need lower per-request latency
Multi-user concurrent requests
Batch data processing
Non-Applicable Scenarios
Model fits single GPU (prefer single GPU)
Ultra-low latency requirements (multi-GPU adds communication overhead)
Low inter-GPU bandwidth (e.g., PCIe 3.0 x8)
Core Knowledge: VRAM Estimation
Inference VRAM components:
Model parameters : parameter count × precision bytes
Activations : forward pass intermediate results
KV Cache : autoregressive generation cache
Gradients/optimizer states : training only, not inference
Formula: Total VRAM ≈ Model Params VRAM + KV Cache VRAM + Activations VRAM + Framework Overhead
Model Parameter VRAM
Model Params VRAM = Parameter Count × Precision BytesPrecision bytes: FP32=4, FP16/BF16=2, INT8=1, INT4=0.5
Examples: LLaMA 70B FP16 = 70B × 2 = 140GB; INT8 = 70GB; INT4 = 35GB
KV Cache VRAM
KV Cache VRAM = 2 × Layers × Hidden Dim × Seq Len × Batch Size × Precision BytesFactor 2 for K and V matrices. Example: LLaMA 70B (80 layers, 8192 hidden), seq len 2048, batch 1, FP16 → 2 × 80 × 8192 × 2048 × 1 × 2 ≈ 5GB
Multi-GPU Parallel Strategies
Data Parallelism (DP)
Each GPU loads full model
Different GPUs process different data
Suitable when model fits single GPU
Less used for inference
Tensor Parallelism (TP)
Split each layer across GPUs
Intra-layer parallelism, frequent communication
Requires high bandwidth (NVLink, NVSwitch)
Frameworks: Megatron-LM, vLLM
Pipeline Parallelism (PP)
Split model by layers across GPUs
Inter-layer parallelism, less communication
Pipeline bubbles (idle time) exist
Frameworks: GPipe, PipeDream
Hybrid Parallelism
Combine TP and PP
For massive models and multi-node
Complex config but optimal performance
GPU Interconnect Topology
NVLink
NVIDIA proprietary high-speed interconnect
Bandwidth: 300-900 GB/s (generation dependent)
Low latency, ideal for TP
PCIe
General-purpose interface
PCIe 4.0 x16 = 32 GB/s, PCIe 3.0 x16 = 16 GB/s
Higher latency, suitable for PP or infrequent TP
NVSwitch
NVIDIA high-speed switch
All-to-all GPU connectivity
Best for large-scale TP
Check topology: nvidia-smi topo -m Output legend: NV12 = NVLink 12 links (high bandwidth), SYS = PCIe + CPU interconnect (low), PHB = PCIe host bridge (medium)
Inference Framework Selection
Transformers (HuggingFace)
Pros: Simple, rich model zoo
Cons: Average perf, limited multi-GPU
Best for: Quick validation, small models
vLLM
Pros: High perf, PagedAttention, TP support
Cons: Limited model support
Best for: Production, large models, multi-user concurrency
TensorRT-LLM
Pros: Highest perf, TP+PP support
Cons: Complex config, long compile time
Best for: Extreme perf needs, NVIDIA GPUs
DeepSpeed-Inference
Pros: TP+PP support, flexible config
Cons: Complex config
Best for: Large models, hybrid parallelism
Megatron-LM
Pros: Good TP perf
Cons: Training-focused, limited inference
Best for: Research, custom inference
Practical Steps
Step 1: Environment Setup
1.1 Check GPU Info
nvidia-smi # GPU count/model
nvidia-smi topo -m # Topology
nvcc --version # CUDA version
nvidia-smi | grep "Driver Version"1.2 Install Dependencies
conda create -n llm python=3.10
conda activate llm
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
python -c "import torch; print(torch.__version__); print(torch.cuda.is_available()); print(torch.cuda.device_count())"1.3 Download Model
pip install huggingface_hub
huggingface-cli download meta-llama/Llama-2-70b-hf --local-dir ./models/llama-2-70b
# or
git lfs install
git clone https://huggingface.co/meta-llama/Llama-2-70b-hf ./models/llama-2-70bStep 2: Transformers + Accelerate
2.1 Install
pip install transformers accelerate2.2 Inference Code
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from accelerate import init_empty_weights, load_checkpoint_and_dispatch
model_path = "./models/llama-2-70b"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(
model_path,
torch_dtype=torch.float16,
device_map="auto",
low_cpu_mem_usage=True,
)
print(f"Model loaded to devices: {model.hf_device_map}")
prompt = "What is the capital of France?"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda:0")
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=100, temperature=0.7, top_p=0.9, do_sample=True)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"Response: {response}")2.3 Run
python multi_gpu_inference.pyExpected output shows model sharded across GPUs (e.g., layers 0-19 on GPU0, 20-39 on GPU1, etc.)
2.4 Monitor VRAM
nvidia-smi # Should show balanced VRAM across GPUs2.5 Custom Device Map
device_map = {
"model.embed_tokens": 0,
"model.layers.0": 0,
# ... first 20 layers on GPU 0
"model.layers.20": 1,
# ... layers 21-40 on GPU 1
"model.layers.40": 2,
# ... layers 41-60 on GPU 2
"model.layers.60": 3,
# ... layers 61-79 on GPU 3
"model.norm": 3,
"lm_head": 3,
}
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16, device_map=device_map, low_cpu_mem_usage=True)Step 3: vLLM
3.1 Install
pip install vllm
# or from source
git clone https://github.com/vllm-project/vllm.git
cd vllm
pip install -e .3.2 Launch TP Server
python -m vllm.entrypoints.openai.api_server \
--model ./models/llama-2-70b \
--tensor-parallel-size 4 \
--dtype float16 \
--max-model-len 4096 \
--port 8000Key args: --model (path), --tensor-parallel-size (GPU count), --dtype (float16/bfloat16), --max-model-len (max seq len), --port
3.3 Test Service
curl http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{"model": "./models/llama-2-70b", "prompt": "What is the capital of France?", "max_tokens": 100, "temperature": 0.7}'Or Python client (OpenAI compatible):
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")
response = client.completions.create(model="./models/llama-2-70b", prompt="What is the capital of France?", max_tokens=100, temperature=0.7)
print(response.choices[0].text)3.4 Performance Tuning
python -m vllm.entrypoints.openai.api_server \
--model ./models/llama-2-70b \
--tensor-parallel-size 4 \
--dtype float16 \
--max-model-len 4096 \
--gpu-memory-utilization 0.9 \
--max-num-seqs 256 \
--port 8000--gpu-memory-utilization: VRAM usage ratio (0-1, default 0.9). --max-num-seqs: max concurrent sequences (affects throughput).
Step 4: DeepSpeed-Inference
4.1 Install
pip install deepspeed4.2 Inference Code
import torch
import deepspeed
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "./models/llama-2-70b"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16)
ds_engine = deepspeed.init_inference(
model,
mp_size=4,
dtype=torch.float16,
replace_with_kernel_inject=True,
)
model = ds_engine.module
prompt = "What is the capital of France?"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=100, temperature=0.7, top_p=0.9, do_sample=True)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"Response: {response}")4.3 Run
deepspeed --num_gpus 4 deepspeed_inference.pyStep 5: Quantization
5.1 GPTQ
pip install auto-gptq
huggingface-cli download TheBloke/Llama-2-70B-GPTQ --local-dir ./models/llama-2-70b-gptqUsage: standard Transformers load with device_map="auto"
5.2 AWQ
pip install autoawq
huggingface-cli download TheBloke/Llama-2-70B-AWQ --local-dir ./models/llama-2-70b-awqUsage: from awq import AutoAWQForCausalLM; model = AutoAWQForCausalLM.from_quantized(path, fuse_layers=True, device_map="auto")
5.3 bitsandbytes (4-bit)
pip install bitsandbytes
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=torch.float16,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
)
model = AutoModelForCausalLM.from_pretrained(model_path, quantization_config=quantization_config, device_map="auto")Quantization Comparison
FP16 : 16-bit, 100% VRAM, baseline speed, 100% accuracy
INT8 : 8-bit, 50% VRAM, slightly slower, 99%+ accuracy
INT4 (GPTQ) : 4-bit, 25% VRAM, slower, 95%+ accuracy
INT4 (AWQ) : 4-bit, 25% VRAM, faster, 97%+ accuracy
Step 6: Performance Testing
6.1 Throughput Benchmark
# benchmark_throughput.py
import time, torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "./models/llama-2-70b"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16, device_map="auto")
prompts = ["What is the capital of France?"] * 100
# Warmup
for _ in range(5):
inputs = tokenizer(prompts[0], return_tensors="pt").to("cuda")
model.generate(**inputs, max_new_tokens=50)
start = time.time()
total_tokens = 0
for prompt in prompts:
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
outputs = model.generate(**inputs, max_new_tokens=50)
total_tokens += outputs.shape[1]
elapsed = time.time() - start
print(f"Throughput: {total_tokens/elapsed:.2f} tokens/sec")6.2 Latency Benchmark
# benchmark_latency.py
latencies = []
for _ in range(100):
start = time.time()
with torch.no_grad():
model.generate(**inputs, max_new_tokens=50)
latencies.append((time.time() - start) * 1000)
print(f"Avg: {sum(latencies)/len(latencies):.2f} ms")
print(f"P50: {sorted(latencies)[len(latencies)//2]:.2f} ms")
print(f"P95: {sorted(latencies)[int(len(latencies)*0.95)]:.2f} ms")
print(f"P99: {sorted(latencies)[int(len(latencies)*0.99)]:.2f} ms")6.3 GPU Monitoring
nvidia-smi dmon -s u -d 1
# or nvtop, gpustat6.4 Bottleneck Analysis
VRAM bottleneck : OOM errors → reduce batch size, quantize, add GPUs
Compute bottleneck : GPU util >90% → optimize model, upgrade GPU
Communication bottleneck : Low GPU util but multi-GPU slower → optimize TP config, use NVLink, reduce TP degree, use PP
I/O bottleneck : Low GPU util, high CPU util → increase preprocessing threads, faster storage
Step 7: Production Deployment
7.1 Docker
FROM nvidia/cuda:12.1.0-cudnn8-runtime-ubuntu22.04
RUN apt-get update && apt-get install -y python3.10 python3-pip git && rm -rf /var/lib/apt/lists/*
RUN pip3 install --no-cache-dir torch transformers accelerate vllm
WORKDIR /app
COPY models/ /app/models/
COPY inference.py /app/
EXPOSE 8000
CMD ["python3", "-m", "vllm.entrypoints.openai.api_server", "--model", "/app/models/llama-2-70b", "--tensor-parallel-size", "4", "--dtype", "float16", "--port", "8000"]Build: docker build -t llm-inference:latest . Run:
docker run --gpus all -p 8000:8000 llm-inference:latest7.2 Kubernetes
apiVersion: apps/v1
kind: Deployment
metadata:
name: llm-inference
spec:
replicas: 1
selector:
matchLabels:
app: llm-inference
template:
metadata:
labels:
app: llm-inference
spec:
containers:
- name: llm
image: llm-inference:latest
ports:
- containerPort: 8000
resources:
limits:
nvidia.com/gpu: 4
env:
- name: CUDA_VISIBLE_DEVICES
value: "0,1,2,3"
---
apiVersion: v1
kind: Service
metadata:
name: llm-inference
spec:
selector:
app: llm-inference
ports:
- protocol: TCP
port: 8000
targetPort: 8000
type: LoadBalancerDeploy:
kubectl apply -f llm-deployment.yaml7.3 HPA
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: llm-inference-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: llm-inference
minReplicas: 1
maxReplicas: 5
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 807.4 Monitoring
Prometheus scrape config for llm-inference:8000. Logs via ELK/Loki.
Common Commands
GPU Management
nvidia-smi
nvidia-smi topo -m
nvidia-smi dmon -s u -d 1
nvidia-smi pmon
export CUDA_VISIBLE_DEVICES=0,1,2,3Model Management
huggingface-cli download model_name --local-dir ./models/model_name
python -c "from transformers import AutoConfig; print(AutoConfig.from_pretrained('./models/model_name'))"
du -sh ./models/model_namePerformance Testing
python benchmark_throughput.py
python benchmark_latency.py
python -m vllm.entrypoints.openai.api_server --model ./models/llama-2-70b --tensor-parallel-size 4 --benchmarkRisk Warnings
High-Risk Operations
OOM : Process crash. Causes: model too large, batch size too large, KV cache too large. Fix: reduce batch size, quantize, add GPUs.
Multi-process conflict : GPU resource contention. Cause: multiple processes on same GPU. Fix: set CUDA_VISIBLE_DEVICES for isolation.
Communication overhead : Increased latency. Cause: low inter-GPU bandwidth, excessive TP degree. Fix: use NVLink, reduce TP degree, use PP.
Model corruption : Wrong outputs. Cause: incomplete download, weight loading errors. Fix: verify checksum, re-download.
Common Misconceptions
More GPUs = faster. False: communication overhead may hurt. Choose parallelism based on model size and topology.
TP degree = GPU count. False: TP degree should divide GPU count; prefer powers of 2 (2,4,8).
Quantization has no accuracy loss. False: trade-off between VRAM and accuracy.
Larger batch size = better. False: causes OOM; tune based on VRAM and perf needs.
Validation
Model Loading
model = AutoModelForCausalLM.from_pretrained("./models/llama-2-70b", torch_dtype=torch.float16, device_map="auto")
print(f"Model device map: {model.hf_device_map}")Inference Correctness
prompt = "1 + 1 ="
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=5)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"Response: {response}") # Should output: 1 + 1 = 2Performance
python benchmark_throughput.py # Should meet expected throughputRollback Plans
Fallback to Single GPU
export CUDA_VISIBLE_DEVICES=0
# or in code
device_map = {"": 0}
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16, device_map=device_map)Rollback Quantization
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16, device_map="auto")Production Considerations
Resource Planning
Plan GPU count by model size and QPS
Reserve VRAM for KV cache
Monitor GPU utilization to avoid waste
High Availability
Deploy multiple replicas
Use load balancing
Health checks and auto-restart
Monitoring & Alerting
GPU utilization, VRAM usage
Inference latency, throughput
Error rate, OOM count
Cost Optimization
Spot instances (cloud)
Quantization to reduce GPU count
Batching for higher throughput
Summary
Multi-GPU inference is essential when single GPU VRAM is insufficient. Key techniques:
Tensor Parallelism (TP) : Intra-layer, frequent comms, needs NVLink
Pipeline Parallelism (PP) : Inter-layer, less comms, works on PCIe
Hybrid Parallelism : TP+PP for massive models
Framework choice:
Transformers+Accelerate : Simple, quick validation
vLLM : High perf, production-ready
DeepSpeed-Inference : Flexible, large models
TensorRT-LLM : Peak perf on NVIDIA
Optimizations: Quantization (INT8/INT4), PagedAttention for KV cache, Flash Attention, Continuous Batching. Success requires matching parallelism strategy and framework to model characteristics, GPU topology, and business requirements.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Ops Community
A leading IT operations community where professionals share and grow together.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
