Multi-GPU LLM Inference: Step-by-Step Deployment Guide for Memory-Constrained Environments

This comprehensive tutorial covers multi-GPU inference for large language models, detailing VRAM estimation formulas, parallel strategies (tensor, pipeline, data), GPU interconnects (NVLink, PCIe), framework comparisons (vLLM, DeepSpeed, Transformers), quantization techniques (GPTQ, AWQ, bitsandbytes), performance benchmarking, and production deployment with Docker and Kubernetes.

Ops Community
Ops Community
Ops Community
Multi-GPU LLM Inference: Step-by-Step Deployment Guide for Memory-Constrained Environments

Problem Background

Large model parameter counts (GPT-3 175B, LLaMA 70B, Qwen 72B) exceed single GPU VRAM (24GB/48GB/80GB), making multi-GPU inference necessary. Target roles: algorithm engineers, ops engineers, MLOps engineers, R&D engineers.

Multi-GPU Inference Characteristics

Model parameters distributed across GPUs

Cross-GPU communication required

VRAM and compute scale linearly

Inference latency may increase

Configuration complexity rises

Applicable Scenarios

Single GPU VRAM insufficient for full model

Need higher inference throughput

Need lower per-request latency

Multi-user concurrent requests

Batch data processing

Non-Applicable Scenarios

Model fits single GPU (prefer single GPU)

Ultra-low latency requirements (multi-GPU adds communication overhead)

Low inter-GPU bandwidth (e.g., PCIe 3.0 x8)

Core Knowledge: VRAM Estimation

Inference VRAM components:

Model parameters : parameter count × precision bytes

Activations : forward pass intermediate results

KV Cache : autoregressive generation cache

Gradients/optimizer states : training only, not inference

Formula: Total VRAM ≈ Model Params VRAM + KV Cache VRAM + Activations VRAM + Framework Overhead

Model Parameter VRAM

Model Params VRAM = Parameter Count × Precision Bytes

Precision bytes: FP32=4, FP16/BF16=2, INT8=1, INT4=0.5

Examples: LLaMA 70B FP16 = 70B × 2 = 140GB; INT8 = 70GB; INT4 = 35GB

KV Cache VRAM

KV Cache VRAM = 2 × Layers × Hidden Dim × Seq Len × Batch Size × Precision Bytes

Factor 2 for K and V matrices. Example: LLaMA 70B (80 layers, 8192 hidden), seq len 2048, batch 1, FP16 → 2 × 80 × 8192 × 2048 × 1 × 2 ≈ 5GB

Multi-GPU Parallel Strategies

Data Parallelism (DP)

Each GPU loads full model

Different GPUs process different data

Suitable when model fits single GPU

Less used for inference

Tensor Parallelism (TP)

Split each layer across GPUs

Intra-layer parallelism, frequent communication

Requires high bandwidth (NVLink, NVSwitch)

Frameworks: Megatron-LM, vLLM

Pipeline Parallelism (PP)

Split model by layers across GPUs

Inter-layer parallelism, less communication

Pipeline bubbles (idle time) exist

Frameworks: GPipe, PipeDream

Hybrid Parallelism

Combine TP and PP

For massive models and multi-node

Complex config but optimal performance

GPU Interconnect Topology

NVLink

NVIDIA proprietary high-speed interconnect

Bandwidth: 300-900 GB/s (generation dependent)

Low latency, ideal for TP

PCIe

General-purpose interface

PCIe 4.0 x16 = 32 GB/s, PCIe 3.0 x16 = 16 GB/s

Higher latency, suitable for PP or infrequent TP

NVSwitch

NVIDIA high-speed switch

All-to-all GPU connectivity

Best for large-scale TP

Check topology: nvidia-smi topo -m Output legend: NV12 = NVLink 12 links (high bandwidth), SYS = PCIe + CPU interconnect (low), PHB = PCIe host bridge (medium)

Inference Framework Selection

Transformers (HuggingFace)

Pros: Simple, rich model zoo

Cons: Average perf, limited multi-GPU

Best for: Quick validation, small models

vLLM

Pros: High perf, PagedAttention, TP support

Cons: Limited model support

Best for: Production, large models, multi-user concurrency

TensorRT-LLM

Pros: Highest perf, TP+PP support

Cons: Complex config, long compile time

Best for: Extreme perf needs, NVIDIA GPUs

DeepSpeed-Inference

Pros: TP+PP support, flexible config

Cons: Complex config

Best for: Large models, hybrid parallelism

Megatron-LM

Pros: Good TP perf

Cons: Training-focused, limited inference

Best for: Research, custom inference

Practical Steps

Step 1: Environment Setup

1.1 Check GPU Info

nvidia-smi          # GPU count/model
nvidia-smi topo -m  # Topology
nvcc --version      # CUDA version
nvidia-smi | grep "Driver Version"

1.2 Install Dependencies

conda create -n llm python=3.10
conda activate llm
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
python -c "import torch; print(torch.__version__); print(torch.cuda.is_available()); print(torch.cuda.device_count())"

1.3 Download Model

pip install huggingface_hub
huggingface-cli download meta-llama/Llama-2-70b-hf --local-dir ./models/llama-2-70b
# or
git lfs install
git clone https://huggingface.co/meta-llama/Llama-2-70b-hf ./models/llama-2-70b

Step 2: Transformers + Accelerate

2.1 Install

pip install transformers accelerate

2.2 Inference Code

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from accelerate import init_empty_weights, load_checkpoint_and_dispatch

model_path = "./models/llama-2-70b"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(
    model_path,
    torch_dtype=torch.float16,
    device_map="auto",
    low_cpu_mem_usage=True,
)
print(f"Model loaded to devices: {model.hf_device_map}")

prompt = "What is the capital of France?"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda:0")
with torch.no_grad():
    outputs = model.generate(**inputs, max_new_tokens=100, temperature=0.7, top_p=0.9, do_sample=True)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"Response: {response}")

2.3 Run

python multi_gpu_inference.py

Expected output shows model sharded across GPUs (e.g., layers 0-19 on GPU0, 20-39 on GPU1, etc.)

2.4 Monitor VRAM

nvidia-smi  # Should show balanced VRAM across GPUs

2.5 Custom Device Map

device_map = {
    "model.embed_tokens": 0,
    "model.layers.0": 0,
    # ... first 20 layers on GPU 0
    "model.layers.20": 1,
    # ... layers 21-40 on GPU 1
    "model.layers.40": 2,
    # ... layers 41-60 on GPU 2
    "model.layers.60": 3,
    # ... layers 61-79 on GPU 3
    "model.norm": 3,
    "lm_head": 3,
}
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16, device_map=device_map, low_cpu_mem_usage=True)

Step 3: vLLM

3.1 Install

pip install vllm
# or from source
git clone https://github.com/vllm-project/vllm.git
cd vllm
pip install -e .

3.2 Launch TP Server

python -m vllm.entrypoints.openai.api_server \
    --model ./models/llama-2-70b \
    --tensor-parallel-size 4 \
    --dtype float16 \
    --max-model-len 4096 \
    --port 8000

Key args: --model (path), --tensor-parallel-size (GPU count), --dtype (float16/bfloat16), --max-model-len (max seq len), --port

3.3 Test Service

curl http://localhost:8000/v1/completions \
    -H "Content-Type: application/json" \
    -d '{"model": "./models/llama-2-70b", "prompt": "What is the capital of France?", "max_tokens": 100, "temperature": 0.7}'

Or Python client (OpenAI compatible):

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")
response = client.completions.create(model="./models/llama-2-70b", prompt="What is the capital of France?", max_tokens=100, temperature=0.7)
print(response.choices[0].text)

3.4 Performance Tuning

python -m vllm.entrypoints.openai.api_server \
    --model ./models/llama-2-70b \
    --tensor-parallel-size 4 \
    --dtype float16 \
    --max-model-len 4096 \
    --gpu-memory-utilization 0.9 \
    --max-num-seqs 256 \
    --port 8000

--gpu-memory-utilization: VRAM usage ratio (0-1, default 0.9). --max-num-seqs: max concurrent sequences (affects throughput).

Step 4: DeepSpeed-Inference

4.1 Install

pip install deepspeed

4.2 Inference Code

import torch
import deepspeed
from transformers import AutoModelForCausalLM, AutoTokenizer

model_path = "./models/llama-2-70b"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16)

ds_engine = deepspeed.init_inference(
    model,
    mp_size=4,
    dtype=torch.float16,
    replace_with_kernel_inject=True,
)
model = ds_engine.module

prompt = "What is the capital of France?"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
    outputs = model.generate(**inputs, max_new_tokens=100, temperature=0.7, top_p=0.9, do_sample=True)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"Response: {response}")

4.3 Run

deepspeed --num_gpus 4 deepspeed_inference.py

Step 5: Quantization

5.1 GPTQ

pip install auto-gptq
huggingface-cli download TheBloke/Llama-2-70B-GPTQ --local-dir ./models/llama-2-70b-gptq

Usage: standard Transformers load with device_map="auto"

5.2 AWQ

pip install autoawq
huggingface-cli download TheBloke/Llama-2-70B-AWQ --local-dir ./models/llama-2-70b-awq

Usage: from awq import AutoAWQForCausalLM; model = AutoAWQForCausalLM.from_quantized(path, fuse_layers=True, device_map="auto")

5.3 bitsandbytes (4-bit)

pip install bitsandbytes
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
quantization_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.float16,
    bnb_4bit_use_double_quant=True,
    bnb_4bit_quant_type="nf4",
)
model = AutoModelForCausalLM.from_pretrained(model_path, quantization_config=quantization_config, device_map="auto")

Quantization Comparison

FP16 : 16-bit, 100% VRAM, baseline speed, 100% accuracy

INT8 : 8-bit, 50% VRAM, slightly slower, 99%+ accuracy

INT4 (GPTQ) : 4-bit, 25% VRAM, slower, 95%+ accuracy

INT4 (AWQ) : 4-bit, 25% VRAM, faster, 97%+ accuracy

Step 6: Performance Testing

6.1 Throughput Benchmark

# benchmark_throughput.py
import time, torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_path = "./models/llama-2-70b"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16, device_map="auto")
prompts = ["What is the capital of France?"] * 100
# Warmup
for _ in range(5):
    inputs = tokenizer(prompts[0], return_tensors="pt").to("cuda")
    model.generate(**inputs, max_new_tokens=50)
start = time.time()
total_tokens = 0
for prompt in prompts:
    inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
    with torch.no_grad():
        outputs = model.generate(**inputs, max_new_tokens=50)
    total_tokens += outputs.shape[1]
elapsed = time.time() - start
print(f"Throughput: {total_tokens/elapsed:.2f} tokens/sec")

6.2 Latency Benchmark

# benchmark_latency.py
latencies = []
for _ in range(100):
    start = time.time()
    with torch.no_grad():
        model.generate(**inputs, max_new_tokens=50)
    latencies.append((time.time() - start) * 1000)
print(f"Avg: {sum(latencies)/len(latencies):.2f} ms")
print(f"P50: {sorted(latencies)[len(latencies)//2]:.2f} ms")
print(f"P95: {sorted(latencies)[int(len(latencies)*0.95)]:.2f} ms")
print(f"P99: {sorted(latencies)[int(len(latencies)*0.99)]:.2f} ms")

6.3 GPU Monitoring

nvidia-smi dmon -s u -d 1
# or nvtop, gpustat

6.4 Bottleneck Analysis

VRAM bottleneck : OOM errors → reduce batch size, quantize, add GPUs

Compute bottleneck : GPU util >90% → optimize model, upgrade GPU

Communication bottleneck : Low GPU util but multi-GPU slower → optimize TP config, use NVLink, reduce TP degree, use PP

I/O bottleneck : Low GPU util, high CPU util → increase preprocessing threads, faster storage

Step 7: Production Deployment

7.1 Docker

FROM nvidia/cuda:12.1.0-cudnn8-runtime-ubuntu22.04
RUN apt-get update && apt-get install -y python3.10 python3-pip git && rm -rf /var/lib/apt/lists/*
RUN pip3 install --no-cache-dir torch transformers accelerate vllm
WORKDIR /app
COPY models/ /app/models/
COPY inference.py /app/
EXPOSE 8000
CMD ["python3", "-m", "vllm.entrypoints.openai.api_server", "--model", "/app/models/llama-2-70b", "--tensor-parallel-size", "4", "--dtype", "float16", "--port", "8000"]

Build: docker build -t llm-inference:latest . Run:

docker run --gpus all -p 8000:8000 llm-inference:latest

7.2 Kubernetes

apiVersion: apps/v1
kind: Deployment
metadata:
  name: llm-inference
spec:
  replicas: 1
  selector:
    matchLabels:
      app: llm-inference
  template:
    metadata:
      labels:
        app: llm-inference
    spec:
      containers:
      - name: llm
        image: llm-inference:latest
        ports:
        - containerPort: 8000
        resources:
          limits:
            nvidia.com/gpu: 4
        env:
        - name: CUDA_VISIBLE_DEVICES
          value: "0,1,2,3"
---
apiVersion: v1
kind: Service
metadata:
  name: llm-inference
spec:
  selector:
    app: llm-inference
  ports:
  - protocol: TCP
    port: 8000
    targetPort: 8000
  type: LoadBalancer

Deploy:

kubectl apply -f llm-deployment.yaml

7.3 HPA

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: llm-inference-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: llm-inference
  minReplicas: 1
  maxReplicas: 5
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 80

7.4 Monitoring

Prometheus scrape config for llm-inference:8000. Logs via ELK/Loki.

Common Commands

GPU Management

nvidia-smi
nvidia-smi topo -m
nvidia-smi dmon -s u -d 1
nvidia-smi pmon
export CUDA_VISIBLE_DEVICES=0,1,2,3

Model Management

huggingface-cli download model_name --local-dir ./models/model_name
python -c "from transformers import AutoConfig; print(AutoConfig.from_pretrained('./models/model_name'))"
du -sh ./models/model_name

Performance Testing

python benchmark_throughput.py
python benchmark_latency.py
python -m vllm.entrypoints.openai.api_server --model ./models/llama-2-70b --tensor-parallel-size 4 --benchmark

Risk Warnings

High-Risk Operations

OOM : Process crash. Causes: model too large, batch size too large, KV cache too large. Fix: reduce batch size, quantize, add GPUs.

Multi-process conflict : GPU resource contention. Cause: multiple processes on same GPU. Fix: set CUDA_VISIBLE_DEVICES for isolation.

Communication overhead : Increased latency. Cause: low inter-GPU bandwidth, excessive TP degree. Fix: use NVLink, reduce TP degree, use PP.

Model corruption : Wrong outputs. Cause: incomplete download, weight loading errors. Fix: verify checksum, re-download.

Common Misconceptions

More GPUs = faster. False: communication overhead may hurt. Choose parallelism based on model size and topology.

TP degree = GPU count. False: TP degree should divide GPU count; prefer powers of 2 (2,4,8).

Quantization has no accuracy loss. False: trade-off between VRAM and accuracy.

Larger batch size = better. False: causes OOM; tune based on VRAM and perf needs.

Validation

Model Loading

model = AutoModelForCausalLM.from_pretrained("./models/llama-2-70b", torch_dtype=torch.float16, device_map="auto")
print(f"Model device map: {model.hf_device_map}")

Inference Correctness

prompt = "1 + 1 ="
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=5)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(f"Response: {response}")  # Should output: 1 + 1 = 2

Performance

python benchmark_throughput.py  # Should meet expected throughput

Rollback Plans

Fallback to Single GPU

export CUDA_VISIBLE_DEVICES=0
# or in code
device_map = {"": 0}
model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16, device_map=device_map)

Rollback Quantization

model = AutoModelForCausalLM.from_pretrained(model_path, torch_dtype=torch.float16, device_map="auto")

Production Considerations

Resource Planning

Plan GPU count by model size and QPS

Reserve VRAM for KV cache

Monitor GPU utilization to avoid waste

High Availability

Deploy multiple replicas

Use load balancing

Health checks and auto-restart

Monitoring & Alerting

GPU utilization, VRAM usage

Inference latency, throughput

Error rate, OOM count

Cost Optimization

Spot instances (cloud)

Quantization to reduce GPU count

Batching for higher throughput

Summary

Multi-GPU inference is essential when single GPU VRAM is insufficient. Key techniques:

Tensor Parallelism (TP) : Intra-layer, frequent comms, needs NVLink

Pipeline Parallelism (PP) : Inter-layer, less comms, works on PCIe

Hybrid Parallelism : TP+PP for massive models

Framework choice:

Transformers+Accelerate : Simple, quick validation

vLLM : High perf, production-ready

DeepSpeed-Inference : Flexible, large models

TensorRT-LLM : Peak perf on NVIDIA

Optimizations: Quantization (INT8/INT4), PagedAttention for KV cache, Flash Attention, Continuous Batching. Success requires matching parallelism strategy and framework to model characteristics, GPU topology, and business requirements.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

quantizationvLLMtensor parallelismDeepSpeedpipeline parallelismLLM deploymentGPU memory optimizationmulti-GPU inference
Ops Community
Written by

Ops Community

A leading IT operations community where professionals share and grow together.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.