Partitioning A100/H100 GPUs with MIG: Creation, Binding & Rollback for Multi-Tenant Workloads
This comprehensive guide covers NVIDIA MIG (Multi-Instance GPU) on A100/H100, detailing hardware-level GPU partitioning, instance creation and binding, Kubernetes/Docker integration, isolation verification, monitoring, and rollback procedures for production multi-tenant model serving.
1. Problem Background
Running large model inference on GPU servers presents a core challenge: maximizing the value of expensive A100/H100 GPUs. An A100-80G is costly; using it for a single 7B model wastes memory and compute. Multiple business units often need dedicated GPUs, leading to procurement pressure. However, naively running multiple models on one GPU causes SM contention, memory pressure, poor isolation, latency jitter, and difficult fault attribution.
Common first thoughts — Docker device limits or scheduler concurrency — do not provide hardware-level isolation. NVIDIA MIG (Multi-Instance GPU) solves this by physically partitioning a GPU into multiple independent GPU Instances (GIs), each with dedicated SMs, L2 cache, memory slices, and bandwidth. Instances are fully isolated and can be assigned to different users/containers.
2. Applicable Scenarios
You have A100/A30/H100/H200 (or other MIG-capable data center GPUs) and want to run multiple isolated model instances on one physical GPU.
Multiple teams/models share a GPU and require fixed memory, fixed compute, and fault isolation.
You need to offer independent GPU slices or API quotas to different customers/service tiers.
You are evaluating MIG vs. MPS (Multi-Process Service) for multi-tenant isolation.
You need to integrate MIG instances into Kubernetes so different Pods use different MIG instances for different models.
Critical reminder: MIG is not supported on all GPUs. Consumer cards (RTX series) and some data center cards (e.g., A16) lack support. The maximum instance count and available profiles depend on the specific GPU model and memory capacity — always verify with nvidia-smi mig -lgip before proceeding.
3. Core Concepts
3.1 What is MIG?
MIG is NVIDIA's hardware-level virtualization for data center GPUs. When enabled, a physical GPU can be split into multiple GPU Instances (GIs), each with independent SMs, L2 cache, memory slices, and bandwidth. On top of GIs, Compute Instances (CIs) can be created for finer-grained compute scheduling.
Key understanding: MIG isolates at the physical slice level — memory is hard-partitioned (instances cannot access each other's memory), and compute is allocated via dedicated SMs. This fundamentally differs from MPS, which shares compute and memory.
3.2 GPU Instance (GI) vs Compute Instance (CI)
GPU Instance: Unit for partitioning memory, L2 cache, and memory bandwidth. A GI has a fixed memory slice and corresponding SM group.
Compute Instance: Unit for partitioning scheduling units (compute engines). A GI can contain multiple CIs, each with independent compute capability slices.
For most "one model per instance" scenarios, one CI per GI suffices. Start with GIs only; add CIs only if finer compute granularity is needed.
3.3 Hardware Boundaries
The number and size of MIG instances depend on the GPU's Compute Capability and memory. Query with nvidia-smi mig -lgip for actual supported profiles. Example output for A100-80G:
GPU instance profiles: GPU 0 Name: NVIDIA A100-SXM4-80GB MIG 1g.10gb (ID 19) x8 MIG 2g.20gb (ID 20) x8 MIG 3g.40gb (ID 26) x8 MIG 4g.40gb (ID 21) x8 MIG 7g.40gb (ID 22) x8 MIG 7g.79gb (ID 23) x8The x8 indicates max instances of that profile on this GPU. H100 has similar 1g/2g/3g/4g/7g series. Never memorize profile counts; always query nvidia-smi mig -lgip . Different driver/firmware versions yield different outputs.
3.4 MIG and Heterogeneous/Multi-Process
A common misconception: enabling MIG does not automatically distribute processes across instances. Processes must explicitly target a MIG instance via CUDA_VISIBLE_DEVICES with the MIG UUID, or via Kubernetes MIG device plugin.
3.5 MPS vs MIG Positioning
MPS enables process-level compute sharing with no memory isolation; MIG provides hardware-level physical isolation of both memory and compute. For "one GPU, multiple models, each gets a small GPU", MIG is appropriate. For "multiple processes of one large application sharing a full GPU", MPS is better. Do not confuse them.
4. Overall Implementation Flow
Verify capability: Confirm GPU supports MIG, driver version satisfies requirements, firmware supports MIG. Skip this and everything fails.
Enable MIG mode: Turn on MIG mode at the physical host level; usually requires reboot or driver reload.
Create GPU/Compute Instances: Use nvidia-smi mig commands to create GIs/CIs, or define profiles for batch creation.
Validate instances: Confirm instance list, UUIDs, memory, compute independence, and CUDA visibility.
Integrate usage: Assign different models/containers to instances via CUDA_VISIBLE_DEVICES or K8s device plugin.
Monitor: Ensure monitoring tools (DCGM Exporter) report per-instance metrics.
Maintain: Manage, destroy, rollback, quota, and access control.
5. Practical Steps
5.1 Step 1: Confirm GPU MIG Support, Driver & Firmware
Three checks before starting: GPU model, MIG support, driver/firmware.
# View GPU model and UUID nvidia-smi -L # View driver and CUDA version nvidia-smi # View GPU supported capabilities (including MIG) nvidia-smi -q -d SUPPORTED_CLOCKSCheck current MIG status per GPU: nvidia-smi -q | grep -i "MIG Mode" If MIG Mode shows Disabled and the GPU doesn't support MIG, hardware partitioning is impossible. The most reliable way to confirm support is checking official specs and nvidia-smi mig -lgip (covered in 5.2). Data center GPUs (A100/A30/H100/H200) typically support MIG; consumer cards do not. Use recent stable data center drivers; old drivers/firmware may lack MIG configurations.
5.2 Step 2: Query Supported GPU Instance Profiles
Before enabling MIG, check available GI profiles: nvidia-smi mig -lgip Output lists available GPU Instance Profiles (example above). The x8 means max 8 instances of that profile. Query Compute Instance profiles: nvidia-smi mig -lgipc Lists CI profiles available for a given GI combination. Profile IDs and names are crucial for creation. If -lgip errors or returns empty, the GPU may not support MIG or the driver lacks support — investigate those first.
5.3 Step 3: Enable MIG Mode
MIG mode is a global switch that typically requires a reboot (or dynamic enable in some scenarios). Standard flow:
# Check current MIG mode nvidia-smi -q | grep -i "MIG Mode" # Enable MIG on GPU 0 (high risk, changes GPU usage, use maintenance window) nvidia-smi -i 0 -mig 1For multiple GPUs, use nvidia-smi mig -i <gpu_id> per GPU. Instance counts and profiles follow nvidia-smi mig --help. Warning: Enabling MIG prevents the GPU from being used in "full GPU" mode; existing containers/tasks expecting a whole GPU will find no device. Notify all dependent teams. Confirm after enable: nvidia-smi -q | grep -i "MIG Mode" # Expect Enabled If reboot required, reboot and re-verify.
5.4 Step 4: Create GPU Instances (GI)
After MIG mode is on, create GIs. Example: create a 1g.10gb profile GI on GPU 0:
nvidia-smi mig -i 0 -cgi 19 19is the profile ID from -lgip. -cgi = create GPU instance. Verify: nvidia-smi mig -lgi Create multiple GIs (comma-separated profile IDs): nvidia-smi mig -i 0 -cgi 19,19,19,19 Confirm with nvidia-smi -L which now lists MIG instances with independent UUIDs:
GPU 0: NVIDIA A100-SXM4-80GB (UUID: GPU-xxx) MIG 1g.10gb Device 0: (UUID: MIG-xxx-1) MIG 1g.10gb Device 1: (UUID: MIG-xxx-2)Each MIG UUID is the target for model/container binding.
5.5 Step 5: Create Compute Instances (CI, Optional)
If a GI needs further compute partitioning, create CIs. First check available CI profiles for a GI: nvidia-smi mig -lgipc Create CI (example):
nvidia-smi mig -i 0 -gi <gpu_instance_id> -cci <compute_profile_id> -gispecifies GI UUID or index; -cci = create compute instance. For single-model workloads, one GI + one CI (or just GI) is enough. Start simple with GI only; add CI later if needed. Full verification: nvidia-smi mig -lgi nvidia-smi -q Confirm each MIG instance's memory, SMs, and usage match expectations.
5.6 Step 6: Run Different Models on Different MIG Instances
Bind models/processes to specific instances using CUDA_VISIBLE_DEVICES with the MIG UUID. Get UUIDs: nvidia-smi -L Bind process: CUDA_VISIBLE_DEVICES=MIG-xxxxx python run_model.py For vLLM inference service:
CUDA_VISIBLE_DEVICES=MIG-xxxxx vllm serve <model> --max-model-len 4096With Docker, use NVIDIA_VISIBLE_DEVICES or --gpus with nvidia-container-toolkit: docker run --gpus 'device=MIG-xxxxx' <image> ... Or NVIDIA_VISIBLE_DEVICES=<uuid>. Ensure toolkit version supports MIG and syntax matches your container runtime.
5.7 Step 7: Kubernetes Integration
Use NVIDIA device plugin with MIG strategy. Key concepts:
Configure device plugin on K8s nodes with MIG granularity policy (e.g., single or mixed mode in plugin config).
Pods request corresponding MIG resources via resource declarations.
Different MIG instances can be scheduled to different Pods/namespaces/teams.
Specific K8s YAML and plugin config depend on your device plugin version — do not copy fields blindly.
5.8 Step 8: Verify Isolation
MIG's value is isolation; verify three aspects: memory isolation, compute independence, and fault containment.
Memory isolation: Confirm each instance's memory usage is independent; one instance filling memory doesn't affect another's available memory. nvidia-smi # MIG mode shows per-instance memory usage Compute independence: Run a stress task on instance A; instance B's utilization should remain independent, latency stable. Critical test: Saturate instance A's compute, observe instance B's P99 latency. If B jitters, isolation failed or config is wrong — revisit configuration.
5.9 Step 9: Manage and Destroy MIG Instances
Track which process occupies which instance. View instance usage:
nvidia-smi mig -lgi nvidia-smi --query-mig-devices=index,gi_name,gi_slice_count,ci_name --format=csvBefore destroying, ensure no process is using the instance; otherwise data corruption or crashes may occur:
# Check for processes on instances nvidia-smi --query-compute-apps=pid,used_gpu_memory --format=csvDestroy specific GI (confirm idle first):
nvidia-smi mig -i 0 -dgi <gpu_instance_id> # Destroy all GIs on GPU 0 nvidia-smi mig -i 0 -dgi -gi <gpu_instance_id>To fully clean up and restore full GPU mode, destroy all instances then disable MIG:
nvidia-smi mig -i 0 -dgi -ci # Destroy instances (confirm idle first) nvidia-smi -i 0 -mig 0 # Disable MIG mode (high risk, affects whole GPU)Both destroy and disable remove devices; require no active usage, maintenance window, and rollback plan.
5.10 Step 10: Plan Instance Quotas & Ownership
Multi-team MIG requires technical + process controls:
Technical: Restrict accounts that can run nvidia-smi mig -mig/-cgi/-dgi; regular accounts get read-only.
Process: Instance creation/deletion via change requests recording time, GPU, profile, purpose, owner.
Ledger: Document each GPU's instance inventory, owning team, SLA, owner; reconcile regularly.
Quota planning: don't size instances to exact memory/compute limits — leave headroom. Assign different profiles per service tier to prevent high-priority tasks being starved by low-priority ones. Planning is the first line of defense for production MIG.
5.11 Step 11: MIG & Multi-GPU/Multi-Node Perspective
MIG solves "single GPU, multiple models" but coexists with multi-GPU planning. Consider holistically:
High compute/large memory needs: full GPU or large profile.
Medium compute, many instances: MIG split into medium profiles.
Hybrid: same node mixes full-GPU and MIG-partitioned GPUs.
Estimate each model's memory/compute needs first, then decide full GPU vs MIG and profile, then map across GPUs. Avoid siloed decisions.
5.12 Step 12: Service Registration Habit
With many models/instances, register each MIG instance: service name, port, owner, monitoring dashboard. Treat each MIG instance as a "small GPU device" in CMDB. Ensures operational continuity despite personnel changes.
5.13 Step 13: Troubleshooting Example — "Model Times Out Nightly"
Symptom: Model on MIG instance shows throughput drop and occasional timeouts at night, but physical GPU appears idle.
Initial hypothesis: MIG isolation issue, neighbor contention, or instance capacity insufficient.
Commands:
nvidia-smi -L nvidia-smi --query-mig-devices=MIG_index,utilization.gpu,utilization.memory,memory.used --format=csv -l 5Key metrics: Instance's utilization.memory frequently maxed, utilization.gpu also high → instance's own capacity saturated by business demand, not neighbor theft.
Root cause: Instance compute/bandwidth insufficient for model, or concurrency spike exceeds planned capacity; unrelated to neighbors.
Fix: Assign larger MIG profile, or rate-limit during peaks, or split workload. Confirm capacity planning gap, not leak.
Verification: After profile change, instance latency normalizes, utilization.memory no longer sustained at max.
Rollback: Revert to original profile; roll back service params.
Retrospective: Incorporate "model demand growth" into instance capacity reviews; periodically reassess MIG planning to avoid gradual saturation.
5.14 Step 10 Supplement: Per-Instance Alerting Strategy
MIG alerts must be per-instance, not just per physical GPU:
Instance utilization chronically exceeds planned ceiling.
Instance memory approaches allocated limit.
Instance chronically "starved" (long-term below expected).
Physical GPU temperature/power still alerted at GPU level.
Thresholds based on per-instance long-term baselines (P95/P99), not hardcoded. Per-instance alerting catches "physical GPU looks fine but one instance is near saturation" early.
5.15 Complete MIG Planning & Multi-GPU Coordination Method
Inventory demand: List all models, record each model's memory and compute requirements.
Map to GPUs: Assign demands to available GPUs; decide which GPUs stay full, which get MIG-split.
Cross-GPU balance: Don't concentrate traffic on one GPU; avoid fragmenting one GPU while others sit idle.
Leave headroom: Reserve capacity on each GPU and instance for variance and growth.
Document ledger: Record each GPU's partitioning scheme.
Multi-GPU coordination value: over-partitioning one A100 into too many tiny slices hurts per-slice performance; planning must be holistic, not just "can we slice?".
5.16 MIG & CUDA_VISIBLE_DEVICES Interaction Details
Understanding CUDA_VISIBLE_DEVICES behavior under MIG is key to binding isolation:
Bind by instance UUID: CUDA_VISIBLE_DEVICES=MIG-xxxx bash run_model.sh restricts process to that single instance.
Bind by index: MIG mode index semantics differ from full GPU; UUID is more stable.
Containers use same UUID env var or device parameter.
Wrong binding causes process to see wrong device or none. Validate binding on test instances before production. Binding is a hard constraint — don't guess.
5.17 MIG Instance Capacity Review & Periodic Assessment
Instances aren't set-and-forget; demand grows. Periodic assessment:
Regularly observe per-instance utilization and memory curves.
If an instance persistently saturates or idles, consider replanning.
Planning changes go through change control: add/remove instances, swap profiles — all require validation and rollback.
Keep ledger current.
Periodic review prevents "demand growth silently saturates instance" or "instance chronically wasted". Capacity is dynamic; ledger and assessment must move with it.
5.18 Case Study: Recurrent OOM on an Instance
Symptom: Model on MIG instance intermittently hits CUDA OOM; restart recovers.
Initial hypothesis: Instance memory under-provisioned, concurrency spike, or fragmentation.
Commands: nvidia-smi for instance memory, nvidia-smi mig -lgi for instance details, app logs for allocation failures.
Key metric: Instance memory peaks at limit, OOMs align with concurrency peaks.
Root cause: Instance memory allocation too tight + concurrency surge; confirmed by memory curve hitting ceiling aligned with error timestamps.
Fix: Assign larger profile, or rate-limit concurrency, or split workload.
Verification: After profile change, instance memory stays within safe margin, stable.
Rollback: Revert to original profile.
Retrospective: Base instance capacity on peak baseline; add rate-limiting as safeguard.
5.19 Incorporate MIG Ops into Daily Rhythm
Daily: Per-instance utilization and memory patrol.
Weekly: Compare baselines, check trends.
Post-change: Re-establish baselines.
Long-term: Integrate monitoring, set alerts, maintain ledger.
5.20 MIG & Multi-Model Service Governance
With many models/instances, governance must keep pace:
Each instance maps to service name, port, owner.
Access control and quotas.
Monitoring and alerting per service.
Changes and handovers per service.
Governance turns "instances" into managed "services", avoiding "who built it uses it" chaos.
6. Monitoring & Metrics Observation
MIG mode requires per-instance metric reporting. nvidia-smi supports per-instance views:
nvidia-smi nvidia-smi --query-mig-devices=gpu_uuid,mig_dev_index,mig_profile,utilization.gpu,utilization.memory,memory.used --format=csvDCGM also reports MIG instance metrics, but metric names and dimensions depend on actual exporter version and collection config — don't rely on outdated field names. Key MIG monitoring: per-instance utilization, memory, temperature (physical GPU level), and Xid errors.
Feeding per-instance utilization, memory, and usage into Prometheus enables per-model-instance alerting and capacity management — critical for production.
7. Configuration Examples
7.1 MIG with Docker
nvidia-container-toolkit exposes GPU/MIG instances to containers. Key config in Docker daemon and toolkit config for device support. Example direction:
{ "nvidia-container-runtime": { "nvidia-visible-devices": "all" } }Container specifies instance via env var:
docker run --rm --env NVIDIA_VISIBLE_DEVICES=<mig_uuid> nvcr.io/nvidia/cuda:12.2-base-ubuntu22.04 nvidia-smi -LToolkit/driver version differences affect MIG support — validate against your environment's docs; this is directional only.
7.2 MIG with Kubernetes
K8s integration relies on device plugin. Two main steps:
Deploy device plugin DaemonSet; configure MIG strategy via plugin config (e.g., single per profile or mixed; exact fields per plugin version).
Pods request instances via resource declarations (example below).
apiVersion: v1 kind: Pod metadata: name: model-a spec: containers: - name: model-a image: my-registry/model-a:1.0 resources: limits: nvidia.com/gpu: 1 env: - name: NVIDIA_VISIBLE_DEVICES value: "all"Note: This is illustrative; MIG resource declaration and granularity control fields depend on the device plugin's documentation — do not copy fields assuming production-ready . K8s + MIG is an integration action; follow plugin docs for canary rollout.
8. Log & Metric Observation Methods
Key MIG observation sources: nvidia-smi -q for MIG Mode status. nvidia-smi mig -lgi for instance list and occupancy.
Kernel logs for MIG-related dmesg.
Per-instance independent utilization, memory, and physical GPU temperature/power.
Business-side throughput, latency, error rates (to verify isolation reality).
Observation principles:
Sample at instance granularity, not just physical GPU.
Align "instance utilization" with "model throughput/latency on that instance" on same timeline.
Monitor long-term instance occupancy trends for capacity planning.
Xid errors still logged at physical GPU level; include in ledger under MIG mode.
Metric analysis must define normal ranges. An instance running consistently at its model's demand is normal; if a sibling instance is busy and this one jitters abnormally, isolation may be broken — investigate.
9. Troubleshooting Path: Complete Partitioning Loop
Loop from "request MIG for two models" to "verify isolation, go production":
Goal: One A100-80G runs two models, each ~half memory/compute, no interference.
Environment: A100-80G, healthy driver, nvidia-container-toolkit or bare metal, maintenance window available.
Pre-check: nvidia-smi -L confirm GPU; nvidia-smi -q | grep -i "MIG Mode" confirm disabled; nvidia-smi mig -lgip query profiles.
Implementation: nvidia-smi -i 0 -mig 1 enable MIG (maintenance window), nvidia-smi mig -i 0 -cgi <id> create two GIs, nvidia-smi -L get UUIDs, launch two models with CUDA_VISIBLE_DEVICES=<uuid>.
Config doc: Explicit per-instance profile, model assignment, team ownership.
Effect mechanism: MIG mode/instances effective immediately on that GPU; containers/K8s must correctly target instances.
Validation: nvidia-smi confirms both instances present, memory independent; stress instance A, verify instance B P99 stable.
Risks: MIG changes full-GPU behavior, instance destruction risky, maintenance window mandatory.
Rollback: Destroy instances, disable MIG to restore full GPU.
Maintenance advice: Per-instance monitoring, quota/permission control, ledger recording.
Only when the full loop completes — one GPU stably serving two models as "two isolated instances" — is it truly production-ready.
10. Risk Warnings
High-risk MIG operations and mitigations:
Enable/disable MIG mode: Changes whole GPU behavior; mandatory maintenance window, notify all dependent teams. Backup current instance/usage records first.
Destroy MIG instance ( nvidia-smi mig -dgi ): Confirm idle first; release any occupying processes to avoid data corruption.
Create instances: Consumes GPU memory/compute slices; don't over-create and exhaust capacity. Plan then create.
Batch create multiple instances: Define purpose per instance to avoid wasted slices; batch commands must be traceable and rollback-able.
Driver reinstall/upgrade: Affects MIG support; follow driver upgrade process with rollback plan.
K8s integration: Device plugin version matching issues; canary rollout.
Iron law for high-risk ops: Any command changing MIG mode/instances — confirm no occupancy, reserve maintenance window, record change, prepare rollback.
11. Verification Methods
After MIG enable: nvidia-smi -q | grep -i "MIG Mode" shows Enabled.
After instance creation: nvidia-smi -L lists expected MIG UUID count.
After binding: inside container/process, CUDA_VISIBLE_DEVICES=<uuid> nvidia-smi -L shows only target instance.
Isolation verification: stress one instance, other instance's utilization and latency remain independent and stable.
After destruction: nvidia-smi mig -lgi empty, state restored as expected.
K8s: Pod's nvidia-smi -L sees corresponding instance, resource declaration effective.
Verification isn't just "it runs"; isolation and metrics must meet expectations, with evidence retained.
12. Rollback Plans
Instance rollback: Destroy extra instances to revert to previous working instance set; re-validate.
Full GPU rollback: Destroy all instances → nvidia-smi -i 0 -mig 0 disable MIG → restore full GPU mode → rebind with full GPU method → verify full GPU works.
K8s rollback: Remove device plugin / revert resource declarations; Pods fall back to full GPU.
Driver rollback: Restore previous driver and re-verify MIG support.
Rollback must achieve "one-click return to last known good state"; post-rollback re-validation mandatory.
13. Production Considerations
All MIG changes require execution window, business notification, impact scope documentation.
Instance planning (which model, profile, team per instance) must land in ledger to avoid "who built it uses it" ambiguity.
MIG instances must integrate monitoring with per-instance alerting; thresholds based on actual model load baselines, not hardcoded.
Permission control: accounts executing -mig / -cgi / -dgi strictly limited.
Multi-GPU MIG planning requires cross-GPU coordination; don't overpack one GPU while others idle.
Handover docs must clarify "which GPUs have MIG enabled, what instances exist, who uses them".
Advanced Troubleshooting, Instance & Handover Checklists
Complete "Single GPU, Two Models" Zero-to-Production Script (Example Skeleton)
Below is an "example skeleton" script illustrating key steps and risks; not guaranteed character-for-character executable. Adapt to your GPU profile IDs and driver version:
#!/usr/bin/env bash # Example skeleton: A100-80G single GPU split into two isolated instances # Dangerous actions (enable/disable MIG, destroy instances) MUST run in maintenance window, backup usage records first # Specific profile IDs MUST be queried via nvidia-smi mig -lgip first; do not hardcode GPU=0 # Step 1: Confirm GPU & MIG support nvidia-smi -L nvidia-smi -q | grep -i "MIG Mode" # Step 2 (maintenance window): Enable MIG (depends on driver whether reboot needed) # nvidia-smi -i "$GPU" -mig 1 # Step 3: Query available profiles, pick your profile id # nvidia-smi mig -i "$GPU" -lgip # Step 4: Create two instances (ids per actual query, placeholder here) # nvidia-smi mig -i "$GPU" -cgi <profile_id> # nvidia-smi mig -i "$GPU" -cgi <profile_id> # Step 5: Get two instance UUIDs nvidia-smi -L # Step 6: Launch models with respective UUIDs (example) # CUDA_VISIBLE_DEVICES=<uuid1> vllm serve <model1> & # CUDA_VISIBLE_DEVICES=<uuid2> vllm serve <model2> & # Step 7: Verify isolation and record ledger nvidia-smiPlaceholders <profile_id>, <uuid1>, <model1> must be replaced with actual values. This only sequences the flow; production requires step-by-step, maintenance window, per-step verification — do not run blindly.
MIG FAQ
Q1: After enabling MIG, why can't original full-GPU containers find the GPU? MIG mode changes whole GPU usage; regular full-GPU processes can't locate a usable device. Must switch to instance UUID binding. This is expected MIG behavior — notify dependent teams in advance.
Q2: How many instances can MIG create? Depends on GPU model/memory/driver; query with nvidia-smi mig -lgip — don't memorize numbers.
Q3: MIG or MPS? Need hardware isolation, multiple models each owning a "small GPU" → MIG. Large app's multiple processes sharing whole GPU → MPS. Isolation is the dividing line.
Q4: Does destroying an instance lose data? Normally doesn't delete model weights (application-managed), but must confirm no process occupies instance before destroy; model weights stored on application side, don't rely on instance persistence.
Q5: Can MIG instances connect to Prometheus? Yes, use MIG-capable monitoring/DCGM config to collect per-instance metrics; metric names per actual version.
Q6: Why does my GPU's nvidia-smi mig -lgip have no output? GPU doesn't support MIG, driver lacks support, or MIG not properly enabled. Check GPU model, driver, MIG Mode status first.
Reading nvidia-smi Output Structure Under MIG
MIG-enabled nvidia-smi output structure differs completely from full GPU. Key reading points:
Physical GPU top row and MIG instance rows are separate.
MIG instance rows show MIG 1g.10gb identifier and independent memory display. Compute M. column shows Default (MIG disabled) or Enabled.
Memory usage accumulates per instance.
Structured query more reliable:
nvidia-smi --query-mig-devices=index,gi_name,ci_name,utilization.gpu,utilization.memory,memory.used --format=csvReading MIG output is a fundamental skill; you must parse it to pinpoint each instance's state.
Complete MIG + Docker Integration Example
Full MIG + Docker integration example (per your environment's nvidia-container-toolkit version):
# Verify container toolchain sees MIG devices (per toolkit version) docker run --rm --gpus all nvcr.io/nvidia/cuda:12.2-base nvidia-smi -L # Assign single MIG instance to container docker run --rm --env NVIDIA_VISIBLE_DEVICES=MIG-xxxxx nvcr.io/nvidia/cuda:12.2-base nvidia-smi -LNotes: MIG device visibility depends on toolkit's MIG support — version must match. Use NVIDIA_VISIBLE_DEVICES with instance UUID. Verify container's nvidia-smi -L shows only target instance before running model. Container integration per your toolkit version docs; validate first.
Deep Dive: MIG + Kubernetes Integration
K8s integration via device plugin. Deep dive:
Device plugin exposes MIG instances as schedulable resources, using official or community plugin.
MIG granularity config (e.g., whole GPU one profile or mixed) declared in plugin config; fields per plugin version.
Pod declares resources and uses env var to make container see corresponding instance.
# Illustrative YAML: Pod requests GPU resource (exact fields per device plugin version) apiVersion: v1 kind: Pod metadata: name: model-a spec: containers: - name: model-a image: my-registry/model-a:1.0 resources: limits: nvidia.com/gpu: 1 env: - name: NVIDIA_VISIBLE_DEVICES value: "all"Note: Illustrative only; MIG device plugin resource declaration and granularity control fields per its documentation — do not copy as production-ready . K8s + MIG is integration; follow plugin for canary rollout and rollback.
MIG Instance Profile Selection Principles
Don't just pick smallest profile. Principles:
First measure model memory and compute needs.
Cross-reference -lgip available profiles; pick "fits with headroom".
Don't over-slice; too small incurs performance penalty.
Target "just enough + headroom" for stability.
Profile principle: "sufficient, with headroom, not minimal". Validate on your actual GPU.
MIG Monitoring Production Practices
Real-world MIG monitoring:
Use MIG-capable exporter/monitoring config to collect per-instance utilization and memory; metrics per actual output.
Physical GPU temperature/power alerted at GPU level; instances alerted on utilization/memory.
Dashboards organized as "physical GPU overview + instance details".
Thresholds based on instance baselines.
Monitoring implementation per your exporter version's actual output; verify before building dashboards.
More MIG FAQ
Q1: Can MIG instance migrate to another GPU? Must recreate instance and reload model; not "move memory". Rebuild via ledger.
Q2: Can I still use full GPU under MIG mode? Enabling MIG changes whole GPU usage; typically cannot use full GPU anymore. Need full GPU → disable MIG.
Q3: Does MIG significantly hurt performance? Properly configured, overhead minimal; but over-slicing or poor adaptation causes penalty — benchmark to decide.
Q4: What's the source of truth for metric names? Your exporter's actual output; don't hardcode fields.
Q5: How many instances is optimal? Depends on models and GPU; stop when sufficient, leave headroom.
MIG Node Handover Checklist
Tick each item:
[ ] Confirm GPU supports MIG, driver/firmware compatible.
[ ] Record which GPUs have MIG enabled and instances created.
[ ] Record each instance's assigned model/team/service.
[ ] Per-instance monitoring and alerting connected (thresholds baseline-based).
[ ] Pre-MIG config and usage records backed up.
[ ] Destroy/close MIG rollback commands documented.
[ ] Maintenance window and responsible accounts defined.
Why Some GPUs Can Partition and Others Can't
Common confusion: "same A100, different behavior". Judgment points:
Memory capacity differs (40G vs 80G) → different available profiles.
Driver/firmware versions differ → MIG capabilities may differ.
Some GPUs have vendor-disabled or unsupported MIG at BIOS/firmware level; verify specs.
System config differences (multi-instance GPU requires physical GPU firmware support).
Judgment method: confirm GPU model via official specs → nvidia-smi mig -lgip actual test → record in ledger. Don't force another environment's profile onto your GPU. "Can it partition?" must be self-verified per environment.
MIG Instance Capacity Planning Calculation Example
Goal: One A100-80G, run one memory-heavy model and one lighter model; which two profiles?
Pre-check: nvidia-smi mig -lgip query actual supported combos; see if two profiles fit simultaneously.
Logic: First quantify each model's memory need (weights + KV cache + headroom), then cross-reference -lgip for each profile's memory and SMs, select combo that fits with headroom. Note: sum of multiple GIs' memory and SMs on same GPU must not exceed GPU capacity.
Implementation: Plan each model's allocation first, then create, bind each model to its instance, verify isolation.
Validation: Both models run simultaneously; confirm independent memory/compute, latency meets expectations.
Rollback: Destroy one to revert, or return to original plan.
Capacity planning core: "measure demand first, pick profile, validate, record" — don't slice by feel.
MIG Security Isolation Boundaries
MIG provides hardware memory and compute isolation, but not necessarily "full security boundary". Clarify boundaries to teams:
Memory: physically isolated between instances; one instance cannot read another's memory.
Compute: independent SMs; no contention.
But driver, physical GPU remain shared; firmware/driver vulnerabilities or management layer need additional protection.
Production advice: sensitive models/customers should use stronger isolation (even dedicated physical GPUs or cloud isolated instances). MIG suits shared but not highly sensitive scenarios. Communicate boundaries clearly to avoid "MIG enabled = fully isolated" misconception.
MIG & Data Lifecycle
MIG instance destruction does not save models/data — a common misunderstanding. Key points:
Instances have no "persistent weights"; weights/checkpoints are application's responsibility to persist to disk/object storage.
Before destroying instance, ensure model saved; don't rely on instance retention.
Migrating model to another GPU: use model files + reload on new instance, not move memory.
Data safety iron law: Data in GPU memory is volatile; persist intermediate states promptly; MIG changes must ensure data already persisted.
MIG Backup & Recovery Approach
MIG "backup" is not memory backup, but "planning" backup: instance inventory, profiles, UUIDs, purposes, owners, monitoring entries. Recovery approach:
Maintain complete MIG ledger per GPU (profile, assignee, quota).
On GPU replacement, recreate instances per ledger and bind same models.
Post-recovery, re-run isolation validation and monitoring checks.
This ledger is MIG's "backup" — documented and verifiable for reliable recovery.
MIG & Cost Attribution Method
Multi-team GPU sharing needs clear cost attribution. Method:
Each team/model maps to fixed MIG instance; instance usage = cost burden.
Use instance monitoring to track each team's occupancy duration and capacity for chargeback.
Ledger records each instance's ownership for finance/ops reconciliation.
Binding cost to instance reflects isolation value and avoids "shared GPU, unclear who uses more".
Common Errors & First Response
Symptom: -lgip no output → First Response: GPU doesn't support MIG or driver/firmware issue; confirm GPU model and versions first.
Symptom: After MIG enable, regular containers can't find GPU → First Response: Must bind via instance UUID.
Symptom: Destroy instance reports busy → First Response: Confirm no process occupying before destroy.
Symptom: Instance utilization "looks shared" → First Response: Verify true isolation via stress test.
Symptom: K8s can't schedule MIG → First Response: Check device plugin config and resource declaration.
First response isn't guessing — it's check commands and ledger first. On error, gather evidence before acting.
MIG Handover Document Template
Minimum handover must include:
Which GPUs have MIG enabled, which don't.
Each GPU's instance list (profile, UUID).
Each instance's assigned model/team/service.
Monitoring dashboards and alert thresholds.
Destroy/close/rollback commands.
Maintenance window and responsible accounts.
Complete handover docs ensure MIG node operations don't lose control due to personnel changes.
MIG Profile Naming & Memory Mapping Explained
Newcomers often confused by 1g.10gb, 2g.20gb, 7g.79gb naming. Pattern: leading number (1g, 2g, 7g) reflects GPU slice partitioning combination and SM/L2 ownership; trailing capacity (10gb, 20gb, 79gb) is memory slice size. Naming per your GPU's nvidia-smi mig -lgip actual output — different generations/memory capacities have different available names and combos.
Understanding naming lets you "read" a profile's approximate compute and memory during planning, then match to model needs. But production always relies on actual measured output; don't memorize names.
Must-Confirm Before Enabling MIG on a GPU
Enabling MIG is a major change altering whole GPU usage. Confirm each before executing:
GPU currently not running critical tasks.
All teams relying on full-GPU usage notified.
Ledger and usage records backed up.
Maintenance window confirmed.
Driver/firmware confirmed MIG-compatible.
Rollback plan confirmed (disable MIG to restore full GPU).
All six confirmed → then run nvidia-smi -i 0 -mig 1.
Production Example: One H100 Serving Two Small Models
Assume MIG-capable H100; Model A and B weights fit chosen profiles. Step 1: Estimate profile per model using stable memory, startup peak, max input length, concurrency, KV cache. Step 2: Cross-check hardware's actual supported instance list. Step 3: Drain device and switch mode. Step 4: Assign different MIG UUIDs per model. Don't just pick 1g profile because "7B FP16 weights ~14GB"; beyond weights there are activations, runtime, KV cache; startup profiling may exceed steady state.
nvidia-smi -L nvidia-smi mig -i 0 -lgip nvidia-smi -i 0 -q | sed -n '/MIG Mode/,+5p' # Record current instances and running processes before change nvidia-smi --query-compute-apps=gpu_uuid,pid,process_name,used_gpu_memory --format=csv nvidia-smi mig -i 0 -lgi nvidia-smi mig -i 0 -lci # After confirming GPU 0 drained, enable MIG mode sudo nvidia-smi -i 0 -mig 1 # Only use example profile if current GPU's -lgip explicitly supports it sudo nvidia-smi mig -i 0 -cgi 1g.10gb -C nvidia-smi -LActual profile may be H100 1g.10gb, A100 1g.10gb, or different memory — cannot copy example name. -C creates default CI after GI; if CI omitted, application's CUDA device may not match expectation. A100 mode switch may require GPU reset; H100 persistence behavior follows official deployment guide and on-site driver state; after change, reboot host and re-verify instances persist.
Three Entry Points: Bare Metal Binding, Containers, K8s
Bare metal processes bind MIG UUID directly; containers need NVIDIA Container Toolkit to pass corresponding device; Kubernetes uses device plugin's exposed MIG resource names. Three paths must not be stacked speculatively — e.g., if Pod already assigned MIG resource by plugin, don't copy a GPU index from host and force-bind inside container.
nvidia-smi -L # Save full MIG-... UUID CUDA_VISIBLE_DEVICES=MIG-<UUID_A> python /opt/model-a/server.py # Container scenario: verify device mapping first, don't launch real production model here docker run --rm --gpus 'device=MIG-<UUID_A>' nvidia/cuda:<supported-tag> nvidia-smi -L kubectl describe node <gpu-node> | sed -n '/Capacity:/,/Allocated resources:/p' # Resource name illustrative only; actual per device plugin's migStrategy and describe node results resources: limits: nvidia.com/mig-1g.10gb: 1If image, runtime, driver combo differ, example Docker command may need adjustment. Post-deploy, from each process/Pod run nvidia-smi -L and record actual UUID; verify both services load, restart independently, rate-limit independently. MIG provides hardware resource partitioning and some fault isolation, but not full multi-tenant security boundary; filesystem, network, process permissions, API auth still need separate config.
Why Throughput May Drop After Partitioning
Instance owns a fraction of GPU's SMs, memory, and bandwidth; actual proportions per profile constrained by GPU spec. On full GPU, model can use all resources regardless of concurrency; split into two profiles, single model's max throughput may drop, though combined GPU utilization may rise. MIG suits scenarios with clear isolation needs and small models leaving idle resources. Large model weights or long-context peaks exceeding single instance capacity make MIG partitioning infeasible.
# Compare two shapes: full GPU single service vs dual MIG services — test business latency and stable concurrency nvidia-smi --query-gpu=timestamp,index,memory.used,utilization.gpu --format=csv -l 1 # Each model checks its own API results curl -fsS http://127.0.0.1:8001/v1/models curl -fsS http://127.0.0.1:8002/v1/modelsReport per-profile, per-MIG-UUID, per-service TTFT/P95, token/s, OOM, error rates. Cannot just sum all instances' memory usage and claim "GPU utilization doubled"; task arrival patterns and inter-slice resource reservation change yield.
Cleaning Instances & Restoring Full GPU
First drain traffic from both services and confirm processes exited, then record UUIDs and GI/CI IDs. Deletion order typically CI then GI; deleting an instance may affect dependent Pods. Don't run targetless "wipe all instances" commands; don't change mode while services running.
nvidia-smi mig -i 0 -lci nvidia-smi mig -i 0 -lgi nvidia-smi --query-compute-apps=pid,process_name,used_gpu_memory --format=csv # Generate delete commands per local nvidia-smi mig --help and recorded instance IDs nvidia-smi mig --help | rg -A8 'delete-compute|delete-gpu' # After all CI/GI cleaned and GPU drained, disable mode per maintenance plan sudo nvidia-smi -i 0 -mig 0 nvidia-smi -i 0 -q | sed -n '/MIG Mode/,+5p'Different drivers may adjust command parameters and UUID display; deletion ops must use local --help and instance IDs — don't reuse previous instance IDs as fixed values for next auto-delete script. Post-restore, confirm full GPU service visible again, original GPU capacity restored, monitoring and scheduler resource names synced.
Official References
NVIDIA MIG User Guide
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Ops Community
A leading IT operations community where professionals share and grow together.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
