Cloud Native 57 min read

GPU Container Troubleshooting: Why nvidia-smi Works But Containers Fail

A systematic guide to debugging GPU container failures across host, runtime, and application layers — covering Docker, containerd, Kubernetes, CDI, permissions, and multi-GPU communication with concrete commands and verification steps.

Ops Community
Ops Community
Ops Community
GPU Container Troubleshooting: Why nvidia-smi Works But Containers Fail

Determine Which Layer the Fault Occurs

The same "no GPU" error can stem from different layers: host driver, container engine, device injection, user-space libraries, or application framework. Record the raw error and timestamp first. Save a timestamped GPU inventory from the host using

nvidia-smi --query-gpu=index,uuid,name,driver_version --format=csv

. UUIDs are stable across reboots; indexes are not. Check container state and logs with docker inspect gpu-app --format '{{json .State}}' and docker logs --timestamps --tail 100 gpu-app. Verify which Docker daemon the client talks to via docker context show and docker context inspect — remote contexts can point to a different host. Confirm server engine version, runtime list, and cgroup driver with docker version and

docker info --format 'Runtimes={{json .Runtimes}} Cgroup={{.CgroupDriver}}/{{.CgroupVersion}}'

. Record OS, kernel, and virtualization type ( cat /etc/os-release; uname -r; systemd-detect-virt || true). Run a minimal CUDA container to isolate runtime injection:

docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

. If this fails, focus on runtime; if it succeeds, investigate the business image, user, environment, and CUDA libraries. Check engine logs near the failure time:

sudo journalctl -u docker --since "15 minutes ago" --no-pager -o short-iso

.

Host-Level Evidence Beyond nvidia-smi

Host nvidia-smi success only proves the management interface works, not compute capability. Check loaded kernel modules and kernel-reported driver version: lsmod | rg "^nvidia"; cat /proc/driver/nvidia/version. Mismatch between disk libraries and kernel module indicates a reboot is needed. Verify PCI devices and bound kernel driver: lspci -nnk -d 10de: — focus on Kernel driver in use. List device nodes and permissions:

ls -l /dev/nvidia*; stat -c "%n %F %a %U:%G" /dev/nvidiactl /dev/nvidia-uvm

. Missing nodes point to module or udev issues, not manual node creation. Query compute mode and memory usage:

nvidia-smi --query-gpu=index,compute_mode,memory.total,memory.used --format=csv

. Prohibited compute mode or full memory mimics runtime failure. Check recent kernel GPU events:

sudo journalctl -k --since "1 hour ago" --no-pager | rg -i "NVRM|Xid|nvidia|AER"

. Record Xid codes, PCI addresses, and timestamps; consult NVIDIA's Xid classification. Inspect ECC, page retirement, and row remapper: nvidia-smi -q -d ECC,PAGE_RETIREMENT,ROW_REMAPER. List GPU processes on host:

nvidia-smi --query-compute-apps=pid,process_name,used_memory --format=csv; ps -eo pid,user,lstart,cmd | rg "python|cuda|vllm"

.

Toolkit Installation and Actual Invocation

NVIDIA Container Toolkit injects device nodes and host driver libraries. Installation, runtime registration, and daemon restart are three separate steps. Verify toolkit binaries:

command -v nvidia-ctk nvidia-container-cli nvidia-container-runtime; nvidia-ctk --version; nvidia-container-cli --version

. Check package installation (Debian: dpkg-query -W "*nvidia-container*" "*libnvidia-container*"; RPM: rpm -qa | rg "nvidia-container|libnvidia-container"). Test toolkit device discovery independently: sudo nvidia-container-cli -k -d /dev/tty info. Inspect Docker daemon.json syntax:

sudo cat /etc/docker/daemon.json; sudo python3 -m json.tool /etc/docker/daemon.json >/dev/null

. Register NVIDIA runtime with backup:

sudo install -d -m 700 /root/gpu-runtime-backup; sudo cp -a /etc/docker/daemon.json /root/gpu-runtime-backup/daemon.json.$(date +%Y%m%d%H%M%S); sudo nvidia-ctk runtime configure --runtime=docker

. Verify systemd drop-ins and actual config path:

systemctl cat docker; systemctl show docker -p ExecStart -p FragmentPath -p DropInPaths

. Restart Docker and confirm runtime registration:

sudo systemctl restart docker; sudo systemctl --no-pager --full status docker; docker info --format '{{json .Runtimes}}'

. Only new containers pick up the new injection path.

Check Device Requests, Not Just Image Names

Docker GPU requests, NVIDIA_VISIBLE_DEVICES, and CUDA device filtering operate at different layers. Inspect saved DeviceRequests and Runtime:

docker inspect gpu-app --format 'Runtime={{.HostConfig.Runtime}} Requests={{json .HostConfig.DeviceRequests}}'

. Missing requests mean the launch definition must be fixed and container recreated. Create a single-GPU diagnostic container by UUID to avoid index drift:

GPU_UUID="GPU-REAL-UUID"; docker run --rm --gpus "device=${GPU_UUID}" nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi -L

. Container-internal device indices restart at zero; map UUID to logical index for framework config. Check NVIDIA/CUDA env vars that may hide devices:

docker inspect gpu-app --format '{{range .Config.Env}}{{println .}}{{end}}' | rg '^(NVIDIA_|CUDA_)'

. Explicitly request compute and utility capabilities for comparison:

docker run --rm --gpus all -e NVIDIA_DRIVER_CAPABILITIES=compute,utility nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

. Verify injected device nodes inside container: docker exec gpu-app sh -c "ls -l /dev/nvidia*; id". Non-root containers need matching supplementary groups. Check mounts that may overlay NVIDIA library or device directories:

docker inspect gpu-app --format '{{json .Mounts}}' | python3 -m json.tool

. For Compose, use docker compose config to see merged GPU declarations.

Management Library Works, Why Compute Library Fails

nvidia-smi

uses NVML; DL frameworks need CUDA driver API, runtime libraries, and compiled kernels. Host driver provides libcuda; image provides runtime and dependencies. Host CUDA Version is the driver's supported upper bound, not the image's toolchain version. In the business container, check dynamic linker cache for management and CUDA driver libs:

docker exec gpu-app sh -c "ldconfig -p 2>/dev/null | grep -E 'libcuda|libnvidia-ml'"

. Watch for stub libraries from CUDA toolkit taking precedence. Examine the application process's library search path:

docker exec gpu-app sh -c 'printf "LD_LIBRARY_PATH=%s
" "$LD_LIBRARY_PATH"'

. Find candidate libraries and distinguish driver-injected from stub:

docker exec gpu-app sh -c "find /usr/local/cuda /usr/lib /lib -name 'libcuda.so*' -o -name 'libnvidia-ml.so*' 2>/dev/null"

. Test CUDA driver initialization directly via Python:

docker exec -i gpu-app python3 - <<'CHECK'
import ctypes
lib=ctypes.CDLL('libcuda.so.1')
lib.cuInit.argtypes=[ctypes.c_uint]
lib.cuInit.restype=ctypes.c_int
print('cuInit:',lib.cuInit(0))
CHECK

. Zero return means success; non-zero must be interpreted per CUDA driver API. Check PyTorch build info and device availability:

docker exec -i gpu-app python3 - <<'CHECK'
import torch
print(torch.__version__,torch.version.cuda)
print(torch.cuda.is_available(),torch.cuda.device_count())
print(torch.__file__)
CHECK

. If version.cuda is None, fix Python deps; if compiled with CUDA but device count zero, continue driver/API/environment checks. Run a small tensor computation with explicit sync:

docker exec -i gpu-app python3 - <<'CHECK'
import torch
x=torch.arange(1024,device='cuda',dtype=torch.float32)
y=(x*x).sum()
torch.cuda.synchronize()
print(y.item(),torch.cuda.get_device_name(0))
CHECK

. Verify supported compute architectures:

docker exec -i gpu-app python3 - <<'CHECK'
import torch
print(torch.cuda.get_arch_list())
if torch.cuda.is_available():
    print(torch.cuda.get_device_capability(0))
CHECK

. Third-party extensions (attention kernels, quantization) may need recompilation.

Permissions, Cgroups, and Security Policies

Traditional file permissions, device cgroups, and MAC (SELinux/AppArmor) jointly gate device access. Inspect container user, supplementary groups, and security options:

docker inspect gpu-app --format 'User={{.Config.User}} Groups={{json .HostConfig.GroupAdd}} Security={{json .HostConfig.SecurityOpt}}'

. If root works but business user fails, fix device group or policy — don't run as root permanently. Get numeric GID of device nodes for group-add:

stat -c "%n uid=%u gid=%g mode=%a" /dev/nvidia0 /dev/nvidiactl /dev/nvidia-uvm

. Test with numeric user and group:

GPU_GID=$(stat -c %g /dev/nvidia0); docker run --rm --gpus all --user 1000:1000 --group-add "$GPU_GID" nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

. Check cgroup version and mount:

docker inspect gpu-app --format '{{.State.Pid}}'; findmnt -t cgroup,cgroup2; cat /proc/$(docker inspect gpu-app --format '{{.State.Pid}}')/cgroup

. On SELinux hosts, collect recent AVC denials: getenforce; sudo ausearch -m AVC,USER_AVC -ts recent -i. On AppArmor, check status and kernel denials:

sudo aa-status; sudo journalctl -k --since "15 minutes ago" | rg -i "apparmor.*denied"

. Verify runc/containerd versions:

runc --version; containerd --version; sudo ls -l /dev/char | head -30

.

Rootless and CDI: Choose the Right Troubleshooting Path

Rootless Docker config lives in user scope; CDI writes device specs explicitly. Identify rootless mode:

docker info --format '{{json .SecurityOptions}}'; systemctl --user status docker --no-pager; printf '%s
' "${XDG_RUNTIME_DIR:-unset}"

. Configure user-level daemon.json:

nvidia-ctk runtime configure --runtime=docker --config="$HOME/.config/docker/daemon.json"; systemctl --user restart docker

. Check Toolkit config for no-cgroups and runmode:

sudo sed -n "1,220p" /etc/nvidia-container-runtime/config.toml

. List CDI device names: nvidia-ctk cdi list. Generate current spec:

sudo install -d -m 755 /var/run/cdi; sudo nvidia-ctk cdi generate --output=/var/run/cdi/nvidia.yaml; nvidia-ctk cdi list

. Test native CDI injection (Podman example):

podman run --rm --device nvidia.com/gpu=all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

. Check auto-refresh service:

systemctl status nvidia-cdi-refresh.path nvidia-cdi-refresh.service --no-pager; journalctl -u nvidia-cdi-refresh.service --since "1 day ago" --no-pager

.

Starts Normal, Loses GPU After Running

Different from creation failure. Known issue: traditional hook loses device access on cgroup updates; systemd daemon-reload can trigger. Record old container timestamps:

docker inspect gpu-app --format 'Created={{.Created}} Started={{.State.StartedAt}} PID={{.State.Pid}}'

. Correlate GPU failure time with deployments, reloads, resource updates. Check Docker events:

docker events --since 1h --until "$(date -Is)" --filter container=gpu-app

. Compare NVML in old vs new container on same host:

docker exec gpu-app nvidia-smi; docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

. Search for daemon-reload near failure:

sudo journalctl --since "2 hours ago" --no-pager | rg -i "Reloading|daemon-reload|docker|containerd"

. Save container resource constraints:

docker inspect gpu-app --format 'Memory={{.HostConfig.Memory}} NanoCPUs={{.HostConfig.NanoCpus}} Cpuset={{.HostConfig.CpusetCpus}}'

. Export container definition and logs before rebuild:

umask 077; mkdir -p gpu-incident; docker inspect gpu-app > gpu-incident/inspect.json; docker logs --timestamps gpu-app > gpu-incident/app.log 2>&1

. Recreate via original deployment (Compose: docker compose up -d --force-recreate gpu-app). Verify NVML, small compute, and business request after recovery.

Kubernetes: Troubleshooting Boundaries Shift

Pod requests extended resources; device plugin reports to kubelet; CRI runtime injects. Docker success ≠ CRI success. Check Pod status and events:

kubectl -n ai get pod gpu-app -o wide; kubectl -n ai describe pod gpu-app

. Pending → check scheduling constraints and node allocatable. Bound but creation fails → check node runtime. Running but framework fails → continue container lib/app tests. Verify node publishes NVIDIA resource:

kubectl get node gpu-node-01 -o json | jq '.status.capacity,.status.allocatable'

. Zero capacity → device plugin registration; reduced → device health/plugin logs. Inspect container resource requests:

kubectl -n ai get pod gpu-app -o json | jq '.spec.containers[] | {name,resources,env}'

. Check initContainers and RuntimeClass. List NVIDIA DaemonSets:

kubectl get daemonset -A | rg -i "nvidia|gpu"; kubectl get pods -A -o wide | rg -i "device-plugin|nvidia"

. Verify RuntimeClass handler registration:

kubectl get runtimeclass -o yaml; kubectl -n ai get pod gpu-app -o jsonpath="{.spec.runtimeClassName}{
}"

. On faulty node, inspect CRI:

sudo crictl info; sudo systemctl cat containerd; sudo containerd config dump | rg -n "nvidia|runtime|imports"

. Exec into business container:

kubectl -n ai exec gpu-app -c app -- nvidia-smi; kubectl -n ai logs gpu-app -c app --since=15m --timestamps

.

Special Devices and Multi-GPU Communication

Single-GPU init success ≠ all devices, MIG instances, or inter-GPU comms work. Check topology: nvidia-smi topo -m. Cross-NUMA affects performance, not disappearance. Inspect MIG mode and instances:

nvidia-smi -q -d MIG; nvidia-smi mig -lgi; nvidia-smi mig -lci

(only on MIG-capable GPUs). Check NVLink status: nvidia-smi nvlink --status. Verify shared memory and IPC config:

docker exec gpu-app df -h /dev/shm; docker inspect gpu-app --format 'ShmSize={{.HostConfig.ShmSize}} Ipc={{.HostConfig.IpcMode}}'

. Test per-GPU context creation:

docker exec -i gpu-app python3 - <<'CHECK'
import torch
for i in range(torch.cuda.device_count()):
    with torch.cuda.device(i):
        t=torch.ones(256,device=f'cuda:{i}')
        print(i,t.sum().item())
        torch.cuda.synchronize(i)
CHECK

. Map failing logical GPU back to host UUID. Check NCCL/UCX/FI env vars:

docker exec gpu-app sh -c 'env | grep -E "^(NCCL_|UCX_|FI_)" | sort'

. Verify RDMA devices for multi-node:

docker exec gpu-app sh -c "ls -l /dev/infiniband 2>/dev/null; ip -br link"

.

Make Recovery a Repeatable Acceptance

Acceptance must cover host, container, and business layers. Export sanitized baseline:

umask 077; { date -Is; uname -r; nvidia-smi -L; docker version; nvidia-ctk --version; } > gpu-baseline.txt

. Automated NVML check with exit code:

if docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi > nvml-check.log 2>&1; then echo "NVML check passed"; else cat nvml-check.log; exit 1; fi

. Compute smoke test in business image:

docker exec -i gpu-app python3 - <<'CHECK'
import torch
a=torch.eye(32,device='cuda')
b=a@a
torch.cuda.synchronize()
assert torch.allclose(a,b)
print('CUDA smoke passed')
CHECK

. Monitor utilization during test:

nvidia-smi --query-gpu=timestamp,uuid,utilization.gpu,memory.used --format=csv -l 1

. End-to-end health request:

curl --fail --max-time 30 -H "Content-Type: application/json" -d '{"input":"gpu-smoke","max_tokens":8}' http://127.0.0.1:8000/test

. Diff runtime config changes:

sudo diff -u /root/gpu-runtime-backup/daemon.json.BACKUP_TIME /etc/docker/daemon.json

. Write incident summary with symptom, scope, evidence, cause, change, verify, rollback.

Deep Dive: Collect Runtime Evidence

Enable debug logging around a controlled repro. Check runtime log config:

sudo rg -n "debug|log-level|mode" /etc/nvidia-container-runtime/config.toml

. Preview debug config:

sudo nvidia-ctk config --set nvidia-container-runtime.log-level=debug --set nvidia-container-runtime.debug=/var/log/nvidia-container-runtime.log

. Apply with --in-place after preview; revert after repro. Inspect entrypoint and cmd:

docker inspect gpu-app --format 'Entry={{json .Config.Entrypoint}} Cmd={{json .Config.Cmd}}'

. Check NVIDIA_REQUIRE_* constraints:

docker inspect gpu-app --format '{{range .Config.Env}}{{println .}}{{end}}' | rg '^NVIDIA_REQUIRE_'

. Save image digest and arch:

docker image inspect nvidia/cuda:12.4.1-base-ubuntu22.04 --format 'Arch={{.Architecture}} Digests={{json .RepoDigests}}'

. Check disk space and inodes:

df -h / /var/lib/docker /var/log; df -i / /var/lib/docker /var/log

. Generate checksums of collected artifacts:

find gpu-incident -maxdepth 1 -type f -print0 | sort -z | xargs -0 sha256sum > gpu-incident-checksums.txt

.

Two Illustrative Cases

Case 1: Admin runs nvidia-smi on node A, sees four GPUs, but Docker test fails with driver selection error. Reinstalling Toolkit doesn't help. docker context show reveals client points to node B (CPU-only). Fix: use correct context or deploy GPU on target node. Case 2: Business container runs hours then reports NVML Unknown Error. Host query OK, new container on same UUID works, old container fails. Timeline shows automation updated memory limit; node uses traditional hook. Hypothesis: cgroup update revoked device access. Rebuild container to restore; long-term: test CDI migration or validated component upgrade on same version test node. Both cases show nvidia-smi success is only local evidence. Reliable troubleshooting links: where executed, who created, what requested, what injected, what loaded, when failed.

References

NVIDIA Container Toolkit Troubleshooting: Cgroups, NVML, Permissions

NVIDIA Container Toolkit Installation and Runtime Configuration

NVIDIA CDI Support and Device Spec Maintenance

Docker Compose GPU Device Reservation

Kubernetes GPU Scheduling

NVIDIA CUDA Compatibility

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

DockerKubernetestroubleshootingGPUcontainer runtimeCDINVMLNVIDIA Container Toolkit
Ops Community
Written by

Ops Community

A leading IT operations community where professionals share and grow together.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.