Operations 50 min read

NVIDIA Container Toolkit Practical Guide: Enable GPU Access for Docker Containers

This comprehensive guide covers installing and configuring NVIDIA Container Toolkit for Docker GPU access, including version pinning, runtime registration, Compose deployment, image building, non-root execution, CDI, offline installation, capacity planning, and upgrade/rollback procedures with concrete commands and verification steps.

MaGe Linux Operations
MaGe Linux Operations
MaGe Linux Operations
NVIDIA Container Toolkit Practical Guide: Enable GPU Access for Docker Containers

Understand the Four Components and Establish Version Boundaries

The host driver provides kernel modules and user-space interfaces; the CUDA image supplies the compute runtime; Docker manages container lifecycles; the Toolkit translates requests into device and library container configurations. All four upgrade independently, so a compatibility matrix must be recorded before installation.

Record OS, architecture, kernel, and Docker context to ensure operations target the correct host. Verify driver recognition of all expected GPUs with

nvidia-smi --query-gpu=index,uuid,name,driver_version --format=csv

. Note that the CUDA Version shown is the driver's capability, not the toolkit version inside the image.

Check existing Docker configuration (

docker info --format 'Root={{.DockerRootDir}} Driver={{.Driver}} Runtimes={{json .Runtimes}}'

) to avoid overwriting custom log drivers, storage backends, or root directories. Create a compatibility matrix file (gpu-compatibility.txt) capturing OS, arch, GPU model, driver, Docker, Toolkit, CUDA image tag and digest, framework, and verification checks.

Configure Trusted Software Sources and Pin Package Versions

For Debian/Ubuntu: install ca-certificates curl gnupg, download the NVIDIA GPG key, verify its fingerprint, dearmor into /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg, then write the stable repository list with signed-by pointing to that keyring. Use sed to uncomment the deb line. For RPM systems: download nvidia-container-toolkit.repo to /etc/yum.repos.d/ and list duplicates with dnf list --showduplicates nvidia-container-toolkit.

Refresh indexes and list candidate versions ( apt-cache madison nvidia-container-toolkit / apt-cache policy libnvidia-container1). Pin the exact approved version for all four packages ( nvidia-container-toolkit, nvidia-container-toolkit-base, libnvidia-container-tools, libnvidia-container1) using a version variable that fails if empty.

Register the NVIDIA Runtime with Docker

Backup /etc/docker/daemon.json (or note its absence). Run sudo nvidia-ctk runtime configure --runtime=docker to merge the NVIDIA runtime configuration. Validate syntax with python3 -m json.tool /etc/docker/daemon.json and, if supported, dockerd --validate --config-file=/etc/docker/daemon.json. Check systemd drop-ins for conflicting options ( systemctl cat docker). Restart Docker and verify the runtime appears in docker info --format '{{json .Runtimes}}'.

From Minimal Test to Specific GPU Selection

Start with a control container without GPU (

docker run --rm nvidia/cuda:12.4.1-base-ubuntu22.04 sh -c "uname -m; cat /etc/os-release"

). Then request all GPUs (

docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi

). Inspect the created container's DeviceRequests to confirm the request was accepted (

docker inspect gpu-toolkit-test --format '{{json .HostConfig.DeviceRequests}}'

).

Select a specific GPU by UUID (

docker run --rm --gpus "device=${GPU_UUID}" nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi -L

). Demonstrate multi-GPU selection with --gpus '"device=0,1"'. Limit driver capabilities to reduce injected libraries ( -e NVIDIA_DRIVER_CAPABILITIES=compute,utility). Clean up test containers explicitly.

Express GPU Requirements in Docker Compose

Use deploy.resources.reservations.devices with driver: nvidia, count (or device_ids), and capabilities: [gpu]. count and device_ids are mutually exclusive. Render the final config with

docker compose -f compose.yaml config > compose.rendered.yaml

to audit overrides and variable interpolation. Pull images before deployment ( docker compose pull) and record digests. Run a one-off check service ( docker compose run --rm gpu-check) to validate device reservation.

Add runtime environment variables, restart policy, shared memory, and stop grace period to the production service definition.

Design Application Images: Don't Install Drivers Inside Containers

Application images should contain business dependencies, CUDA runtime, and tools; host driver libraries are injected by the Toolkit. Use runtime base images (e.g., nvidia/cuda:12.4.1-runtime-ubuntu22.04) for production. Inspect architecture and digests (

docker image inspect ... --format 'Arch={{.Architecture}} Digests={{json .RepoDigests}}'

).

For compiling CUDA extensions, use a multi-stage build: devel stage to produce wheels, then copy into a runtime stage. Verify framework CUDA compatibility at runtime (

python3 -c "import torch; print(torch.__version__, torch.version.cuda, torch.cuda.is_available())"

) and run a small compute test with synchronization to expose async errors. Ensure LD_LIBRARY_PATH does not prioritize stub libraries. Exclude secrets, model weights, and caches via .dockerignore.

Run as Non-Root User and Control Access Scope

Create a fixed UID/GID user in the Dockerfile (

groupadd -g 10001 app && useradd -u 10001 -g app -m app

), chown app directories, and set USER 10001:10001. On the host, check device node permissions and numeric group (

stat -c "%n %a uid=%u gid=%g" /dev/nvidia0 /dev/nvidiactl /dev/nvidia-uvm

). Run a test container with --user 10001:10001 --group-add ${DEVICE_GID} to verify access. Add

--read-only --tmpfs /tmp:rw,nosuid,size=256m --cap-drop ALL --security-opt no-new-privileges

for hardening. On SELinux hosts, audit denials with ausearch -m AVC,USER_AVC -ts recent -i. Verify the deployed container's actual user, groups, and security options via docker inspect.

CDI and Rootless Are Two Distinct Paths

CDI (Container Device Interface) encodes device nodes, library mounts, and edits into a spec for compatible runtimes. List available CDI devices with nvidia-ctk cdi list. Generate a spec (

sudo nvidia-ctk cdi generate --output=/var/run/cdi/nvidia.yaml

) and monitor the refresh service (

systemctl status nvidia-cdi-refresh.path nvidia-cdi-refresh.service

). On Docker with native CDI support, request devices via --device nvidia.com/gpu=all. Do not mix traditional --gpus and CDI unconsciously.

For rootless Docker, confirm the engine runs in user mode ( docker info --format '{{json .SecurityOptions}}', systemctl --user status docker). Register the runtime in the user config (

nvidia-ctk runtime configure --runtime=docker --config=$HOME/.config/docker/daemon.json

) and restart the user service. The no-cgroups setting (

sudo nvidia-ctk config --set nvidia-container-cli.no-cgroups=true

) may be required; document why and how to revert.

Offline Deployment: Deliver Dependencies with Provenance

On a matching-arch preparation machine, download exact package versions ( apt-get download ... or dnf download --resolve --alldeps --destdir ...). Verify architecture and signatures ( dpkg-deb -f / rpm --checksig). Generate SHA256 manifests (

find ... -name "*.deb" -print0 | sort -z | xargs -0 sha256sum > toolkit.sha256

). Test installation on a cloned baseline ( apt-get install ./offline-toolkit/*.deb). Save approved images with docker save and verify digests on load.

Distinguish Device Availability from Capacity Sufficiency

Successful device injection does not guarantee enough VRAM. Query total, used, free memory and compute-app usage (

nvidia-smi --query-gpu=uuid,memory.total,memory.used,memory.free --format=csv

). Build a memory budget (weights, workspace, cache, reserve) and validate with real workloads. Configure --shm-size for multi-process data loading. Mount model and cache volumes read-only/read-write explicitly, ensuring host directories exist with correct permissions. Implement a health check script that performs a trivial CUDA operation and exits non-zero on failure; integrate into Compose with appropriate start_period.

Upgrade, Rollback, and Controlled Troubleshooting

Before upgrades, snapshot running containers and engine baseline ( docker ps --format ..., docker info, nvidia-ctk --version). Tail Docker logs and run nvidia-container-cli -k -d /dev/tty info to separate host device discovery from engine integration. Enable debug logging temporarily (

nvidia-ctk config --set nvidia-container-runtime.log-level=debug --set nvidia-container-runtime.debug=/var/log/nvidia-container-runtime.log

). Compare pre/post configs for runtime path changes. Rollback by restoring daemon.json and restarting Docker. Use apt-mark showhold to track pinned packages. Recreate Compose services with docker compose up -d --force-recreate to apply new device requests.

Acceptance Delivery: Make the Next Deployment Equally Successful

Deliver a deployment package containing compose.yaml, Dockerfile, requirements.lock, check_gpu.py, compatibility.txt, image-digest.txt, rollback.md, acceptance.md. Run automated acceptance tests: NVML query on target GPU, non-root device visibility, application health endpoint, and record image digests and device requests. Document maintenance gaps (reboot, CDI refresh, driver upgrade) explicitly. Generate a SHA256 manifest of the delivery directory for integrity verification.

Illustrative Case Studies (Fictional)

Case 1: 80 GiB GPU, minimal nvidia-smi passes but model load fails. Root cause: CPU-only framework build. After fixing dependencies, model loads but OOM under concurrency. Budget: weights 48 GiB, workspace 6 GiB, cache 16 GiB, reserve 5 GiB = 75 GiB used, 5 GiB headroom. Variable input lengths push cache higher. Fix: correct framework build and adjust capacity config.

Case 2: Offline ARM server. Packages collected from x86 host, only filenames checked. Installation fails; copying binaries bypassing package manager yields non-executable tools. Correct path: same-arch preparation environment, full dependency resolution, signature and hash manifests, end-to-end test on cloned system.

Reference Materials

NVIDIA Container Toolkit Installation Guide

NVIDIA Docker Configuration and Capability Variables

NVIDIA CDI Support

NVIDIA Runtime Troubleshooting

Docker Compose GPU Support

Docker Rootless Mode

CUDA Compatibility Documentation

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

dockercudacapacity-planningtroubleshootinggpucomposecontainer-runtimeoffline-deploymentnvidia-container-toolkit
MaGe Linux Operations
Written by

MaGe Linux Operations

Founded in 2009, MaGe Education is a top Chinese high‑end IT training brand. Its graduates earn 12K+ RMB salaries, and the school has trained tens of thousands of students. It offers high‑pay courses in Linux cloud operations, Python full‑stack, automation, data analysis, AI, and Go high‑concurrency architecture. Thanks to quality courses and a solid reputation, it has talent partnerships with numerous internet firms.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.