ITPUB
Jul 25, 2026 · Artificial Intelligence
Running a 35B MoE Model at 32 tokens/s on $600 Tesla P100 GPUs
This article details how to repurpose two second‑hand Tesla P100 GPUs (≈$600 each) in a Dell R730XD server with ESXi 8.0, AlmaLinux 10, NVIDIA drivers, Docker and Ollama to run the qwen3.6:35b model at 32.77 tokens/s, including step‑by‑step configuration, performance benchmarks, and practical tips.
DockerESXiGPU passthrough
0 likes · 14 min read
