Tagged articles

multi-GPU inference

2 articles · Page 1 of 1
Ops Community
Ops Community
Oct 7, 2026 · Artificial Intelligence

Multi-GPU LLM Inference: Step-by-Step Deployment Guide for Memory-Constrained Environments

This comprehensive tutorial covers multi-GPU inference for large language models, detailing VRAM estimation formulas, parallel strategies (tensor, pipeline, data), GPU interconnects (NVLink, PCIe), framework comparisons (vLLM, DeepSpeed, Transformers), quantization techniques (GPTQ, AWQ, bitsandbytes), performance benchmarking, and production deployment with Docker and Kubernetes.

DeepSpeedGPU memory optimizationLLM deployment
0 likes · 32 min read
Multi-GPU LLM Inference: Step-by-Step Deployment Guide for Memory-Constrained Environments
AIWalker
AIWalker
Jan 17, 2025 · Artificial Intelligence

How CLEAR Cuts Attention Compute by 99.5% and Enables Efficient On‑Device Text‑to‑Image Diffusion

The CLEAR method linearizes pretrained Diffusion Transformers by restricting attention to a local window, reducing attention FLOPs by 99.5%, accelerating 8K image generation 6.3× while preserving quality, and supporting multi‑GPU patch‑wise inference for high‑resolution text‑to‑image synthesis.

Diffusion TransformersLinear Attentionclear
0 likes · 21 min read
How CLEAR Cuts Attention Compute by 99.5% and Enables Efficient On‑Device Text‑to‑Image Diffusion