Tagged articles

INT4

3 articles · Page 1 of 1
MaGe Linux Operations
MaGe Linux Operations
Jul 16, 2026 · Artificial Intelligence

How to Choose Between INT8, FP8, and INT4 Quantization for Large Models

This guide explains how to evaluate INT8, FP8, and INT4 quantization strategies for large language models on NVIDIA GPUs, covering precision trade‑offs, memory consumption, kernel support, KV‑Cache considerations, and detailed deployment, testing, and rollback procedures to ensure performance and quality.

FP8GPU deploymentINT4
0 likes · 48 min read
How to Choose Between INT8, FP8, and INT4 Quantization for Large Models
MaGe Linux Operations
MaGe Linux Operations
Jun 17, 2026 · Artificial Intelligence

Model Quantization: INT8, INT4, and AWQ/GPTQ – Choosing the Right Compression for Production

This article explains how INT8, INT4, bitsandbytes, GPTQ, and AWQ quantization methods can dramatically cut memory usage, boost inference speed, and lower costs for large language models, while detailing their trade‑offs, practical workflows, benchmark results, and common pitfalls to help engineers decide which technique best fits their production scenario.

AWQGPTQINT4
0 likes · 22 min read
Model Quantization: INT8, INT4, and AWQ/GPTQ – Choosing the Right Compression for Production
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Mar 6, 2026 · Artificial Intelligence

How Baidu’s End‑to‑End Quantization Stack Supercharges Large‑Model Inference on Kunlun XPU

Baidu Baige built a full‑stack quantization pipeline that integrates model‑level, framework‑level, and hardware‑level optimizations on the Kunlun XPU platform, enabling FP16/BF16 large models to be compressed to 25‑50% of their original size while boosting inference speed by 30‑50% and dramatically reducing memory consumption for enterprise deployments.

AI InferenceHardware AccelerationINT4
0 likes · 16 min read
How Baidu’s End‑to‑End Quantization Stack Supercharges Large‑Model Inference on Kunlun XPU