Old Zhang's AI Learning
Sep 25, 2026 · Artificial Intelligence
Ternary Bonsai 2: Qwen3.8-27B Compressed to 5.9GB at 98.2% Performance
PrismML's Ternary Bonsai 2 27B uses rotated weight basis and FP16 group-wise scaling to ternary-quantize Qwen3.8-27B to 1.76 bits (5.9GB), achieving 98.2% benchmark retention across coding, reasoning, and agent tasks, with custom CUDA/MLX kernels enabling fast inference on consumer GPUs and Apple Silicon.
Bonsai 2GGUFLLM deployment
0 likes · 15 min read
