Tagged articles

PrismML

1 articles · Page 1 of 1
Old Zhang's AI Learning
Old Zhang's AI Learning
Sep 25, 2026 · Artificial Intelligence

Ternary Bonsai 2: Qwen3.8-27B Compressed to 5.9GB at 98.2% Performance

PrismML's Ternary Bonsai 2 27B uses rotated weight basis and FP16 group-wise scaling to ternary-quantize Qwen3.8-27B to 1.76 bits (5.9GB), achieving 98.2% benchmark retention across coding, reasoning, and agent tasks, with custom CUDA/MLX kernels enabling fast inference on consumer GPUs and Apple Silicon.

Bonsai 2GGUFLLM deployment
0 likes · 15 min read
Ternary Bonsai 2: Qwen3.8-27B Compressed to 5.9GB at 98.2% Performance