Tagged articles

Qwen3.8-27B

8 articles · Page 1 of 1
Old Zhang's AI Learning
Old Zhang's AI Learning
Sep 25, 2026 · Artificial Intelligence

Ternary Bonsai 2: Qwen3.8-27B Compressed to 5.9GB at 98.2% Performance

PrismML's Ternary Bonsai 2 27B uses rotated weight basis and FP16 group-wise scaling to ternary-quantize Qwen3.8-27B to 1.76 bits (5.9GB), achieving 98.2% benchmark retention across coding, reasoning, and agent tasks, with custom CUDA/MLX kernels enabling fast inference on consumer GPUs and Apple Silicon.

Bonsai 2GGUFLLM deployment
0 likes · 15 min read
Ternary Bonsai 2: Qwen3.8-27B Compressed to 5.9GB at 98.2% Performance
Old Zhang's AI Learning
Old Zhang's AI Learning
Aug 23, 2026 · Artificial Intelligence

Qwen3.8-27B Quantization Selection Guide: Match Your Hardware to the Right GGUF Version

This guide analyzes Unsloth's Dynamic v3.0 quantizations of Qwen3.8-27B, showing Mean KLD divergence across versions, recommending UD-Q4_K_XL for 24GB GPUs, detailing hardware requirements, sampling parameters for thinking modes, and explaining why 4-bit is the baseline for tool use while 8-bit offers diminishing returns.

Dynamic v3.0GGUFMean KLD
0 likes · 10 min read
Qwen3.8-27B Quantization Selection Guide: Match Your Hardware to the Right GGUF Version
AI Programming Lab
AI Programming Lab
Aug 20, 2026 · Artificial Intelligence

Testing Qwen3.8-27B on Dual 4090 GPUs and Connecting to Claude Code for Unlimited Tokens

The article details a hands‑on deployment of Alibaba's Qwen3.8-27B model on a server with two 48 GB turbo‑variant RTX 4090 GPUs using vLLM, discusses hardware and software constraints, configuration tweaks like FP8 KV cache, and integration with Claude Code via a custom Anthropic‑compatible router, while sharing performance observations and community benchmark scores.

Anthropic APIClaude CodeFP8 KV cache
0 likes · 9 min read
Testing Qwen3.8-27B on Dual 4090 GPUs and Connecting to Claude Code for Unlimited Tokens
AI Engineering
AI Engineering
Aug 20, 2026 · Artificial Intelligence

Qwen3.8-27B 1‑bit Quantization Fits in 8 GB RAM with 77% Accuracy

Unsloth’s new Dynamic V3 quantization for Qwen3.8‑27B compresses the 27‑billion‑parameter model to as little as 6.2 GB using 1‑bit, preserving about 77 % of the original Top‑1 accuracy and allowing inference on devices with 8 GB of combined RAM and VRAM, while higher‑bit versions require proportionally more memory.

1-bitDynamic V3GGUF
0 likes · 5 min read
Qwen3.8-27B 1‑bit Quantization Fits in 8 GB RAM with 77% Accuracy
IT Xianyu
IT Xianyu
Aug 15, 2026 · Artificial Intelligence

Running Alibaba’s Open‑Source Qwen3.8‑27B on a Consumer GPU: Unexpected Performance

The author tests Alibaba’s newly released open‑source Qwen3.8‑27B model, quantizes it to 4‑bit GGUF to fit a 14 GB VRAM slot on a 20 GB consumer GPU, and finds it matches or exceeds larger closed‑source models like Opus 4.6 Max and Claude on coding and helper tasks, all under an Apache 2.0 license.

Apache-2.0LLM quantizationQwen3.8-27B
0 likes · 7 min read
Running Alibaba’s Open‑Source Qwen3.8‑27B on a Consumer GPU: Unexpected Performance
Old Zhang's AI Learning
Old Zhang's AI Learning
Aug 14, 2026 · Artificial Intelligence

Why Qwen3.8-27B Is the World’s New Favorite Open‑Source LLM and How to Deploy It Locally

The article introduces Qwen3.8-27B, a dense multimodal LLM with up to 256K tokens (extendable to 1M), highlights its benchmark gains over previous Qwen models, discusses model size, quantization options, and provides step‑by‑step instructions for local deployment using vLLM, Docker, and LMStudio.

BenchmarkMultimodalQwen3.8-27B
0 likes · 8 min read
Why Qwen3.8-27B Is the World’s New Favorite Open‑Source LLM and How to Deploy It Locally