Alibaba Open-Sources Three Qwen3.5 Models that Set New Mid-Size Performance Records and Run on Consumer GPUs

Alibaba has released three new Qwen3.5 models (35B‑A3B, 122B‑A10B, 27B) that, thanks to hybrid attention and high‑sparse MoE architecture, achieve record performance for mid‑size LLMs, support up to 1 M token context, and can be deployed on consumer‑grade GPUs with low inference cost.

Smart Sea Tide
Smart Sea Tide
Smart Sea Tide
Alibaba Open-Sources Three Qwen3.5 Models that Set New Mid-Size Performance Records and Run on Consumer GPUs

Alibaba has open‑sourced three new Qwen3.5 models—Qwen3.5‑35B‑A3B, Qwen3.5‑122B‑A10B, and Qwen3.5‑27B—positioned as mid‑size large language models that set new performance benchmarks while remaining runnable on consumer‑grade graphics cards.

The performance leap stems from a hybrid attention mechanism combined with a high‑sparse Mixture‑of‑Experts (MoE) architecture and training on a larger corpus of mixed text‑and‑vision tokens. These innovations allow the models to achieve higher accuracy with fewer total and activation parameters.

On multiple authoritative leaderboards covering instruction following, doctoral‑level reasoning, mathematical reasoning, and multilingual knowledge, Qwen3.5‑35B‑A3B and Qwen3.5‑122B‑A10B outperform the much larger predecessor Qwen3‑235B‑A22B and the flagship Qwen3‑VL, and they also surpass comparable models such as GPT‑5 mini, gpt‑oss‑120b, and Claude Sonnet 4.5.

Qwen3.5‑27B, the first dense model in the Qwen 3.5 family, demonstrates especially strong agent capabilities and native multimodal support. In tool‑calling and programming agent benchmarks it exceeds GPT‑5 mini, while in visual reasoning, document recognition, and video reasoning it outperforms both the Qwen3‑VL flagship and Claude Sonnet 4.5, all while being runnable on a single GPU.

All three models retain near‑lossless precision under 4‑bit quantization. Qwen3.5‑27B supports context lengths over 800 K tokens, and Qwen3.5‑35B‑A3B can handle more than 1 M tokens on a consumer GPU with 32 GB of VRAM, enabling efficient processing of long documents and complex tasks.

Deployment barriers are lowered further: each model can be run on consumer‑grade GPUs, dramatically reducing hardware costs. Alibaba also launched a managed service, Qwen3.5‑Flash, on Alibaba Cloud Bailei, which inherits the 1 M‑token context, built‑in tool‑calling, and offers an input cost of only 0.2 CNY per million tokens. The previously released Qwen3.5‑Plus matches Gemini 3’s performance at just 5 % of Gemini’s API price, creating a tiered service ecosystem.

The release has sparked strong community response. Alibaba now has over 400 open‑source Qwen models, more than 1 billion total downloads, and over 200 k derivative models. The earlier Qwen3.5‑397B‑A17B topped Hugging Face’s global ranking, keeping Qwen as the leading open‑source model and advancing the democratization and industrialization of large‑scale AI.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Agentopen sourcelarge language modelbenchmarkmultimodalQwen3.5consumer GPU
Smart Sea Tide
Written by

Smart Sea Tide

Sharing cutting‑edge big data and AI technologies, with occasional lifestyle insights.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.