Alibaba Unveils Qwen3‑Max‑Thinking: Trillion‑Parameter Model Joins Global AI Elite

Alibaba's Tongyi team released the Qwen3‑Max‑Thinking model, surpassing one trillion parameters and 36 T tokens of pre‑training, introducing adaptive tool‑calling and test‑time expansion techniques that boost benchmark scores across 19 tests, positioning it alongside top international models and now available via app, web, and API.

Smart Sea Tide
Smart Sea Tide
Smart Sea Tide
Alibaba Unveils Qwen3‑Max‑Thinking: Trillion‑Parameter Model Joins Global AI Elite

Alibaba's Tongyi team announced the official release of its flagship language model Qwen3‑Max‑Thinking, whose total parameter count exceeds one trillion and is trained on 36 T tokens of high‑quality data. The model comes in Base, Instruct, and Thinking variants and targets comprehensive upgrades in inference, knowledge, tool usage, and agent capabilities, directly competing with GPT‑5.2‑Thinking, Claude‑Opus‑4.5, and Gemini‑3 Pro.

The core breakthroughs are twofold. First, an adaptive tool‑calling ability lets the model autonomously invoke built‑in search, memory, and code interpreter modules during dialogue, unlike traditional models that require manual tool selection. This is achieved through a dedicated training pipeline that fine‑tunes the model with tool‑use tasks and employs dual feedback from rules and the model itself, reducing hallucinations, enabling real‑time information retrieval, and allowing complex computational reasoning via code execution.

Second, a test‑time expansion technique replaces naïve parallel‑path scaling. By limiting the number of parallel trajectories and reallocating compute to an experience‑accumulating multi‑round self‑reflection process, the model extracts key insights from prior reasoning steps, avoids redundant derivations, and improves context utilization under the same token budget, yielding notable gains on benchmarks such as GPQA and LiveCodeBench.

In extensive benchmark evaluations covering 19 tests of scientific knowledge, complex reasoning, and programming, Qwen3‑Max‑Thinking matches or surpasses leading closed‑source models. It achieved perfect scores (25/25) on AIME and HMMT, 83.9 on IMO‑AnswerBench, 85.9 on LiveCodeBench v6, 75.3 on SWE‑bench Verified, 93.7 on C‑Eval, 87.4 on GPQA, 82.1 on Tau2 Bench, and 90.2 on Arena‑Hard v2, demonstrating strong performance across mathematics, coding, and tool‑use dimensions.

Following the launch, developers worldwide praised the Qwen series for its rapid iteration and transparent communication, noting that its update cadence now exceeds that of OpenAI and other international providers. Users expressed anticipation for the new version and called for Android‑specific UI improvements to fully leverage the model’s capabilities.

Qwen3‑Max‑Thinking is now accessible for free through the Qwen App, PC client, and web interface, with the API endpoint qwen3-max-2026-01-23 open to developers. The release marks a milestone for Chinese AI, showcasing a domestically built trillion‑parameter model that competes at the highest global tier.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

large language modelbenchmark performanceAI model comparisonQwen3-Max-Thinkingtrillion-parameter modeladaptive tool calling
Smart Sea Tide
Written by

Smart Sea Tide

Sharing cutting‑edge big data and AI technologies, with occasional lifestyle insights.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.