Qwen3.8-Max Official Release: Open‑Source Model Beats Expectations in Benchmarks

The newly released 2.4 T‑parameter Qwen3.8‑Max model adds native multimodal support, a 1 M‑token context window and an open‑source API, and it outperforms or matches top closed‑source LLMs on SWE‑Bench, Terminal‑Bench and FrontierQA while handling diverse real‑world tasks.

Old Zhang's AI Learning
Old Zhang's AI Learning
Old Zhang's AI Learning
Qwen3.8-Max Official Release: Open‑Source Model Beats Expectations in Benchmarks

Alibaba announced the official release of Qwen3.8‑Max, a 2.4 T‑parameter large language model with native multimodal capabilities, a 1 M‑token context window, and an API that went live on the Qwen AI platform before being open‑sourced a week later.

Benchmark results show Qwen3.8‑Max leading or matching top closed‑source models such as Opus and various GPT versions on SWE‑Bench, Terminal‑Bench, and FrontierQA, indicating strong performance across programming, office, research and long‑running tasks.

Test 1

The author evaluated the model with a classic reading‑comprehension passage ("背影"), SVG code generation and an aesthetic test. Qwen3.8‑Max solved the tasks more reliably than GPT‑Sol‑Max, which struggled with the same prompt.

Test 2

Using a securities‑research article that listed 50 recommended books, the model extracted the titles and recommendation notes, fetched detailed information and covers from Doubao, and generated a web page that also included a staged one‑year reading plan (basic, macro, master investors, trading discipline, and personal development).

Test 3

A complex workflow—article → script → Remotion video → dubbing → cover → BGM → final video—was executed. The author notes that Qwen3.7‑Max handled the pipeline without issues, and Qwen3.8‑Max completed it successfully in a single run.

Test 4

Feeding a markdown article to Qwen3.8‑Max produced a hand‑written‑style PPT in the author’s preferred “Wang Hong” visual style, demonstrating the model’s ability to generate custom presentation layouts.

Test 5

Within the Qoder IDE, the Better Harness feature used Qwen3.8‑Max to automatically analyze an in‑progress project and suggest numerous fixes, revealing many hidden problems that the author plans to address later.

Test 6

When asked to explain the principle of a lunar eclipse to a child, the model decomposed the concept and generated a clear, illustrated explanation using only a simple prompt.

Test 7

The model created a full 3‑D disassembly of a device, modeling 13 components with basic geometry and procedural materials, and added interactive controls for exploded view, part‑by‑part animation, hover highlighting, click‑to‑view parameters, and label tracking.

Qwen3.8‑Max can be invoked via the Qwen AI platform and integrates seamlessly with major agent frameworks and coding assistants. Its API follows the Anthropic protocol (compatible with Claude Code) and the OpenAI Responses protocol (compatible with Codex), requiring only a config change. The model is already available in Qoder, the Qwen Office client, and the Qwen App.

Overall, the author finds the real‑world tests exceed expectations and hopes that smaller variants (e.g., the announced 27B model) will also be open‑sourced for local deployment.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

multimodal AIlarge language modelopen source AIAI assistantbenchmarkingQwen3.8-Max
Old Zhang's AI Learning
Written by

Old Zhang's AI Learning

AI practitioner specializing in large-model evaluation and on-premise deployment, agents, AI programming, Vibe Coding, general AI, and broader tech trends, with daily original technical articles.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.