Qwen3.8-Max: Open‑Source Max‑Level LLM That Automates Programming, Office Work, and Research

Qwen3.8-Max, a 2.4 trillion‑parameter open‑weight LLM, ranks fourth on Frontend Code Arena, autonomously completes multi‑day coding projects, reproduces and surpasses a research paper, dominates a multimodal dialogue contest, and tackles real‑world office, chip‑design, and e‑commerce tasks, showcasing unprecedented self‑directed capability.

SuanNi
SuanNi
SuanNi
Qwen3.8-Max: Open‑Source Max‑Level LLM That Automates Programming, Office Work, and Research

Introduction

Qwen3.8-Max, a 2.4‑trillion‑parameter model, is the strongest in the Qwen family and the first open‑weight Max‑level large language model. It ranks fourth on the Frontend Code Arena leaderboard with a score of 1668, just behind Claude Opus 5 (Max) and Kimi K3 (Max).

Autonomous Capability Overview

The model can work continuously for days with minimal human intervention, handling programming, office, research, and long‑cycle tasks.

Challenge 1: Ten‑Day Autonomous Coding

Qwen3.8-Max created the “oh‑my‑cli” project from scratch, building a self‑evolving harness that turns user feedback and test results into issues, which agents automatically claim, implement, test, and iterate. After 16 days it produced 265 commits, 127 pull requests, and 151 issues.

Challenge 2: Reproducing and Surpassing a Research Paper

Tasked with reproducing “Unified Data Selection for LLM Reasoning”, the model wrote all data processing, training, and evaluation code from zero. Over ~5 days (125 h) it wrote ~7 600 lines, executed >1 100 steps, and ran 33 GPU‑training rounds, fully replicating six major conclusions and improving the AIME24 benchmark by 2.7 points.

Challenge 3: 24‑Hour Multimodal Dialogue Intent Recognition Contest

In the WWW2025 competition with 526 human teams, Qwen3.8-Max independently read the rules, assembled a solution using fine‑tuned Chinese BERT, MacBERT, RoBERTa models for text and a Qwen2.5‑VL‑7B visual model with Chinese‑CLIP fallback. After 45 submissions its accuracy rose from 0.60 to 0.853, beating 458 human teams and capturing 87 % of the leaderboard.

Real‑World Office and Productivity Tasks

Beyond coding, the model excels at multi‑step, tool‑heavy office tasks. Scaling RL environments and compute, Qwen3.8-Max shows steady gains across dozens of in‑house and public benchmarks. On various harnesses (QwenWork, Claude Code, Codex, OpenClaw, Hermes) its general work ability is comparable.

In quantitative‑strategy development, starting from a brief description it built data pipelines, factor libraries, and iterative back‑testing loops, identifying over‑fit signals, pruning them, and converging on robust strategies with excess Sharpe ratios between 0.64 and 1.48, compressing weeks of analyst work into a single dialogue.

Long‑Term Complex Tasks

Chip‑design case: given only a task description, an empty RTL skeleton, and evaluation scripts, the model autonomously performed ~500 interaction rounds, 71 evaluations, crossing 13 milestones, reducing chip area by 81 %.

E‑Commerce Bench: using de‑identified Taobao/Tmall data, the model managed multiple stores with ¥10 k capital, handling product selection, negotiation, inventory, dynamic pricing, and returns across 600 suppliers, detecting fraud in 152 cases. It achieved a final cash balance of ¥416 252 (4.16× ROI), a 152 % improvement over Qwen3.7‑Max.

Multimodal End‑to‑End Ability

Qwen3.8‑Max processes >200‑page reports and >100‑hour videos, extracting structured summaries, building video graphs, and generating deliverable web pages. It also creates visual outputs such as Vlogs, teaching animations, front‑end projects from screenshots, and 3D interior renders.

The model continuously monitors its intermediate outputs, correcting layout or orientation errors, forming a native visual feedback loop during execution.

Benchmarks and Ecosystem Extensions

The team released RecreationBench, a hybrid‑agent benchmark covering Ubuntu, macOS, Windows, Android, and Web platforms, and Qwen‑MM‑Plugins, a library that adds image/video processing, multimodal memory, dynamic resolution, and professional tool support (Blender, CAD) to existing agent harnesses.

A smaller 27‑billion‑parameter version of Qwen3.8 will also be open‑sourced, promising further capability gains for developers.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

open-sourcelarge language modelBenchmarkmultimodalAI researchautonomous AIQwen3.8-Max
SuanNi
Written by

SuanNi

A community for AI developers that aggregates large-model development services, models, and compute power.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.