5 AI Models Launch in 48 Hours: Tiered Pricing, Cost Cuts, Open Agent Training
Within 48 hours, xAI, Xiaomi, Anthropic, and OpenAI released Grok 4.7, MiMo V2.6, Claude Opus 5.5, GPT-6 Sol, and Luna, revealing a shift toward tiered model pricing, significant cost reductions, and open-sourcing of agent training methodologies alongside benchmark improvements.
Five Major AI Models Released in Under Two Days
Between September 22 and 23 (Beijing time), four vendors shipped new models in rapid succession: xAI's Grok 4.7, Xiaomi's MiMo V2.6 series (Pro, Flash, Pro-UltraSpeed), Anthropic's Claude Opus 5.5, and OpenAI's GPT-6 Sol and GPT-6 Luna. The releases signal three structural shifts: models are being sold in tiers, competition is moving from exam-style benchmarks to real-world task execution, and open-source projects are beginning to release full agent training pipelines.
Grok 4.7: Larger Parameters, Modest Gains
Grok 4.7 is a ~2.1T parameter pre-trained model, up from ~1.5T for Grok 4.6. Context window remains 500K tokens. Standard pricing stays at $2/M input and $6/M output tokens; long-context (>200K input) rises to $4/$12. Its AA index reaches ~46, close to GPT-5.6 Sol, but the score is driven by knowledge and legal tasks. On Terminal-Bench 4.0 (coding agent benchmark) it scores around DeepSeek V4.1 Flash level, behind GLM 5.3 Flash. CritPt scientific reasoning shows no clear improvement over 4.6, and front-end design tasks sometimes regress. The release reads as a routine scale-up without breakthrough in coding, agent, or long-context capabilities.
Xiaomi MiMo V2.6: Three Variants, Aggressive Pricing, Open RL Pipeline
Xiaomi launched three models simultaneously:
MiMo-V2.6-Pro – 1.02T total / 42B active parameters, high-performance flagship.
MiMo-V2.6-Flash – 309B total / 15B active, low-cost high-concurrency.
MiMo-V2.6-Pro-UltraSpeed – Pro variant with up to ~20× output speed.
Both Pro and Flash support 1M context and native multimodal input (text, image, video, audio). MiMo V2.6 Pro scores 46.32 on the Artificial Analysis Intelligence Index, matching Grok 4.7 and jumping from V2.5's 26. Xiaomi's model card reports Terminal-Bench 4.0 34.9, Terminal-Bench 2.1 89.9, DeepSWE v1.1 71.9 (note: different test setups). CritPt scientific reasoning nears the top tier. The model behaves as a "bucket model" covering bug fixes, feature work, front-end, and back-end tasks.
Pricing is notably low: Pro cached input ~¥0.025/M tokens, Flash ~¥0.02/M; batch inference can approach half price. UltraSpeed mode targets ~20× throughput (actual speed depends on deployment).
Crucially, Xiaomi open-sourced weights on Hugging Face the same day, plus a technical report, 7,000+ high-quality RL task environments, mini-harnesses, and a distilled MiMo-V2.6-Distill-Qwen-9B model enabling small teams to reproduce agentic RL. The post-training RL cost was disclosed: ~$850K for Flash, ~$2.62M for Pro, total ~$3.47M in under six days. The team also live-streamed the post-training process. MiMo Desktop launched with custom API key support; invite-only testing ends in one week.
Claude Opus 5.5: Coding Execution at Lower Cost
API identifier: claude-opus-5-5. Anthropic positions it as matching Fable 5.1 on most tasks while cutting Opus 5 running cost by 40%. Specs: 1M context, 128K max output, knowledge cutoff June 2026, Adaptive Thinking enabled by default at medium effort. Pricing: $4/M input, $20/M output, $0.20/M cache read. Output speed improved >30% over Opus 5.
Key benchmark: Terminal-Bench 4.0 66.4% (new high), excelling at complex coding and multi-step terminal execution. Terminal-Bench Science 58.7% (below GPT-6 Astra's 64.6%), AutomationBench 40.0% (slightly below Astra's 41.4%). This reveals a divergence: Claude emphasizes coding, code aesthetics, and detail execution; OpenAI leans toward cross-software agents, research, and broad automation.
Communication style changed: key information surfaced earlier, less hedging and jargon, stronger "human feel" in writing and code review. Best suited for daily feature work, debugging, code review; high-stakes, unsupervised, long-horizon planning still demands higher-tier models.
GPT-6 Sol and Luna: Splitting the Flagship into Two Tiers
OpenAI released Sol and Luna together, keeping Astra as the flagship. Sol targets daily execution: coding, tool use, business process automation, general agent tasks. Luna serves as a backend engine for massive information processing, classification, extraction, summarization, and high-volume request workloads.
Pricing is the headline: Sol and Luna are up to 50% cheaper than GPT-5.6 equivalents. Luna's raw I/O pricing undercuts some DeepSeek tiers, though cache pricing remains higher. AutomationBench data shows Sol achieves higher completion rates at markedly lower per-task cost versus Opus 5 (note: Opus 5.5 reached 40.0% on its own AutomationBench run; test setups differ, so treat as cost-capability signal, not strict ranking). Sol's factual error rate approaches Astra; for factual tasks, set reasoning level to high or xhigh.
Trade-off: Sol lags Astra in high-end aesthetics and complex detail generation (e.g., Blender motorcycle, Temple of Heaven modeling show structural errors). Luna's knowledge cutoff updated to May 18, suited for high-concurrency backend; factual tasks also benefit from higher reasoning levels. Rollout began in ChatGPT Work and Codex; standard Chat not fully switched at press time.
Three Structural Shifts
1. Models Are Now Sold in Tiers
Vendors no longer push a single flagship for everything. Claude offers Fable 5.1 and Opus 5.5; OpenAI has Astra, Sol, Luna; Xiaomi provides Pro, Flash, UltraSpeed. Flagships handle complex tasks, mid-tier models cover daily work, low-cost tiers serve massive-scale calls. Developers now choose the tier that matches the task's cost-capability profile.
2. Evaluation Shifts from Exams to Execution
Benchmarks now measure real work: Terminal-Bench tests terminal coding, AutomationBench tests cross-tool automation, CritPt tests scientific reasoning. The question changes from "How accurate is the answer?" to "Can it finish the job end-to-end?" — writing code, calling tools, executing dozens of steps, delivering a result.
3. Open Source Releases Training Methods, Not Just Weights
MiMo V2.6 open-sourced RL environments, training framework, mini-harnesses, and a distilled model — moving from "here is the model" to "here is how to train the agent." Meanwhile, proprietary vendors push stronger capabilities into cheaper daily tiers (Sol, Luna, Opus 5.5). The immediate outcome for developers: same budget runs more tasks; same task has more model options.
The real signal from this release wave isn't who topped a leaderboard. It's that models are layering, prices are falling, agent capability is becoming the battleground, and training recipes are opening. Intelligence isn't free yet, but the competition is shifting from "whose model is strongest?" to "how much for equivalent capability?"
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
AI Code to Success
Focused on hardcore practical AI technologies (OpenClaw, ClaudeCode, LLMs, etc.) and HarmonyOS development. No hype—just real-world tips, pitfall chronicles, and productivity tools. Follow to transform workflows with code.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
