Opus 5.5 vs GPT-6 Sol: Cost, Capability, and the AI Pacing Paradox

Anthropic and OpenAI released new models days after their CEOs advocated for slower AI development; independent benchmarks show Opus 5.5 excels at complex knowledge work while GPT-6 Sol cuts task costs by half, but Luna trades coding ability for price, revealing tension between safety rhetoric and commercial competition.

Design Hub
Design Hub
Design Hub
Opus 5.5 vs GPT-6 Sol: Cost, Capability, and the AI Pacing Paradox

Placing the Three New Models in Context

On September 12, 2026, Anthropic CEO Dario Amodei called for "pacing" frontier AI capability progress; OpenAI CEO Sam Altman supported external audits. Ten days later, on September 22, Anthropic released Claude Opus 5.5, and OpenAI released GPT-6 Sol and GPT-6 Luna. The releases test whether "slow down" rhetoric holds under commercial pressure.

One of the official release animations from both companies
One of the official release animations from both companies

Opus 5.5: Capability Ceiling as a Mature Work Partner

Opus 5.5 is the first model in Anthropic's 5.5 family. Anthropic claims it matches Claude Fable 5.1 on most tasks, with typical task runtime cost down ~40% and output speed up >30% versus Opus 5. API pricing: $4 per million input tokens, $20 per million output tokens (20% lower than Opus 5); cache reads at $0.20. The "40% runtime cost reduction" combines price and token usage changes, not a uniform per-token drop.

Anthropic highlights strengths in agentic coding, computer use, and knowledge work. Reported benchmarks: Terminal-Bench 4.0 66.4%, FrontierCode 54.4%, GDPval-AA knowledge work 1846 Elo. Early customer cases: one migrated ~680k lines of code in a day; another reduced a 200k-line audit and fix from 20+ hours to under 3 hours. These cases are illustrative, not guaranteed averages.

Second segment of Claude official video
Second segment of Claude official video

Independent evaluation by Artificial Analysis gave Opus 5.5's highest reasoning tier a score of 58 (highest at time of testing) and 1822 Elo on AA-Briefcase professional knowledge work, leading in analysis quality and deliverable presentation. The evaluation used Anthropic's default safety fallback, so results reflect the actual service configuration, not every request handled solely by Opus 5.5.

Independent evaluation intelligence index and cost positioning
Independent evaluation intelligence index and cost positioning
Opus 5.5 knowledge work performance
Opus 5.5 knowledge work performance

Developer Alex Albert demonstrated Opus 5.5 driving Blender to build a 1906 San Francisco Market Street scene, showing improved spatial reasoning and tool orchestration. However, a single demo does not guarantee historical accuracy, model structure, material quality, or reusability for production pipelines.

Opus 5.5 building scene in Blender
Opus 5.5 building scene in Blender

GPT-6 Sol and Luna: Real Cost Savings, Capability Gains Not Universal

OpenAI positions Sol as the workhorse for complex daily tasks, Luna as a low-cost tier for high-volume calls, while GPT-6 Astra remains the top capability tier. Sol API pricing: $2/$10 per million input/output tokens; Luna: $0.10/$0.50. OpenAI claims ~50% price reduction versus GPT-5.6 series promotional prices. Both available via API, with gradual rollout to ChatGPT Work and Codex users; free and Go users can try Luna in the desktop app.

GPT-6 Sol and Luna official release video
GPT-6 Sol and Luna official release video

List price alone is misleading: Sol's input/output prices are half of Opus 5.5's, Luna an order of magnitude cheaper. Actual cost depends on reasoning tier, cache hits, tool calls, and retry rates.

OpenAI reports Sol improvements on AutomationBench, DeepSWE, OSWorld. In DeepSWE, Sol's highest reasoning tier scores 68.8%, Luna 66.6%; Sol scores 60.5% on OSWorld 2.0 offline reward. Internal factuality eval shows ~50% fewer errors, with more direct, less verbose answers.

Critical caveat: Official comparisons use older baselines: OpenAI compares against old Claude Opus 5 or Fable 5.1; Anthropic's tables compare against old GPT-5.6 Sol. Neither release provides a head-to-head "same-day exam" of the two new models.

Artificial Analysis finds GPT-6 Sol's highest tier single-task cost ~$1.06, roughly half of GPT-5.6 Sol; Luna ~$0.07. Sol's coding agent index rises from 55 to 57, but Luna's falls from 43 to 41. "Cheaper" is more certain than "across-the-board stronger."

Sol and Luna independent evaluation cost changes
Sol and Luna independent evaluation cost changes
Coding agent index changes
Coding agent index changes

Sol's hallucination rate drops significantly, but attempt rate falls from 99% to 83%, and accuracy drops from 59% to 54%. This trade-off between caution and usefulness matters for research, support, or medical products: lower hallucinations alone are insufficient without considering refusal and completion rates.

Factuality and refusal trade-off
Factuality and refusal trade-off

Real-World Testing: Experience Adds Contradictions Beyond Leaderboards

Developer Simon Willison integrated all three models into daily work. He praises Luna's pricing and sets Sol and Opus 5.5 as defaults in Codex and Claude Code. However, he recorded a failure: Opus 5.5 at max reasoning failed twice to draw a pelican on a bicycle, hitting output limits after ~20 minutes and $2.56 each. Higher reasoning tiers do not guarantee better results; runaway thinking costs are part of product experience.

Design and delivery performance in independent testing
Design and delivery performance in independent testing

Product lead Claire Vo ran a blind test across email, PRD, frontend prototype, backend, long tasks, SVG, and video editing. Her verdict: Astra most surprising; Opus 5.5 more stable on long tasks and B2B frontend; Sol attractive for clear writing, readable PRDs, and price. Human aesthetic preferences diverged from model judge preferences. Delivery, editability, and maintainability outweigh "first-look wow."

Some users quickly judged Sol "worse, just cheaper" from screenshots. Independent benchmarks show knowledge work delivery regressions alongside improvements in coding agent, automation, and factuality metrics. Results vary by reasoning tier, tools, and scoring method. Labeling a single regression as "across-the-board decline" is as unreliable as treating marketing as "total victory."

Early adopter Lovable's internal eval: Sol beats GPT-5.6 Sol by 6-12% on zero-to-app builds, with notable gains in brief adherence and design fundamentals. Opus 5.5 completes tasks in existing codebases with fewer steps. These are product-team workflow data with specific samples, criteria, and competitive interests—not universal leaderboards.

The Real Tension Between "Pacing" and Continued Releases

Amodei's original text does not demand immediate training or release halts. "Pacing" means capability growth should not outrun alignment, auditing, monitoring, and third-party evaluation; he proposes continuous, near-employee-level access for external evaluators. Altman publicly supports this and clarifies pacing ≠ stopping.

Anthropic's Opus 5.5 release notes address the question: current model risks are manageable with existing tests and guards; METR and Frontier Design conducted pre-release external evaluations. For future systems that could automate AI research, Anthropic does not believe current measures are sufficient.

This explains "why they can still release," but does not prove the industry has solved verification. Key questions remain: What is the scope of external evaluation? What negative results can third parties publish? Is there a clear incident response and rollback mechanism? These matter more than whether a release contradicts a speech. Commercial competition speed has not slowed; verifiable progress should be measured by transparency and enforceability of safety evidence.

Practical Selection Guide for Designers, Product Managers, and Developers

Long-horizon research, complex code migration, professional delivery requiring repeated verification: Try Opus 5.5 first. Higher cost than Sol, but independent benchmarks and early cases show stronger performance on hard tasks and deliverable quality. Start with default or medium reasoning; do not treat "max" as automatic gain.

Continuous product iteration, daily coding, PRD rewriting, multi-round prototyping: Sol is competitive. Measure not single answers but total rounds from brief to shippable delivery, human rework time, and omissions. Embed "completion criteria" in prompts and acceptance checklists to prevent brief answers from missing key requirements.

High-volume classification, extraction, database QA, low-risk automation: Luna's price is compelling. But it regresses on some coding and knowledge work benchmarks; do not offload tasks needing full judgment just because it's cheap. Better: use Luna for triage, escalate uncertain or high-value tasks to Sol or Opus.

Independent evaluation comparison of model capabilities and costs
Independent evaluation comparison of model capabilities and costs

The most interesting aspect of this release cycle is not "who won" but that two directions hold simultaneously: top-tier models raise the ceiling for complex work, while cheap models make previously uneconomical product flows viable. Safety discussions must therefore look beyond training speed; as cheaper agents embed into more real systems, deployment scale, permission boundaries, and accountability will shape the next risk frontier.

Data as of September 23, 2026. Official scores, independent benchmarks, and personal tests use different conditions; noted separately, not combined into a single leaderboard. Animated images are short excerpts from public demos.
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

model comparisonindustry analysisAI SafetyAI pricingAI benchmarksClaude Opus 5.5GPT-6 LunaGPT-6 Sol
Design Hub
Written by

Design Hub

Periodically delivers AI‑assisted design tips and the latest design news, covering industrial, architectural, graphic, and UX design. A concise, all‑round source of updates to boost your creative work.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.