GPT-6 Sol, Luna & Claude Opus 5.5 Released: Price Drops, Benchmarks & How to Choose
OpenAI launched GPT-6 Sol and Luna while Anthropic released Claude Opus 5.5, all emphasizing lower costs and improved capabilities; the article compares pricing, benchmarks, cache economics, and independent evaluations, advising developers to test models on their own tasks to verify real-world cost per qualified result.
Three Major Model Releases in One Night
On September 23 (Beijing time), Anthropic released Claude Opus 5.5 and roughly 100 minutes later OpenAI added GPT-6 Sol and GPT-6 Luna to the GPT-6 family. Both vendors highlight stronger, faster, or more cost‑efficient models, with pricing reductions featured prominently.
GPT-6 Sol and Luna: New Low‑Price Tiers
OpenAI positions gpt-6-sol for complex coding and agent workflows, and gpt-6-luna for budget‑sensitive, high‑volume tasks; the existing gpt-6-astra continues to serve the most demanding workloads.
Price cuts: Sol’s regular input and output prices are halved versus GPT‑5.6 Sol. Luna’s input is halved and output drops from $1.20 to $0.50 per million tokens (≈58% reduction).
20× price gap: Sol and Luna differ by a factor of 20 in per‑token pricing, making Luna a viable candidate for classification, extraction, and format‑conversion tasks.
AutomationBench chart: OpenAI’s own chart shows Sol at xhigh reasoning intensity scoring 33.2% at $0.27 per task, while Opus 5 at max scores 26.9% at roughly 11.1× the cost. The comparison uses Opus 5 (not 5.5), different tools/reasoning settings, and cost‑per‑task (not cost‑per‑successful‑delivery).
Reduced misleading completions: OpenAI reports fewer cases where the model claims a coding task is done when it is not, based on a targeted “deception” evaluation.
Availability: Plus, Pro, Business, Enterprise, and Edu users can access Sol/Luna in ChatGPT Work and Codex; Free and Go users get Luna in the desktop app. API model names are gpt-6-sol and gpt-6-luna.
Claude Opus 5.5: Capability, Speed, and Writing Style Improvements
Anthropic’s first Claude 5.5 family model, Opus 5.5 , targets agent coding, computer operation, and knowledge work. Sonnet 5.5 and Haiku 5.5 are slated for later release.
Benchmark claims: Terminal‑Bench 4.0: Opus 5.5 66.4% ( xhigh ) vs Opus 5 52.3%; FrontierCode v1.1: 54.4% vs 48.0% (both at max reasoning). Different reasoning tiers are used, so direct column‑by‑column comparison is cautioned.
Default‑tier efficiency: Anthropic emphasizes that Opus 5.5 at default settings beats GPT‑6 Astra at roughly one‑fifth the per‑task cost on FrontierCode, and reaches similar Terminal‑Bench 4.0 performance at about 40% of the per‑attempt cost.
Writing style: Opus 5.5 presents key information first, uses less jargon, and adheres better to user‑specified writing rules. An official example shows Opus 5.5 explaining a billing discrepancy before diving into code.
Safety: Best score in Anthropic’s automated behavioral audit; pre‑release external evaluations by METR and Frontier Design (participation ≠ endorsement of all capability/cost claims).
Availability: On AWS, Google Cloud, Microsoft Azure, and Claude Platform (model name claude-opus-5-5). Pro, Max, Team, and per‑seat Enterprise plans receive increased 5‑hour quotas plus a one‑time quota reset.
Head‑to‑Head Pricing Comparison
Standard API prices (USD per million tokens):
GPT‑6 Sol: input $2, cache read $0.20, cache write $2.50, output $10.
GPT‑6 Luna: input $0.10, cache read $0.01, cache write $0.125, output $0.50.
Claude Opus 5.5: input $4, cache read $0.20, cache write $5, output $20.
Sol/Luna: requests exceeding 272k input tokens double input/cache prices and multiply output by 1.5×. Opus 5.5 regular input/output are 20% lower than Opus 5; cache read is 60% lower. Anthropic’s “typical workload cost down 40%” combines price cuts with fewer tokens used — not a guaranteed 40% bill reduction for every user.
Speed: Opus 5.5 output >30% faster; Fast mode (Claude Code/Platform) up to 2.5× speed but at $8/$40 per million tokens input/output.
Cache Economics: Hit Rate Matters More Than Unit Price
Both Sol and Opus 5.5 charge $0.20 per million tokens for cache reads, but cache write prices differ ($2.50 vs $5).
OpenAI’s “up to 90% cache discount” applies only to hit input tokens; output is billed normally, and the first cache write costs 1.25× regular input.
Worked example (Sol): 100k input / 10k output per turn, 80k fixed prefix cached. Turn 1: $0.340. Turn 2 (full hit, no new write): $0.156. Two‑turn total $0.496 vs $0.600 without cache ≈17% savings — not 90%. Cache prefix reusable for 30 minutes after last write/reuse.
Continuous edits around the same project benefit most; one‑off short requests may not justify cache engineering.
Independent Evaluation: Price Cuts ≠ Uniform Capability Gains
Artificial Analysis testing reveals nuance:
Intelligence Index (max tier): Sol per‑task cost fell from $1.99 to $1.06; Luna from $0.18 to $0.07. Output token usage slightly increased, so savings stem mainly from price reductions.
Coding Agent Index: Sol improved by 2 points; Luna regressed by 2 points. Some knowledge‑work tasks showed regressions (missed requirements).
Environment and task sets differ from vendor benchmarks.
Practical Recommendation: Test With Your Own Workload
Run a representative batch of your recurring tasks with the new models and your current model under identical acceptance criteria.
Simple classification/extraction → try Luna first.
Complex coding → include both Sol and Opus 5.5 in the candidate set.
Heavy repeated document/code reading → measure cache hit rates.
Track total spend (model calls, tool use, human verification) divided by accepted tasks. Count retries, model upgrades, and human cleanup in total cost; also monitor latency and unacceptable errors.
The real question: with the same budget, how many more qualified results can you deliver?
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
ShiZhen AI
Tech blogger with over 10 years of experience at leading tech firms, AI efficiency and delivery expert focusing on AI productivity. Covers tech gadgets, AI-driven efficiency, and leisure— AI leisure community. 🛰 szzdzhp001
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
