Claude Haiku 5.5 Outperforms GPT-6 Luna, Undercuts DeepSeek Flash on Price
Anthropic's new Claude Haiku 5.5 surpasses GPT-6 Luna and Chinese rivals on benchmarks while matching Luna's pricing and undercutting DeepSeek-V4.1 Flash, though tokenizer changes increase token usage and API breaking changes require migration effort.
Anthropic has released Claude Haiku 5.5, a new small model that the article evaluates across benchmarks, pricing, capabilities, and migration considerations.
Benchmark Performance
On official benchmarks, Haiku 5.5 beats GPT-6 Luna across all reported metrics. It also surpasses two previous "gatekeeper" small models: DeepSeek V4.1 Flash and GLM-5.3-Flash, becoming the new small-model gatekeeper. The article notes the pelican-on-bicycle test is saturated, so the newer running-zebra test is now used.
Pricing Analysis
Haiku 5.5 pricing is aligned with GPT-6 Luna. For prompts ≤100k tokens (covering 90% of Haiku 4.5 requests), input/output prices are cut 90% versus the previous generation; above 100k tokens the discount is 50%. Haiku 4.5 launched a year ago at 10× the price of 5.5. Even excluding cached-input costs, Haiku 5.5 is cheaper than DeepSeek-V4.1 Flash's off-peak pricing for prompts within 100k tokens.
Capability Gains and Effort Tiers
Haiku 5.5 is the first Haiku to support effort adjustment (low to max, five tiers). Haiku 4.5 scored near zero on computer use and coding; 5.5 shows dramatic improvement. On OSWorld 2.1 (real-computer multi-step tasks): Low tier 42.0% accuracy at $0.07 per run, Max tier 72.4% at $0.61. Haiku 4.5 Max managed only 15.7% at $1.45 — 5.5 Low outperforms 4.5 Max at half the cost. On GDPval-AA v2.1 (44 professional roles): Low Elo 1125 at $0.01, Max Elo 1620 at $0.87; 4.5 Max reached only Elo 735 at $0.24.
Tokenizer Change Impact
Haiku 5.5 adopts the same tokenizer as Sonnet 5.5 and Opus 5.5. Identical tasks consume more tokens, so actual savings are slightly less than list-price discounts suggest. Inflation is especially pronounced for code, tables, and non-English content. In AA evaluation, average task-completion cost did not reach the ideal range despite lower API prices.
Cannot Replace Sonnet for Complex Work
On Terminal-Bench 4.0, Haiku 5.5 scores 39.2% versus Sonnet 5.5's 70.6%. Gaps remain large in complex multi-step coding, cross-file refactoring, and long-horizon autonomous planning. Anthropic recommends Sonnet 5.5 or Opus 5.5 for complex agent coding. Haiku 5.5's value lies in the execution layer: well-defined, parallelizable tasks with clear acceptance criteria. Example: Cognition's Devin uses Opus 5.5 as primary model and Haiku 5.5 as sub-agents, achieving 66.2% on FrontierCode — higher than either model alone.
Breaking API Changes (Migration Required)
Upgrading from Haiku 4.5 is not a drop-in replacement. Five breaking changes: (1) budget_tokens manual thinking config errors; must use adaptive thinking with effort parameter. (2) temperature, top_p, top_k locked to defaults; sampling-based creativity/diversity logic must be rewritten. (3) Assistant message prefill removed; forced-JSON-openers must switch to tool calls or structured output. (4) Computer-use tool version updated from computer_20250124 to computer_toolset_20260801 with different interface/response structure. (5) Adaptive thinking enabled by default; first response block may be a thinking block, requiring parsing by type field. Anthropic provides a migration guide.
Additional Updates
Sonnet 5.5 cache-read price drops from $0.20/M to $0.10/M, cutting typical agent-task costs ~20%. Max and Team subscribers now receive monthly API credits: Max 5× $100, Max 20× $200, Team up to $500 shared, usable across all models for experimentation. The article closes noting OpenAI's response was a "reset card" giveaway.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
