Claude Haiku 5.5 Outperforms GPT-6 Luna, Undercuts DeepSeek Flash on Price

Anthropic's new Claude Haiku 5.5 surpasses GPT-6 Luna and Chinese rivals on benchmarks while matching Luna's pricing and undercutting DeepSeek-V4.1 Flash, though tokenizer changes increase token usage and API breaking changes require migration effort.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Claude Haiku 5.5 Outperforms GPT-6 Luna, Undercuts DeepSeek Flash on Price

Anthropic has released Claude Haiku 5.5, a new small model that the article evaluates across benchmarks, pricing, capabilities, and migration considerations.

Benchmark Performance

On official benchmarks, Haiku 5.5 beats GPT-6 Luna across all reported metrics. It also surpasses two previous "gatekeeper" small models: DeepSeek V4.1 Flash and GLM-5.3-Flash, becoming the new small-model gatekeeper. The article notes the pelican-on-bicycle test is saturated, so the newer running-zebra test is now used.

Pricing Analysis

Haiku 5.5 pricing is aligned with GPT-6 Luna. For prompts ≤100k tokens (covering 90% of Haiku 4.5 requests), input/output prices are cut 90% versus the previous generation; above 100k tokens the discount is 50%. Haiku 4.5 launched a year ago at 10× the price of 5.5. Even excluding cached-input costs, Haiku 5.5 is cheaper than DeepSeek-V4.1 Flash's off-peak pricing for prompts within 100k tokens.

Capability Gains and Effort Tiers

Haiku 5.5 is the first Haiku to support effort adjustment (low to max, five tiers). Haiku 4.5 scored near zero on computer use and coding; 5.5 shows dramatic improvement. On OSWorld 2.1 (real-computer multi-step tasks): Low tier 42.0% accuracy at $0.07 per run, Max tier 72.4% at $0.61. Haiku 4.5 Max managed only 15.7% at $1.45 — 5.5 Low outperforms 4.5 Max at half the cost. On GDPval-AA v2.1 (44 professional roles): Low Elo 1125 at $0.01, Max Elo 1620 at $0.87; 4.5 Max reached only Elo 735 at $0.24.

Tokenizer Change Impact

Haiku 5.5 adopts the same tokenizer as Sonnet 5.5 and Opus 5.5. Identical tasks consume more tokens, so actual savings are slightly less than list-price discounts suggest. Inflation is especially pronounced for code, tables, and non-English content. In AA evaluation, average task-completion cost did not reach the ideal range despite lower API prices.

Cannot Replace Sonnet for Complex Work

On Terminal-Bench 4.0, Haiku 5.5 scores 39.2% versus Sonnet 5.5's 70.6%. Gaps remain large in complex multi-step coding, cross-file refactoring, and long-horizon autonomous planning. Anthropic recommends Sonnet 5.5 or Opus 5.5 for complex agent coding. Haiku 5.5's value lies in the execution layer: well-defined, parallelizable tasks with clear acceptance criteria. Example: Cognition's Devin uses Opus 5.5 as primary model and Haiku 5.5 as sub-agents, achieving 66.2% on FrontierCode — higher than either model alone.

Breaking API Changes (Migration Required)

Upgrading from Haiku 4.5 is not a drop-in replacement. Five breaking changes: (1) budget_tokens manual thinking config errors; must use adaptive thinking with effort parameter. (2) temperature, top_p, top_k locked to defaults; sampling-based creativity/diversity logic must be rewritten. (3) Assistant message prefill removed; forced-JSON-openers must switch to tool calls or structured output. (4) Computer-use tool version updated from computer_20250124 to computer_toolset_20260801 with different interface/response structure. (5) Adaptive thinking enabled by default; first response block may be a thinking block, requiring parsing by type field. Anthropic provides a migration guide.

Additional Updates

Sonnet 5.5 cache-read price drops from $0.20/M to $0.10/M, cutting typical agent-task costs ~20%. Max and Team subscribers now receive monthly API credits: Max 5× $100, Max 20× $200, Team up to $500 shared, usable across all models for experimentation. The article closes noting OpenAI's response was a "reset card" giveaway.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Model ComparisonAnthropicAI PricingAPI MigrationComputer UseLLM BenchmarksAgent BenchmarksClaude Haiku 5.5
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.