Claude Haiku 5.5: 90% Price Cut, Beats GPT-6 Luna on Agent Benchmarks

Anthropic launches Claude Haiku 5.5 with 90% lower input token pricing at $0.10 per million, outperforming GPT-6 Luna on OSWorld (72.4% success), GDPval-AA, and Terminal-Bench, while introducing adjustable reasoning effort, 1M context window, and enterprise adoption for high-volume agent workflows.

Machine Heart
Machine Heart
Machine Heart
Claude Haiku 5.5: 90% Price Cut, Beats GPT-6 Luna on Agent Benchmarks

Release Overview

On October 7, Anthropic released Claude Haiku 5.5 , the newest and cheapest model in the Claude 3.5 family. The company claims it is the fastest, most capable, and lowest-cost Haiku to date. This completes a 15-day rollout of the full Claude 3.5 lineup: Opus 3.5 on September 22, Sonnet 3.5 on September 28, and now Haiku 3.5.

Benchmark Performance

Anthropic published benchmark results showing Haiku 3.5 surpassing OpenAI's GPT-6 Luna on several agent-centric tests:

OSWorld 2.1 (offline subset) : Haiku 3.5 achieves 72.4% success rate , compared to 15.7% for Haiku 3.5 and 48.9% for GPT-6 Luna. OSWorld measures cross-application, multi-step computer operation.

GDPval-AA v2.1 (knowledge work) : Haiku 3.5 scores 1620 vs. GPT-6 Luna's 1437.

Terminal-Bench 4.0 (agent coding) : Haiku 3.5 reaches 39.2% vs. GPT-6 Luna's 16.4%.

However, the peak scores require higher reasoning effort. VentureBeat noted that the 39.2% Terminal-Bench result corresponds to maximum reasoning effort; at the default medium effort the score drops to ~20%. Artificial Analysis observed a similar pattern: maximum effort scores 43 on their Intelligence Index (cost ~$0.21/task), medium effort scores 34 (cost ~$0.05/task). Their independent Terminal-Bench 4.0 runs gave 33% (max) and 15% (medium), differing from Anthropic's numbers due to evaluation setup differences.

Adjustable Reasoning and Cost Trade-offs

Haiku 3.5 introduces an effort parameter (low/medium/high) letting developers control compute allocation. Higher effort improves task success but increases actual spend. The article emphasizes that evaluating small models now requires weighing reasoning effort, task completion rate, and real cost together .

Pricing Details and Tokenizer Change

Anthropic set two API price tiers:

Prompts ≤ 100K tokens: $0.10 per million input tokens, $0.50 per million output tokens — a 90% reduction from Haiku 3.5.

Prompts > 100K tokens: prices remain 50% below Haiku 3.5.

Anthropic estimates ~90% of prior Haiku requests fell under 100K tokens, projecting an average 75% cost reduction for equivalent workloads.

Caveat : Haiku 3.5 uses a new tokenizer that produces ~30% more tokens for the same text, so the 90% per-token price drop does not translate to a 90% bill reduction per task.

Competitive Landscape

GPT-6 Luna matches the base $0.10/$0.50 pricing but applies long-context surcharge only after 272K input tokens, giving Luna a unit-cost advantage in the 100K–272K range. Google Gemini models are pricier: Gemini 3.5 Flash-Lite at $0.30/$2.50, Gemini 3.8 Flash at $0.75/$3.75 per million tokens (per VentureBeat ). The article concludes that final task cost depends on actual token usage, reasoning budget, and success rates.

Enterprise Use Cases

Early adopters report gains in high-volume, well-scoped agent tasks:

Rogo (financial AI): Uses Haiku 3.5 as a subagent to extract revenue data from 10-K filings, feeding results to a stronger lead model. Values the accuracy/speed/price balance for repeated lookups.

Box : Internal eval shows +11 points and ~50% latency reduction vs. Haiku 3.5, enabling cost reports, financial summaries, and periodic reviews at scale. Full benchmark methodology not disclosed.

Asana : On bug triage, project creation, and portfolio search, Haiku 3.5 cuts task latency >30% and boosts single-step agent inference up to 2.5× versus their current model.

AlphaSense : Runs ~8 million weekly calls for document QA. On 400 test queries, Haiku 3.5 scores 0.84 vs. Haiku 3.5's 0.76. At millions of calls, even small per-call savings compound significantly.

All enterprise figures come from Anthropic-disclosed customer feedback; workloads, baselines, and metrics are not fully standardized.

Sonnet 3.5 Price Cut and Ecosystem Updates

Concurrently, Anthropic cut Sonnet 3.5 cache-read pricing by 50% (from $0.20 to $0.10 per million tokens), estimating ~20% lower cost for typical agent workloads that reuse long contexts. New monthly API credits for Claude Max/Team subscribers ($100–$500) encourage agent experimentation.

Haiku 3.5 is available via Anthropic API, Amazon Bedrock, Google Cloud, and Microsoft Azure (model ID: claude-haiku-5-5). Updated Python/TypeScript SDKs add beta browser and computer-use tooling.

Conclusion

The Claude 3.5 family now has clear role separation: Opus for hardest reasoning, Sonnet for daily complex work and coding, Haiku for high-frequency, cost-sensitive, fast-response tasks. The 15-day full-line refresh underscores a shift: as agents move from demos to production, developers must optimize which steps run on which model to achieve commercial viability. Haiku 3.5 makes previously uneconomical high-volume tasks financially feasible, pushing small-model competition into every execution layer of agent systems.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI-agentsbenchmarkenterprise AIAnthropictoken pricingAI model pricingGPT-6 LunaClaude Haiku 5.5
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.