The cheapest MiniMax H3 is here,$0.013/s
Haiku 5.5 Benchmark Model Record
One record in Kingy AI's source-backed model directory: Claude Haiku 5.5 benchmark results with provider, family, modality and API access notes.
None

None

None

Long Story Video Skill

Turn your idea into 10s–10min video—AI writes scripts, prompts, and generates footage automatically.

Ads Video Skill

Ads Video Skill

Generate professional ads and sales videos—AI auto-generates scripts, prompts, and footage.

3D Science Explainer Video Skill

Convert scientific concepts into stunning 3D explain animations

AI Video Prompt Generator

Feedback

😢 Every free AI generation costs us real GPU money. We pay out of pocket to keep this free for you.

💎 Upgrade to Pro and unlock:

  • ✓HD Download
  • ✓No Watermark
  • ✓Unlimited Generations

Haiku 5.5 Benchmark

A source-checked read of the Haiku 5.5 Benchmark: launch scores, price bands and published limits, with no matched hands-on test by Kingy.ai.

All Tools

Discover our comprehensive AI-powered animation toolkit

Why Read the Haiku 5.5 Benchmark

Published launch results, price bands and per-effort costs, each kept with its configuration and retrieval date.

  • What the Launch Numbers Actually Measure
    Elo, partial-credit, coding and accuracy figures use incompatible scales and cannot be averaged.
  • Small-Model Economics for Repeated Work
    Built for extraction, classification and summaries that repeat thousands of times.
  • Two Caveats That Change the Score Tables
    Cheap rates stop at 100K tokens; 72.4% computer use is partial credit, not a pass rate.

How to Read the Haiku 5.5 Benchmark

Three steps for turning published Haiku 5.5 benchmark numbers into a workload decision.

Haiku 5.5 Benchmark: Key Findings

Published specifications, price bands, coding results, computer-use scores and independent index data for Claude Haiku 5.5.

Context Window and Output Contract

Inputs are text plus images and output is text, with adaptive thinking on. One million tokens of context and a 128,000-token output cap dwarf Haiku 4.5's 200,000 and 64,000.

Two Prompt-Length Price Bands

Cross 100,000 prompt tokens and input, output and caching each jump to a band five times the cheap one. OpenAI's base Standard band stretches to 272,000 input tokens.

Terminal-Bench 4.0 Completed Work

Haiku's 39.2% launch figure beats Luna clearly but stays under the published Sol rows. Task runs get an eight-hour agent timeout, and 2.1 or 3.0 numbers belong to an older experiment.

FrontierCode Mergeability and Effort Cost

Grading covers correctness, tests, scope, style and repo conventions. Haiku lands at 46.4% on max effort and 41.6% at medium for $0.13 a rollout, while xhigh to max doubles spend for half a point.

OSWorld 2.1 Partial Credit Versus Strict Pass

Eighty-two offline tasks, five attempts apiece, 1080p screenshots and a 500-step ceiling define the run. For a workflow that must finish, the 37.1% strict rate is the honest number.

Independent Index Scores and Tool Uplift

Artificial Analysis posts Haiku at all five efforts on Intelligence Index v4.3.2 with rounded task costs. HLE reads 45.9% without tools and 57.4% with them — a gain from the surrounding system.

FAQ

Haiku 5.5 Benchmark — FAQ

Common questions about Claude Haiku 5.5 benchmark results, pricing and limits.

1

How does Haiku 5.5 compare with GPT-6 Luna and GPT-6 Sol?

On knowledge work, screenshot tasks and narrow coding, Haiku trades blows with Luna. Sol and Sol 6.1 cost more and earn it on coding, while Sonnet 5.5 still leads on most rows.

2

Why is the 72.4% computer-use score not a success rate?

The same run reports only 37.1% when every checkpoint must be met. An agent can locate and edit the right file, then save it in the wrong format or stop before the last action. Add an end-state check.

3

When does Haiku 5.5 switch to its higher price band?

Only prompts up to 100,000 tokens get the cheapest rates; above that, every category costs five times more. Anthropic's 90%, 50% and ~75% savings figures compare different baselines.

4

Can Haiku 5.5 run locally, and how much GPU memory would it need?

No parameter count, active-parameter figure, layer count or training compute appears in the published specs. 'Small' labels a product tier, not an architecture, so local GPU estimates have no basis.

5

Does spending more reasoning effort always improve results?

Not always. Sol 6.1 scores better at medium than at max, and Haiku's 0.2-point edge over Sonnet at max is too small to call a win. Extra effort brings more tokens, latency and chances to drift off task.

6

Why do the same models show different numbers in different tables?

Sonnet reads 70.6% in Anthropic's launch table and 61.8% on the public leaderboard snapshot, where Haiku had no visible row. Artificial Analysis dropped its 'Default Fallback' label, yet Anthropic's migration guide still says no server-side fallback exists.

Test the Haiku 5.5 Benchmark Against Your Workload

Pick one workload, set the effort level, and test Haiku 5.5 against your own quality floor.