
None
None
Long Story Video Skill
Turn your idea into 10s–10min video—AI writes scripts, prompts, and generates footage automatically.

Ads Video Skill
Generate professional ads and sales videos—AI auto-generates scripts, prompts, and footage.
3D Science Explainer Video Skill
Convert scientific concepts into stunning 3D explain animations
Feedback
😢 Every free AI generation costs us real GPU money. We pay out of pocket to keep this free for you.
💎 Upgrade to Pro and unlock:
- ✓HD Download
- ✓No Watermark
- ✓Unlimited Generations
Haiku 5.5 Benchmark
A source-checked read of the Haiku 5.5 Benchmark: launch scores, price bands and published limits, with no matched hands-on test by Kingy.ai.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator
Seedance 2.5
The Future of AI Video Is Here.

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator
Why Read the Haiku 5.5 Benchmark
Published launch results, price bands and per-effort costs, each kept with its configuration and retrieval date.
- What the Launch Numbers Actually MeasureElo, partial-credit, coding and accuracy figures use incompatible scales and cannot be averaged.
- Small-Model Economics for Repeated WorkBuilt for extraction, classification and summaries that repeat thousands of times.
- Two Caveats That Change the Score TablesCheap rates stop at 100K tokens; 72.4% computer use is partial credit, not a pass rate.
How to Read the Haiku 5.5 Benchmark
Three steps for turning published Haiku 5.5 benchmark numbers into a workload decision.
Haiku 5.5 Benchmark: Key Findings
Published specifications, price bands, coding results, computer-use scores and independent index data for Claude Haiku 5.5.
Context Window and Output Contract
Inputs are text plus images and output is text, with adaptive thinking on. One million tokens of context and a 128,000-token output cap dwarf Haiku 4.5's 200,000 and 64,000.
Two Prompt-Length Price Bands
Cross 100,000 prompt tokens and input, output and caching each jump to a band five times the cheap one. OpenAI's base Standard band stretches to 272,000 input tokens.
Terminal-Bench 4.0 Completed Work
Haiku's 39.2% launch figure beats Luna clearly but stays under the published Sol rows. Task runs get an eight-hour agent timeout, and 2.1 or 3.0 numbers belong to an older experiment.
FrontierCode Mergeability and Effort Cost
Grading covers correctness, tests, scope, style and repo conventions. Haiku lands at 46.4% on max effort and 41.6% at medium for $0.13 a rollout, while xhigh to max doubles spend for half a point.
OSWorld 2.1 Partial Credit Versus Strict Pass
Eighty-two offline tasks, five attempts apiece, 1080p screenshots and a 500-step ceiling define the run. For a workflow that must finish, the 37.1% strict rate is the honest number.
Independent Index Scores and Tool Uplift
Artificial Analysis posts Haiku at all five efforts on Intelligence Index v4.3.2 with rounded task costs. HLE reads 45.9% without tools and 57.4% with them — a gain from the surrounding system.
Haiku 5.5 Benchmark — FAQ
Common questions about Claude Haiku 5.5 benchmark results, pricing and limits.
How does Haiku 5.5 compare with GPT-6 Luna and GPT-6 Sol?
On knowledge work, screenshot tasks and narrow coding, Haiku trades blows with Luna. Sol and Sol 6.1 cost more and earn it on coding, while Sonnet 5.5 still leads on most rows.
Why is the 72.4% computer-use score not a success rate?
The same run reports only 37.1% when every checkpoint must be met. An agent can locate and edit the right file, then save it in the wrong format or stop before the last action. Add an end-state check.
When does Haiku 5.5 switch to its higher price band?
Only prompts up to 100,000 tokens get the cheapest rates; above that, every category costs five times more. Anthropic's 90%, 50% and ~75% savings figures compare different baselines.
Can Haiku 5.5 run locally, and how much GPU memory would it need?
No parameter count, active-parameter figure, layer count or training compute appears in the published specs. 'Small' labels a product tier, not an architecture, so local GPU estimates have no basis.
Does spending more reasoning effort always improve results?
Not always. Sol 6.1 scores better at medium than at max, and Haiku's 0.2-point edge over Sonnet at max is too small to call a win. Extra effort brings more tokens, latency and chances to drift off task.
Why do the same models show different numbers in different tables?
Sonnet reads 70.6% in Anthropic's launch table and 61.8% on the public leaderboard snapshot, where Haiku had no visible row. Artificial Analysis dropped its 'Default Fallback' label, yet Anthropic's migration guide still says no server-side fallback exists.
Test the Haiku 5.5 Benchmark Against Your Workload
Pick one workload, set the effort level, and test Haiku 5.5 against your own quality floor.
