Benchmarks · round 2 · live results

Models, priced by the job.

A trial is a task plus its test suite. A model passes only when every test goes green — retries priced in, wall clock running. Each trial is a public page you can share. Tasks run easy → hard — watch the strikeouts grow.

Round 2 — 5 trials, easy → hardrank by:
taskhaiku-4-5gpt-5-minigpt-4.1sonnet-4-6gpt-5opus-4-8fable-5gpt-5.1gpt-5.2gpt-5.4gpt-5.5sonnet-5
T01 · Fix the failing testR
T02 · Access-log parserR
T03 · CSV with messy inputR
T04 · Token-bucket rate limiter
T05 · Mini alerting DSLR

✓ passed · ✗ DNF · R refused · highlight = winner by the chosen metric. Hover a task name for the full ranking.

Model totals — all 5 trialsgreen / DNF / R = refused
modeltrialsgreentotal costtotal time
#1claude-haiku-4-55/5$0.030556s
#2gpt-5-mini5/5$0.04032m02s
#3gpt-5.45/5$0.058544s
#4gpt-5.25/5$0.065782s
#5gpt-4.15/5$0.077261s
#6claude-opus-4-85/5$0.096377s
#7gpt-5.55/5$0.112759s
#8claude-sonnet-4-65/5$0.15952m48s
#9gpt-5.15/5$0.16301m40s
#10gpt-55/5$0.28292m41s
#11claude-sonnet-54/5$0.39024m30s
#12claude-fable-5RRRR1/5$0.03512m05s

Ranked by jobs completed first, always — then by the metric you pick. Cost and time are independent columns: the cheapest model is not the fastest.

Run through real Cerver sessions — metered, capped, transcript kept. The benchmark is a byproduct of the product.

Round 2 correction: round 1 showed cheap models "hitting a wall" — six of those DNFs were OUR bug (a transcript race made retries blind; the benchmark itself caught it). Fixed, re-run: 10 of 11 models now finish 5/5, and the cheapest model wins outright. This is why benchmarks need receipts.

claude-fable-5 (1/5): its failures are REFUSED verdicts — the model's API-side safety layer stochastically declines parser-writing tasks (it refused "parse a date string"; rich system prompts don't help). The same weights inside Claude Code complete all five tasks. Deployment context is part of the benchmark.

Make your own trial
your task, your tests, any model — free tier, no card