Benchmarks · round 2 · live results

Models, priced by the job.

A trial is a task plus its test suite. A model passes only when every test goes green — retries priced in, wall clock running. Each trial is a public page you can share. Tasks run easy → hard — watch the strikeouts grow.

Round 2 — 5 trials, easy → hardrank by:
taskhaiku-4-5gpt-5-minigpt-4.1sonnet-4-6gpt-5opus-4-8fable-5gpt-5.1gpt-5.2gpt-5.4gpt-5.5sonnet-5
T01 · Fix the failing test✓✓✓✓✓✓R✓✓✓✓✓
T02 · Access-log parser✓✓✓✓✓✓R✓✓✓✓✓
T03 · CSV with messy input✓✓✓✓✓✓R✓✓✓✓✓
T04 · Token-bucket rate limiter✓✓✓✓✓✓✓✓✓✓✓✓
T05 · Mini alerting DSL✓✓✓✓✓✓R✓✓✓✓✗

✓ passed · ✗ DNF · R refused · highlight = winner by the chosen metric. Hover a task name for the full ranking.

Model totals — all 5 trialsgreen / DNF / R = refused
modeltrialsgreentotal costtotal time
#1claude-haiku-4-55/5$0.030556s
#2gpt-5-mini5/5$0.04032m02s
#3gpt-5.45/5$0.058544s
#4gpt-5.25/5$0.065782s
#5gpt-4.15/5$0.077261s
#6claude-opus-4-85/5$0.096377s
#7gpt-5.55/5$0.112759s
#8claude-sonnet-4-65/5$0.15952m48s
#9gpt-5.15/5$0.16301m40s
#10gpt-55/5$0.28292m41s
#11claude-sonnet-54/5$0.39024m30s
#12claude-fable-5RRRR1/5$0.03512m05s

Ranked by jobs completed first, always — then by the metric you pick. Cost and time are independent columns: the cheapest model is not the fastest.

Run through real Cerver sessions — metered, capped, transcript kept. The benchmark is a byproduct of the product.

Round 2 correction: round 1 showed cheap models "hitting a wall" — six of those DNFs were OUR bug (a transcript race made retries blind; the benchmark itself caught it). Fixed, re-run: 10 of 11 models now finish 5/5, and the cheapest model wins outright. This is why benchmarks need receipts.

claude-fable-5 (1/5): its failures are REFUSED verdicts — the model's API-side safety layer stochastically declines parser-writing tasks (it refused "parse a date string"; rich system prompts don't help). The same weights inside Claude Code complete all five tasks. Deployment context is part of the benchmark.

Make your own trial
your task, your tests, any model — free tier, no card