gradientsmith

Cost / quality frontier

Campaign frontier-calibration · 6 models · 200 rollouts each · scored on public and hidden tests

2026-08-12
0%25%50%75%100%$0.00015/call$0.00950/callunit cost per attempt (log scale) →claude-opus-5claude-sonnet-5claude-haiku-4-5gpt-oss-20bqwen3.5-27bqwen3.5-9b

marks a model on the Pareto frontier, meaning no other model is both cheaper and better at the same time. Dimmed points are dominated by something above and to the left of them. Click a point for the per-model breakdown.

Leaderboard

pass@1 clean is measured on the public tests the solver can see. pass@1 adv adds the hidden and adversary-mined tests. The gap between the two numbers is how much a model overfits to what it can see.

modelpass@1 cleanpass@1 advgapp50p95$/attempt$/solved
claude-opus-594%89%5%2400 ms5200 ms$0.00950$0.0107
claude-sonnet-590%83%7%1500 ms3400 ms$0.00380$0.0046
claude-haiku-4-578%66%12%700 ms1600 ms$0.00070$0.0011
gpt-oss-20b69%61%8%950 ms2200 ms$0.00085$0.0014
qwen3.5-27b72%58%14%1100 ms2600 ms$0.00045$0.0008
qwen3.5-9b55%40%15%600 ms1400 ms$0.00015$0.0004