Back to CAD-Bench

Leaderboard

Per-cell results across the full benchmark run. Each row is one (agent, model) pair scored on 100 tasks. Scores are 0–100 percentage points; the combined score is the harmonic mean of geometry similarity and CAD/spec consistency.

RankModelAgentAgent Ver.Geom ScoreSpec ScoreCombinedTokens (mean)Tokens (total)Cost (mean)Cost (total)Task CountDate
1
claude-opus-5(max)claude-code2.1.22688.997.290.6690.35K69.03M$1.132$113.241002026-08-11
2
grok-4.6(high)grok-build0.2.5283.597.287.8910.39K91.04M$0.692$69.231002026-08-11
3
gpt-5.6-sol(max)codex0.147.084.894.386.5684.24K68.42M$0.976$97.581002026-08-11
4
grok-4.5(high)grok-build0.2.5280.294.183.7621.8K62.18M$0.366$36.591002026-08-11
5
gpt-5.5(high)codex0.130.080.893.983.21.12M112.39M$1.700$170.001002026-05-13
6
gemini-3.1-pro-preview(high)mini-swe-agent2.2.874.295.479.0589.37K58.94M$0.708$70.821002026-05-13
7
gpt-5.5(high)mini-swe-agent2.2.871.489.374.4132.47K13.25M$0.423$42.351002026-05-13
8
claude-opus-4-7(max)mini-swe-agent2.2.869.389.073.4240.7K24.07M$0.298$29.841002026-05-13
9
gemini-3.1-pro-preview(high)gemini-cli0.42.068.983.972.9741.76K74.18M$0.513$51.281002026-05-13
10
claude-opus-4-7(max)claude-code2.1.14062.483.365.5709.03K70.9M$0.732$73.251002026-05-13
11
claude-sonnet-4-6(max)claude-code2.1.14047.868.351.81.17M117.13M$1.014$96.341002026-05-13
12
claude-haiku-4-5claude-code2.1.14023.354.128.42.09M209.02M$0.247$24.671002026-05-13
13
gemini-3.1-flash-lite-previewmini-swe-agent2.2.819.647.624.1177.73K17.77M$0.030$3.001002026-05-13
14
gemini-3.1-flash-lite-previewgemini-cli0.42.010.024.211.91.4M139.84M$0.076$7.631002026-05-13