Back to CAD-Bench
Parametric CAD Bench V3 Leaderboard
Current100 tasks: 30 CAD Creation from Text, 30 CAD Create And then Edit from Text, 40 CAD Create from Engineering Drawings. Overall score is the mean continuous task reward; failed or unscored trials count as zero. Scores are shown on a 0–100 scale.
| Rank | Model | Agent | Effort | Create | Create + Edit | Image-to-CAD | Overall (95% CI) | Scored | Perfect | Cost (USD) | Run |
|---|---|---|---|---|---|---|---|---|---|---|---|
1 | GPT-6 Astra | Codex 0.154.0 | max | 52.26 | 44.37 | 69.70 | 56.87% ± 6.85 | 100/100 | 5/100 | $317.81 | View run |
2 | Claude Fable 5.1 | Claude Code 2.1.270 | max | 55.48 | 36.75 | 72.71 | 56.75% ± 7.18 | 98/100 | 7/100 | $1,056.04 | View run |
3 | Claude Opus 5 | Claude Code 2.1.270 | max | 49.48 | 36.52 | 62.64 | 50.86% ± 7.45 | 99/100 | 6/100 | $1,012.21 | View run |
4 | Gemini 3.8 Flash | mini-swe-agent 2.4.6 | high | 47.94 | 31.22 | 31.10 | 36.19% ± 6.97 | 94/100 | 4/100 | $271.85 | View run |
5 | Grok 4.7 | Grok Build 1.0.30 | high | 45.57 | 34.63 | 26.91 | 34.82% ± 6.64 | 100/100 | 3/100 | $670.22 | View run |
6 | GPT-5.6 Sol | Codex 0.154.0 | max | 36.23 | 24.78 | 29.20 | 29.98% ± 6.17 | 100/100 | 1/100 | $291.99 | View run |
7 | Grok 4.6 | Grok Build 1.0.30 | high | 41.61 | 33.45 | 13.00 | 27.72% ± 6.30 | 100/100 | 2/100 | $390.94 | View run |
8 | Gemini 3.8 Flash | Antigravity 1.2.7 | high | 43.28 | 25.32 | 17.45 | 27.56% ± 6.72 | 100/100 | 4/100 | $196.74 | View run |
9 | Kimi K3 | mini-swe-agent 2.4.6 | max | 39.11 | 23.49 | 15.36 | 24.92% ± 6.40 | 96/100 | 1/100 | $314.55 | View run |
10 | Muse Spark 1.3 | mini-swe-agent 2.4.6 | max | 33.61 | 24.33 | 8.84 | 20.92% ± 5.81 | 99/100 | 1/100 | $357.03 | View run |
11 | GPT-5.6 Terra | Codex 0.154.0 | max | 23.32 | 23.23 | 8.18 | 17.24% ± 5.84 | 99/100 | 0/100 | $215.12 | View run |
Harbor snapshot: September 21, 2026. Costs are reported for the full cohort and depend on the model and agent configuration. V3 scores are separate from V1 and V2.
Grok 4.7: Harbor reports metrics over 100 tasks and currently links 96 trials. Review the source run.
Read the V3 release report →