Back to CAD-Bench

Parametric CAD Bench V3 Leaderboard

Current

100 tasks: 30 CAD Creation from Text, 30 CAD Create And then Edit from Text, 40 CAD Create from Engineering Drawings. Overall score is the mean continuous task reward; failed or unscored trials count as zero. Scores are shown on a 0–100 scale.

RankModelAgentEffortCreateCreate + EditImage-to-CADOverall (95% CI)ScoredPerfectCost (USD)Run
1
GPT-6 AstraCodex
0.154.0
max52.2644.3769.7056.87% ± 6.85100/1005/100$317.81View run
2
Claude Fable 5.1Claude Code
2.1.270
max55.4836.7572.7156.75% ± 7.1898/1007/100$1,056.04View run
3
Claude Opus 5Claude Code
2.1.270
max49.4836.5262.6450.86% ± 7.4599/1006/100$1,012.21View run
4
Gemini 3.8 Flashmini-swe-agent
2.4.6
high47.9431.2231.1036.19% ± 6.9794/1004/100$271.85View run
5
Grok 4.7Grok Build
1.0.30
high45.5734.6326.9134.82% ± 6.64100/1003/100$670.22View run
6
GPT-5.6 SolCodex
0.154.0
max36.2324.7829.2029.98% ± 6.17100/1001/100$291.99View run
7
Grok 4.6Grok Build
1.0.30
high41.6133.4513.0027.72% ± 6.30100/1002/100$390.94View run
8
Gemini 3.8 FlashAntigravity
1.2.7
high43.2825.3217.4527.56% ± 6.72100/1004/100$196.74View run
9
Kimi K3mini-swe-agent
2.4.6
max39.1123.4915.3624.92% ± 6.4096/1001/100$314.55View run
10
Muse Spark 1.3mini-swe-agent
2.4.6
max33.6124.338.8420.92% ± 5.8199/1001/100$357.03View run
11
GPT-5.6 TerraCodex
0.154.0
max23.3223.238.1817.24% ± 5.8499/1000/100$215.12View run

Harbor snapshot: September 21, 2026. Costs are reported for the full cohort and depend on the model and agent configuration. V3 scores are separate from V1 and V2.

Grok 4.7: Harbor reports metrics over 100 tasks and currently links 96 trials. Review the source run.

Read the V3 release report →