August 12, 2026
•
Benchmark Update
August 2026 Run: Opus 5 Leads at 0.906, Grok 4.6 good balance of cost and performance
We added four new (agent, model) cells to Parametric CAD Bench, each scored on the same 100 tasks as the existing entries. All four land in the top four, pushing the board to 14 cells. Claude Opus 5 via Claude Code takes first place with a score of 0.906 — the first result to break 0.9 on this benchmark.
Grok 4.6 offers the best balance of cost and performance — delivering 96.9% of the top score at just 61% of the cost.
The new entries
Scores are 0–100 percentage points across 100 tasks per cell. Combined is the harmonic mean of geometry similarity and CAD/spec consistency, computed per task and then averaged.
| Rank | Model | Agent | Geom | Spec | Combined | Cost / task | Cost (total) |
|---|---|---|---|---|---|---|---|
| 1 | claude-opus-5(max) | claude-code 2.1.226 | 88.9 | 97.2 | 90.6 | $1.132 | $113.24 |
| 2 | grok-4.6(high) | grok-build 0.2.52 | 83.5 | 97.2 | 87.8 | $0.692 | $69.23 |
| 3 | gpt-5.6-sol(max) | codex 0.147.0 | 84.8 | 94.3 | 86.5 | $0.976 | $97.58 |
| 4 | grok-4.5(high) | grok-build 0.2.52 | 80.2 | 94.1 | 83.7 | $0.366 | $36.59 |
A new ceiling, and it costs less than the old one
Claude Opus 5 scores 0.906 against the previous best of 0.832 — a 7.4 point gain. Measured as remaining headroom rather than raw points, that closes 44% of the gap to a perfect score (0.168 → 0.094). It also does it for $113.24 against the old leader's $170.00, so the top of the board got both better and cheaper in the same run. That is not the usual pattern: through the May run, higher scores had tracked higher spend fairly closely.
Three months moved the engineering slice a long way
The two runs are 90 days apart — May 13 to August 11 — scored on the same 100 tasks with the same deterministic grader, so the envelope comparison is apples to apples even where per-model attribution isn't. In that single quarter the best combined score went from 0.832 to 0.906, and the board shifted further than that one number suggests: the weakest of the four new cells would have ranked first in May. Grok 4.5's 0.837 beats the score that led the benchmark three months earlier, for $36.59 against $170.00.
In May, exactly one cell out of ten cleared 0.80. This run added four more in a single batch, and their mean of 0.872 sits 9.7 points above the mean of May's top four (0.775). For a task family as unforgiving as parametric CAD — where the output has to rebuild without error, honor a written spec, and resolve to the right solid — that is a steep three-month slope.
Where the gain landed is the more interesting part. Spec consistency was already near saturation in May at 0.954 and improved 1.8 points to 0.972. Geometry similarity — the harder axis, and the one that had been lagging all along — moved 8.1 points, from 0.808 to 0.889, or 4.5× as far as spec. Frontier progress over this quarter went almost entirely into getting the shape right rather than into reading the specification.
Grok 4.6 is the best balance of cost and performance
Grok 4.6 via Grok Build takes second place at 0.878 for $69.23 — that is 96.9% of the leader's score for 61% of its cost. It is also the only new entry that beats another cell on both axes at once: against GPT-5.6-Sol it scores 1.3 points higher and costs $28.35 less. Across the full board it strictly dominates four cells — GPT-5.6-Sol, Gemini 3.1 Pro under mini-swe-agent, Claude Opus 4.7 under Claude Code, and Claude Sonnet 4.6 — each of which scores lower while costing more.
Plotted as score against spend, only three cells above 0.5 sit on the efficiency frontier at all: Claude Opus 5 if you want the maximum score, Grok 4.6 for the balance, and Grok 4.5 as the price floor. Grok 4.5 clears 0.837 for $36.59 against the old leader's $170.00 — a 4.6× saving for 0.5 points more — making it the cheapest route to a score above 0.8. The gap between the two Grok cells is a steep one: 4.6 costs 1.9× what 4.5 does and returns 4.1 additional points, so which of them is the right default depends entirely on whether those points are worth the money for your workload.
Geometry is still the binding constraint
Even after this quarter's 8.1-point gain, geometry is still what holds the top score down. Every cell on the board — all 14, across both runs — scores higher on spec consistency than on geometry similarity. At the top the gap is 8.3 points (0.889 geometry vs 0.972 spec), and two of the new entries reach 0.972 on spec. Models are now very good at producing CAD that agrees with the written specification; they are still measurably worse at producing the intended shape. Since the combined score is a harmonic mean, geometry is what caps the headline number, and closing the remaining 0.094 is largely a geometry problem.
Caveat: we haven't tested more model–agent combinations yet
In the May run, several models were evaluated under multiple agents, which is what exposed the harness effect — the same model scoring differently depending on the scaffold driving it. Each of the four new models ran under exactly one agent: Claude Opus 5 only under claude-code, both Grok models only under grok-build, and GPT-5.6-Sol only under codex.
So these four results cannot separate model capability from harness quality. The caution is concrete rather than theoretical: in May, Claude Opus 4.7 scored 0.734 under mini-swe-agent but only 0.655 under Claude Code — the same model, 7.9 points apart. Claude Opus 5's 0.906 was measured on Claude Code at a newer version (2.1.226 vs 2.1.140), so some unknown share of the jump belongs to the harness rather than the model. Read the four new numbers as model-plus-harness results, not as model rankings. We plan to fill in the missing combinations and publish them as they land.
The other gap is which model families are on the board at all. All 14 cells draw on just four — Claude, GPT, Gemini, and Grok. We are actively working to evaluate more model families, and welcome contributions of your own evaluation results.
View the full leaderboard →