The Angry Millimeter — a dated token and artifact-metric snapshot

MARB v0.9 recap, updated 2026-06-12. One task (the M3-CRETE printer frame, ~100 authored parts), one blind kit, one automated grader (CADCLAW). Drivers: Claude Opus 4.7, Claude Fable 5 (four effort settings), GPT-5 Codex, qwen3-coder-next 80B (local text, two cohorts now at ten seeds each), and the sighted local cells (qwen3-vl 32B with the goal image in-loop; Nemotron 3 Nano Omni).

Archive note. This page preserves the 2026-06-12 board and token-accounting snapshot. The repository's current scoring specification is v0.12; historical run labels and values remain unchanged.

This page preserves the MARB board snapshot updated 2026-06-12 and the token values available under that study's accounting method. Frontier cells are single runs, local cohorts have the sample sizes shown, tool paths differ, and several hosts expose different or no usage fields. The chart supports descriptive comparisons within the recorded conditions; it is not a causal efficiency study, price forecast, or purchasing recommendation.

Dated scatter chart of recorded billed tokens versus GAP median for six frontier workflow cells; usage accounting differs by host and the chart is not a normalized efficiency ranking
Dated billed-token versus GAP view for cells with available or recovered usage. GAP is one answer-key metric; token accounting is host-specific and incomplete.

The task, in one paragraph

Mechanical-track cells record a blind-kit version, authored STEP parts, goal renders, task brief, tool path, model, assistance conditions, and scorer. CADCLAW reports GAP, ORIENT, and POS against the answer key under MARB v0.9. Kit versions, hints, interfaces, attempts, seeds, and token capture differ across cells and must remain part of the comparison.

The board

MARB v0.9 scoreboard: eleven AI cells ranked by GAP median, from Claude Opus 4.7 at 0.0 mm down to the sighted 32B vision model at 873 mm
The full board, ranked by GAP median. Lower is better for GAP and POS; higher is better for ORIENT.
#Model · toolEffortGAP medianORIENTPOS rel.Time
1Claude Opus 4.7 · CadQuerymax0.0 mm51%49.9 mm49 min
2Claude Opus 4.7 · Fusionmax2.0 mm47%47.7 mm34 min
3Claude Fable 5 · CadQueryultra (multi-agent)3.0 mm47%30.4 mm83 min
4Claude Opus 4.8 · Fusionkit v1.3 (hint)5.7 mm39%52.5 mm41 min
5Claude Fable 5 · CadQuerymedium6.5 mm59%48.5 mm39 min
6Claude Fable 5 · CadQuerylow7.0 mm53%68.0 mm38 min
7Claude Fable 5 · CadQueryhigh7.0 mm49%38.1 mm45 min
8GPT-5 Codex · CadQuerymax7.8 mm69%47.2 mm13 min
9Local qwen3-coder-next 80B (n=9)mechanics v2272 ± 149 mm12%118 ± 47 mm~4.5 min/run
10Local qwen3-coder-next 80B (n=8)lean v5341 ± 133 mm20%233 ± 139 mm~4.5 min/run
11Sighted qwen3-vl 32B (n=5)lean v5 + goal image873 ± 174 mm0%1005 ± 613 mm~30 min/run
·Reference (answer key)0.0 mm100%0.0 mm

None of the included artifacts met the benchmark's declared geometric threshold. That does not by itself determine physical buildability. The board is a dated collection of cells, not a fitted capability curve or universal ranking.

Findings to date

1. Presence and location separated in this task. The included frontier runs placed roughly the supplied part count while recording a spread in answer-key position and gap metrics. The result shows why inventory alone is insufficient; it does not establish that placement is “solved” generally or that human review cannot detect the same conditions.

2. Founding-cohort time/attempt association. One nine-attempt workflow recorded 0.0 mm median GAP in 49 minutes; one single-attempt workflow recorded 7.8 mm GAP and 69% ORIENT in 13 minutes. With n=1 and different tool paths, MARB does not infer a general iteration-versus-speed tradeoff or compare the products overall.

3. The three n=1 effort cells were not monotonic. Medium recorded the lowest GAP and highest ORIENT, while high recorded the lowest relative POS. Repeated seeds are required before treating this as an effort effect or cost recommendation.

4. The ultra cell changed the harness. That run added multiple declared audit agents and recorded 3.0 mm GAP and 30.4 mm relative POS. It also reported shaft, belt, and pulley conditions for review. A single run cannot isolate the harness effect from stochastic or other workflow differences.

5. Local prompt cohorts recorded different loadable-export rates. The cohort given the correct CadQuery export idiom produced more loadable files than the initial cohort; later guidance and turn-budget cohorts varied. Extension seeds reduced the initial five-of-five counts. These small cohorts do not establish a universal prompt law or physical buildability.

Three renders from one camera: the goal machine, the blind 80B text model's loose 110-solid build, and the sighted 32B vision model's near-empty 17-solid build
One descriptive render comparison: goal, selected blind-text artifact, and selected sighted artifact. Model size and modality differ.

6. Selected sighted and blind cohorts differed. The sighted 32B cohort produced five loadable exports with fewer placed parts and higher reported errors than the selected blind 80B cohort; a second sighted model produced no loadable export in five attempts. Size, modality, and protocol are confounded, so no causal image-token conclusion is supported.

7. Extension seeds changed the estimates. Initial five-of-five loadable-export counts became 9/10 and 8/10, and the reported spreads widened. This illustrates sampling uncertainty without implying that any particular initial cohort was misleading.

8. The local cohort provides a dated anchor. The recorded local artifacts had larger answer-key errors and more reported overlaps than the included frontier cells. Results apply to the named models, quantization, hardware, prompts, kits, and dates; they do not define a permanent “floor.”

The token ledger: recorded usage, with incompatible meters

The harness captured local usage under its host meter. Claude values were reconstructed by summing available per-message usage blocks in retained transcripts. The recorded Codex CLI session did not expose a comparable count. Cache accounting, token definitions, prices, and missing fields differ by host, so the table is an audit record rather than a normalized efficiency ranking.

RunBilled tokensOutput tokensCaptureAttemptsGAP result
Opus 4.7 · CadQuery22.1M321Krecovered from transcript90.0 mm
Opus 4.7 · Fusion34.2M1.15Mrecovered from transcript122.0 mm
Fable 5 · ultra23.0M614Krecovered from transcript83.0 mm
Fable 5 · high18.5M442Krecovered from transcript47.0 mm
Fable 5 · medium4.7M374Krecovered from transcript36.5 mm
Fable 5 · low4.2M330Krecovered from transcript37.0 mm
GPT-5 Codexnot exposed by host, in-session or on diskunavailable17.8 mm
Local 80B (per run)35–139K total (median ~75K)full, native1272–341 mm
Sighted 32B (per run)67–122K totalfull, native1873 mm

Within the incomplete, host-specific ledger, four descriptive comparisons are visible.

The Fusion and CadQuery Opus cells differed. The recorded Fusion cell had 55% more billed tokens and higher error values on the three displayed metrics. The study did not isolate the interface as the cause; attempts, interaction pattern, and tool maturity also differ.

The medium-effort cell recorded a lower billed-token count. It recorded 4.7M billed tokens, 6.5 mm GAP, and 59% ORIENT. One run and one host meter do not establish a “value” winner, expected future cost, or suitable purchasing choice.

The high-effort cell was non-monotonic in this run. It recorded 3.9× the billed tokens of medium with higher GAP and lower ORIENT, while recording lower relative POS. Repeated seeds are needed before estimating an effort or cost effect.

The ultra and Opus-CadQuery cells had similar recorded bills. They produced different metric profiles under different workflows. The comparison motivates a seeded harness experiment; it does not prove that audit tokens caused precision or that the recorded price is fair.

No cross-vendor efficiency winner is supported by this dataset. The useful operational finding is the measurement gap: MARB records unavailable token usage as unavailable and requires the capture method to travel with each result.

Honest limits

What's next

Planned work includes frontier seeds, a declared connectivity gate, controlled modality comparisons, and stronger token capture. The site pages are dated snapshots; consult the run registry and each scorer version rather than assuming this page is current.

Method, grades, run registry with the available provenance fields, and the CADCLAW engine: GitHub. Reviewers welcome — [email protected].