The Angry Millimeter — a dated token and artifact-metric snapshot
This page preserves the MARB board snapshot updated 2026-06-12 and the token values available under that study's accounting method. Frontier cells are single runs, local cohorts have the sample sizes shown, tool paths differ, and several hosts expose different or no usage fields. The chart supports descriptive comparisons within the recorded conditions; it is not a causal efficiency study, price forecast, or purchasing recommendation.
The task, in one paragraph
Mechanical-track cells record a blind-kit version, authored STEP parts, goal renders, task brief, tool path, model, assistance conditions, and scorer. CADCLAW reports GAP, ORIENT, and POS against the answer key under MARB v0.9. Kit versions, hints, interfaces, attempts, seeds, and token capture differ across cells and must remain part of the comparison.
The board
| # | Model · tool | Effort | GAP median | ORIENT | POS rel. | Time |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.7 · CadQuery | max | 0.0 mm | 51% | 49.9 mm | 49 min |
| 2 | Claude Opus 4.7 · Fusion | max | 2.0 mm | 47% | 47.7 mm | 34 min |
| 3 | Claude Fable 5 · CadQuery | ultra (multi-agent) | 3.0 mm | 47% | 30.4 mm | 83 min |
| 4 | Claude Opus 4.8 · Fusion | kit v1.3 (hint) | 5.7 mm | 39% | 52.5 mm | 41 min |
| 5 | Claude Fable 5 · CadQuery | medium | 6.5 mm | 59% | 48.5 mm | 39 min |
| 6 | Claude Fable 5 · CadQuery | low | 7.0 mm | 53% | 68.0 mm | 38 min |
| 7 | Claude Fable 5 · CadQuery | high | 7.0 mm | 49% | 38.1 mm | 45 min |
| 8 | GPT-5 Codex · CadQuery | max | 7.8 mm | 69% | 47.2 mm | 13 min |
| 9 | Local qwen3-coder-next 80B (n=9) | mechanics v2 | 272 ± 149 mm | 12% | 118 ± 47 mm | ~4.5 min/run |
| 10 | Local qwen3-coder-next 80B (n=8) | lean v5 | 341 ± 133 mm | 20% | 233 ± 139 mm | ~4.5 min/run |
| 11 | Sighted qwen3-vl 32B (n=5) | lean v5 + goal image | 873 ± 174 mm | 0% | 1005 ± 613 mm | ~30 min/run |
| · | Reference (answer key) | 0.0 mm | 100% | 0.0 mm |
None of the included artifacts met the benchmark's declared geometric threshold. That does not by itself determine physical buildability. The board is a dated collection of cells, not a fitted capability curve or universal ranking.
Findings to date
1. Presence and location separated in this task. The included frontier runs placed roughly the supplied part count while recording a spread in answer-key position and gap metrics. The result shows why inventory alone is insufficient; it does not establish that placement is “solved” generally or that human review cannot detect the same conditions.
2. Founding-cohort time/attempt association. One nine-attempt workflow recorded 0.0 mm median GAP in 49 minutes; one single-attempt workflow recorded 7.8 mm GAP and 69% ORIENT in 13 minutes. With n=1 and different tool paths, MARB does not infer a general iteration-versus-speed tradeoff or compare the products overall.
3. The three n=1 effort cells were not monotonic. Medium recorded the lowest GAP and highest ORIENT, while high recorded the lowest relative POS. Repeated seeds are required before treating this as an effort effect or cost recommendation.
4. The ultra cell changed the harness. That run added multiple declared audit agents and recorded 3.0 mm GAP and 30.4 mm relative POS. It also reported shaft, belt, and pulley conditions for review. A single run cannot isolate the harness effect from stochastic or other workflow differences.
5. Local prompt cohorts recorded different loadable-export rates. The cohort given the correct CadQuery export idiom produced more loadable files than the initial cohort; later guidance and turn-budget cohorts varied. Extension seeds reduced the initial five-of-five counts. These small cohorts do not establish a universal prompt law or physical buildability.
6. Selected sighted and blind cohorts differed. The sighted 32B cohort produced five loadable exports with fewer placed parts and higher reported errors than the selected blind 80B cohort; a second sighted model produced no loadable export in five attempts. Size, modality, and protocol are confounded, so no causal image-token conclusion is supported.
7. Extension seeds changed the estimates. Initial five-of-five loadable-export counts became 9/10 and 8/10, and the reported spreads widened. This illustrates sampling uncertainty without implying that any particular initial cohort was misleading.
8. The local cohort provides a dated anchor. The recorded local artifacts had larger answer-key errors and more reported overlaps than the included frontier cells. Results apply to the named models, quantization, hardware, prompts, kits, and dates; they do not define a permanent “floor.”
The token ledger: recorded usage, with incompatible meters
The harness captured local usage under its host meter. Claude values were reconstructed by summing available per-message usage blocks in retained transcripts. The recorded Codex CLI session did not expose a comparable count. Cache accounting, token definitions, prices, and missing fields differ by host, so the table is an audit record rather than a normalized efficiency ranking.
| Run | Billed tokens | Output tokens | Capture | Attempts | GAP result |
|---|---|---|---|---|---|
| Opus 4.7 · CadQuery | 22.1M | 321K | recovered from transcript | 9 | 0.0 mm |
| Opus 4.7 · Fusion | 34.2M | 1.15M | recovered from transcript | 12 | 2.0 mm |
| Fable 5 · ultra | 23.0M | 614K | recovered from transcript | 8 | 3.0 mm |
| Fable 5 · high | 18.5M | 442K | recovered from transcript | 4 | 7.0 mm |
| Fable 5 · medium | 4.7M | 374K | recovered from transcript | 3 | 6.5 mm |
| Fable 5 · low | 4.2M | 330K | recovered from transcript | 3 | 7.0 mm |
| GPT-5 Codex | not exposed by host, in-session or on disk | unavailable | 1 | 7.8 mm | |
| Local 80B (per run) | 35–139K total (median ~75K) | full, native | 1 | 272–341 mm | |
| Sighted 32B (per run) | 67–122K total | full, native | 1 | 873 mm | |
Within the incomplete, host-specific ledger, four descriptive comparisons are visible.
The Fusion and CadQuery Opus cells differed. The recorded Fusion cell had 55% more billed tokens and higher error values on the three displayed metrics. The study did not isolate the interface as the cause; attempts, interaction pattern, and tool maturity also differ.
The medium-effort cell recorded a lower billed-token count. It recorded 4.7M billed tokens, 6.5 mm GAP, and 59% ORIENT. One run and one host meter do not establish a “value” winner, expected future cost, or suitable purchasing choice.
The high-effort cell was non-monotonic in this run. It recorded 3.9× the billed tokens of medium with higher GAP and lower ORIENT, while recording lower relative POS. Repeated seeds are needed before estimating an effort or cost effect.
The ultra and Opus-CadQuery cells had similar recorded bills. They produced different metric profiles under different workflows. The comparison motivates a seeded harness experiment; it does not prove that audit tokens caused precision or that the recorded price is fair.
No cross-vendor efficiency winner is supported by this dataset. The useful operational finding is the measurement gap: MARB records unavailable token usage as unavailable and requires the capture method to travel with each result.
Honest limits
- Every included frontier cell is a single run. Local cohort sizes vary and are shown with each result. Do not infer trends from the frontier deltas.
- Token capture and accounting differ by host, and at least one displayed frontier cell lacks a comparable count.
- The grader scores per-part pose and interface gaps. It does not yet score connectivity — whether parts join into one rigid structure. That gate (L1) is the next grader level, and the local study is exactly why: it is possible to score parts near their homes while building nothing a wrench could touch.
- None of the eleven summarized artifacts met the declared benchmark threshold. That threshold is not a physical buildability certification.
What's next
Planned work includes frontier seeds, a declared connectivity gate, controlled modality comparisons, and stronger token capture. The site pages are dated snapshots; consult the run registry and each scorer version rather than assuming this page is current.
Method, grades, run registry with the available provenance fields, and the CADCLAW engine: GitHub. Reviewers welcome — [email protected].