MARB studies

Each entry records its task, kit, driver, prompt or note, tool path, run count, scorer, and date. Shared conditions improve comparison, while seed count, model size, modality, tool interface, timeout behavior, and protocol changes remain visible confounders.

A note on terms. Artifact-completion metrics record whether a run produced a loadable STEP and how many part instances it placed. “Loadable” does not mean physically buildable. The separate grade reports position, orientation, and answer-key gap metrics under the named scorer version.

StudyDriverDateRunsHeadlineStatus
Sighted local cells qwen3-vl 32B + Nemotron 3 Nano Omni (vision, goal image in-loop) 2026-06-12 15 Descriptive cross-model comparison: the sighted 32B cohort placed fewer parts and recorded higher GAP than the blind 80B cohort; size and modality are confounded. Nemotron Omni: 0/5 loadable exports. Published
Fable 5 effort sweep Claude Fable 5 (CadQuery) 2026-06-11 4 Four single-run cells: three effort settings plus a materially different multi-agent harness. Results are descriptive until repeated seeds are available. Published
Local open-weight anchor qwen3-coder-next (80B, on one local box) 2026-05-30 30 Prompt-cohort study with extension seeds. Loadable-export rates and positional grades varied; the resulting artifacts remained far from the reference. Published
First results, frontier models Claude (Fusion, CadQuery), OpenAI Codex (CadQuery) 2026-05-26 3 Each founding run used authored parts and four goal renders. None met the declared geometric threshold; see the dated scoring method. Published

Study: the Claude Fable 5 effort sweep

Driver: Claude Fable 5, reasoning effort low / medium / high, plus an ultra run (multi-agent: build, 4-agent adversarial audit, targeted fixes). CAD tool: CadQuery. Kit and grader identical to the first results. Blind runs, one per setting. Positional grades (MARB v0.9).

Three single-agent cells varied a recorded reasoning-effort setting under the same task and kit. A fourth “ultra” cell also changed the harness by adding multiple audit agents, so it is not an effort-only comparison. The question is what these four recorded workflows produced—not what more spending generally buys.

EffortGAP medianORIENT alignedPOS relative medianWall-clockBilled tokens
Ultra (multi-agent)3.0 mm47%30.4 mm83 min23.0M
Medium6.5 mm59%48.5 mm39 min4.7M
Low7.0 mm53%68.0 mm38 min4.2M
High7.0 mm49%38.1 mm45 min18.5M

Token figures were recovered post-run by summing recorded per-message usage blocks. In these n=1 cells, high effort recorded 3.9× the billed tokens of medium and higher GAP/lower ORIENT values. Host accounting and cache treatment limit cross-system cost comparisons.

Within these three single-agent runs, the metrics were not monotonic with the named effort setting: medium recorded the lowest GAP and highest ORIENT, high recorded the lowest relative POS, and low recorded the highest relative POS. One run per setting is insufficient for a causal or purchasing conclusion.

The ultra run used a materially different multi-agent harness that probed kit parts, ran declared audits, and applied targeted fixes. It recorded 3.0 mm GAP and 30.4 mm relative POS in 83 minutes with a recovered 23.0M-token bill. The audit reported specific shaft, belt, and pulley conditions for review. This single cell cannot isolate whether the harness, stochastic variation, or another condition produced the metric differences.

On the dated board, these cells occupy different positions by the primary GAP metric. Because each frontier cell is a single run and the ultra harness changes more than one factor, treat the differences as hypotheses for repeated testing. The recap retains the token-accounting limitations.

Study: the sighted local cells

Drivers: qwen3-vl:32b (vision, 8 turns n=5 + 12 turns n=2 preliminary) and nemotron3:33b Omni (vision, n=5), goal image inlined on turn 1, lean-v5 guidance. Same kit, grader, and harness as every other cell. Token capture native. Graded MARB v0.9.

These cells supplied a goal image in-loop to vision-capable models. The recorded sighted cohorts did not outperform the selected blind 80B text cohort on the reported metrics, but model size, modality, run budget, and other factors differ; this is not an image-on/image-off ablation.

CellLoadable exportParts placed (median)GAP medianORIENTPOS rel.
Blind text 80B (qwen3-coder-next, selected 14-turn cohort, n=9)9/10~62272 ± 149 mm12%118 ± 47 mm
Sighted 32B (qwen3-vl, 8 turns, n=5)5/512–17873 ± 174 mm0%1005 ± 613 mm
Sighted 32B, 12 turns (n=2, preliminary)2/2 graded16–20335 ± 138 mm5%260 ± 249 mm
Sighted Nemotron 3 Nano Omni (n=5)0/5
Goal vs blind text build vs sighted vision build, one camera

Study: the local open-weight anchor

Driver: qwen3-coder-next:q4_K_M (80B total, 3B active), text only, on a single local AI supercomputer. CAD tool: CadQuery 2.7.0. Kit: v1.1. Blind run, no internet, no memory of past work. 40 runs in six cohorts plus extension seeds. Buildability metrics plus positional grades (MARB v0.9; graded cohorts now n = 9 and n = 8 loadable of 10 attempts each).

This study records a self-hosted 80B coding model on one local system under the same mechanical task. “Self-hosted” does not mean free: hardware, electricity, setup, operator time, and opportunity cost remain outside the token table.

The study began with thirty runs across six five-run prompt cohorts and later added extension seeds, for forty recorded attempts in the stated aggregate. Notes clarified the tool or task without intentionally revealing the gated reference design; this is a protocol statement, not proof against all contamination.

CohortNote givenTurn budgetLoadable STEPParts placed (median)
ANone (control)81 of 515
BCorrect CAD export idiom85 of 598
CB, plus a build-volume clarification83 of 528
DB, plus design-goal requests82 of 524
ESame as D144 of 530
FLean note, the export idiom sharpened85 of 584

Target is about 100 placed part instances. The kit holds authored STEP parts; the model places each one and exports a single STEP file.

What the recursion taught us

The grades: placed, not jointed

Updated 2026-06-12 with five extension seeds per selected cohort: the initial five-of-five loadable-export counts became 9/10 (mechanics v2) and 8/10 (lean v5), and the reported spreads widened. The aggregates below use the loadable seeds.

The two cohorts were graded with the same GAP, POS, and ORIENT definitions used for the mechanical frontier track. Metric definitions are shared, but model, prompt, kit version, interface, seed count, and execution conditions must remain part of any comparison.

Cohort (n = 5)GAP medianORIENT alignedPOS relative median
Mechanics v2 (n=9 of 10)272 ± 149 mm12%118 ± 47 mm
Lean v5 (n=8 of 10)341 ± 133 mm20%233 ± 139 mm
Frontier range, for scale0.0 to 7.8 mm47 to 69%38 to 68 mm
Three-panel figure: the goal machine next to two selected local artifacts, which place a broadly similar part inventory without reproducing the reference assembly
The goal next to two selected local artifacts. The recorded inventory and scale are broadly similar; the submitted assembly relationships differ from the reference.

What this study does not yet answer

Method, harness, grades, and the run catalogue are in the open-source repository. These mechanical-task cohorts use the stated MARB v0.9 metrics; the separate Pascal lane uses its own grader and is not pooled with them.