MARB studies
Each entry records its task, kit, driver, prompt or note, tool path, run count, scorer, and date. Shared conditions improve comparison, while seed count, model size, modality, tool interface, timeout behavior, and protocol changes remain visible confounders.
A note on terms. Artifact-completion metrics record whether a run produced a loadable STEP and how many part instances it placed. “Loadable” does not mean physically buildable. The separate grade reports position, orientation, and answer-key gap metrics under the named scorer version.
| Study | Driver | Date | Runs | Headline | Status |
|---|---|---|---|---|---|
| Sighted local cells | qwen3-vl 32B + Nemotron 3 Nano Omni (vision, goal image in-loop) | 2026-06-12 | 15 | Descriptive cross-model comparison: the sighted 32B cohort placed fewer parts and recorded higher GAP than the blind 80B cohort; size and modality are confounded. Nemotron Omni: 0/5 loadable exports. | Published |
| Fable 5 effort sweep | Claude Fable 5 (CadQuery) | 2026-06-11 | 4 | Four single-run cells: three effort settings plus a materially different multi-agent harness. Results are descriptive until repeated seeds are available. | Published |
| Local open-weight anchor | qwen3-coder-next (80B, on one local box) | 2026-05-30 | 30 | Prompt-cohort study with extension seeds. Loadable-export rates and positional grades varied; the resulting artifacts remained far from the reference. | Published |
| First results, frontier models | Claude (Fusion, CadQuery), OpenAI Codex (CadQuery) | 2026-05-26 | 3 | Each founding run used authored parts and four goal renders. None met the declared geometric threshold; see the dated scoring method. | Published |
Study: the Claude Fable 5 effort sweep
Three single-agent cells varied a recorded reasoning-effort setting under the same task and kit. A fourth “ultra” cell also changed the harness by adding multiple audit agents, so it is not an effort-only comparison. The question is what these four recorded workflows produced—not what more spending generally buys.
| Effort | GAP median | ORIENT aligned | POS relative median | Wall-clock | Billed tokens |
|---|---|---|---|---|---|
| Ultra (multi-agent) | 3.0 mm | 47% | 30.4 mm | 83 min | 23.0M |
| Medium | 6.5 mm | 59% | 48.5 mm | 39 min | 4.7M |
| Low | 7.0 mm | 53% | 68.0 mm | 38 min | 4.2M |
| High | 7.0 mm | 49% | 38.1 mm | 45 min | 18.5M |
Token figures were recovered post-run by summing recorded per-message usage blocks. In these n=1 cells, high effort recorded 3.9× the billed tokens of medium and higher GAP/lower ORIENT values. Host accounting and cache treatment limit cross-system cost comparisons.
Within these three single-agent runs, the metrics were not monotonic with the named effort setting: medium recorded the lowest GAP and highest ORIENT, high recorded the lowest relative POS, and low recorded the highest relative POS. One run per setting is insufficient for a causal or purchasing conclusion.
The ultra run used a materially different multi-agent harness that probed kit parts, ran declared audits, and applied targeted fixes. It recorded 3.0 mm GAP and 30.4 mm relative POS in 83 minutes with a recovered 23.0M-token bill. The audit reported specific shaft, belt, and pulley conditions for review. This single cell cannot isolate whether the harness, stochastic variation, or another condition produced the metric differences.
On the dated board, these cells occupy different positions by the primary GAP metric. Because each frontier cell is a single run and the ultra harness changes more than one factor, treat the differences as hypotheses for repeated testing. The recap retains the token-accounting limitations.
Study: the sighted local cells
These cells supplied a goal image in-loop to vision-capable models. The recorded sighted cohorts did not outperform the selected blind 80B text cohort on the reported metrics, but model size, modality, run budget, and other factors differ; this is not an image-on/image-off ablation.
| Cell | Loadable export | Parts placed (median) | GAP median | ORIENT | POS rel. |
|---|---|---|---|---|---|
| Blind text 80B (qwen3-coder-next, selected 14-turn cohort, n=9) | 9/10 | ~62 | 272 ± 149 mm | 12% | 118 ± 47 mm |
| Sighted 32B (qwen3-vl, 8 turns, n=5) | 5/5 | 12–17 | 873 ± 174 mm | 0% | 1005 ± 613 mm |
| Sighted 32B, 12 turns (n=2, preliminary) | 2/2 graded | 16–20 | 335 ± 138 mm | 5% | 260 ± 249 mm |
| Sighted Nemotron 3 Nano Omni (n=5) | 0/5 | — | — | — | — |
- Observed sighted 32B cohort. All five runs produced loadable exports with 12–17 placed parts and higher reported errors than the selected blind 80B cohort. The study does not identify the cause.
- Preliminary 12-turn variant. Two completed seeds recorded lower GAP than the eight-turn cohort while still placing fewer than 20 parts; three other attempts were lost to a documented host timeout. No general turn-budget conclusion is supported.
- Observed Nemotron cohort. Five recorded sessions produced no loadable export. This is an artifact-completion result under the stated harness, not a claim about physical buildability or the model generally.
- Caveat, stated plainly: this compares a 32B vision model against an 80B text model — different sizes — because those are the strongest local options per modality. The cell answers a shop's question ("does my best local vision model beat my best local text model here?"), not a controlled ablation. A same-model image-on/off ablation needs a frontier vision model and is queued.
Study: the local open-weight anchor
This study records a self-hosted 80B coding model on one local system under the same mechanical task. “Self-hosted” does not mean free: hardware, electricity, setup, operator time, and opportunity cost remain outside the token table.
The study began with thirty runs across six five-run prompt cohorts and later added extension seeds, for forty recorded attempts in the stated aggregate. Notes clarified the tool or task without intentionally revealing the gated reference design; this is a protocol statement, not proof against all contamination.
| Cohort | Note given | Turn budget | Loadable STEP | Parts placed (median) |
|---|---|---|---|---|
| A | None (control) | 8 | 1 of 5 | 15 |
| B | Correct CAD export idiom | 8 | 5 of 5 | 98 |
| C | B, plus a build-volume clarification | 8 | 3 of 5 | 28 |
| D | B, plus design-goal requests | 8 | 2 of 5 | 24 |
| E | Same as D | 14 | 4 of 5 | 30 |
| F | Lean note, the export idiom sharpened | 8 | 5 of 5 | 84 |
Target is about 100 placed part instances. The kit holds authored STEP parts; the model places each one and exports a single STEP file.
What the recursion taught us
- Export-call association. The cohort given the correct export idiom recorded more loadable exports than the initial five-run cohort. Small cohort sizes and stochastic variation prevent a causal estimate.
- Additional guidance cohorts. Cohorts with build-volume and design-goal notes recorded lower initial loadable-export rates. The study does not establish why.
- Longer-budget cohort. The 14-turn cohort recorded a higher initial export rate than the preceding design-goal cohort but did not establish improved physical or engineering quality.
- Lean-note cohort. The lean note matched the initial five-of-five loadable-export count; extension seeds later reduced the combined rates.
The grades: placed, not jointed
Updated 2026-06-12 with five extension seeds per selected cohort: the initial five-of-five loadable-export counts became 9/10 (mechanics v2) and 8/10 (lean v5), and the reported spreads widened. The aggregates below use the loadable seeds.
The two cohorts were graded with the same GAP, POS, and ORIENT definitions used for the mechanical frontier track. Metric definitions are shared, but model, prompt, kit version, interface, seed count, and execution conditions must remain part of any comparison.
| Cohort (n = 5) | GAP median | ORIENT aligned | POS relative median |
|---|---|---|---|
| Mechanics v2 (n=9 of 10) | 272 ± 149 mm | 12% | 118 ± 47 mm |
| Lean v5 (n=8 of 10) | 341 ± 133 mm | 20% | 233 ± 139 mm |
| Frontier range, for scale | 0.0 to 7.8 mm | 47 to 69% | 38 to 68 mm |
- Parts land 100 to 400 mm off on a 2000 mm machine. Joints that should sit flush are 300 to 400 mm apart at the median, and only 11 to 14% of interfaces fall within 1 mm. About 30% of asymmetric parts get the correct rotation.
- The shape is closer than the anchor. Relative position error is far smaller than absolute (104 mm vs roughly 500 mm for mechanics v2): the model captures local clustering better than it anchors the whole frame.
- Loadable export and positional grade are different outcomes. Lean v5 recorded a higher loadable-export rate while mechanics v2 recorded lower positional errors in this sample.
- The structural gates agree. Both graded builds fail inventory on the part mix (not the count) and show 20 to 28 pairs of clipping solids, while nothing floats free. Parts grouped together, not a connected frame; a connectivity gate that scores seated joints is the next grader level.
What this study does not yet answer
- Design judgement at scale. Asking a small model to weigh manufacturability, thermal load, vibration, fatigue, and kinematics is the real goal of the benchmark. It pays off only with a larger model, a longer budget, and a grader that can score those modes. That work belongs in the shared task definition, not in a local note.
Method, harness, grades, and the run catalogue are in the open-source repository. These mechanical-task cohorts use the stated MARB v0.9 metrics; the separate Pascal lane uses its own grader and is not pooled with them.