MARB — Mechanical Assembly Readiness Benchmark

How close can an AI workflow come to a declared assembly target? MARB compares exported artifacts under a versioned geometric method: the same recorded task and grader, with GAP, ORIENT, POS, and configured gate results reported separately. It measures the published run—not physical buildability, safety, or certified readiness.

Benchmark, grades, and write-ups are open (papers & data). White paper review copies: [email protected]

Sunnyday Technologies
L0–L7Capability ladder
0.0 mmBest GAP median (Claude · CadQuery)
~54%Interfaces within 1 mm (best run)
~100Parts supplied in the founding task

The gap

AI-CAD benchmarking includes geometry-generation, part-classification, joint-prediction, and other assembly-oriented methods. MARB focuses on one narrower question: how a submitted multi-part artifact compares with a declared answer key and gate configuration.

The method reports inventory, interference, floating-part, GAP, ORIENT, and POS evidence where the applicable scorer version supports them. It does not prove that a machine is functional, safe, manufacturable, physically buildable, or ready for use.

Prior-art boundary. The earlier survey snapshot is dated 2026-05 and is not a priority or uniqueness claim. New benchmarks may exist; the maintained value is the published task, method, run registry, and reproducible comparison surface.

The ladder

The maintainers use an L0–L7 research roadmap to separate increasingly broad capability claims. A rung can be evaluated only after its task, evidence, and acceptance rules are defined; the ladder is not a certification scheme.

L0 · Component

One part is exactly as specified.

L1 · Assemble the kit Today

Parts placed, aligned, no collisions, nothing floating. The Model T.

L2 · Constraint-robust

Re-solves correctly when parameters change.

L3 · Mechanically valid

Full-travel kinematics + load-holding, measured.

L4 · Engineering change

Re-design to a new requirement, no regressions.

L5 · Design from intent

Pick & arrange parts from functional goals. Tesla.

L6 · Optimization research

Candidate multi-physics comparison against a declared human baseline.

L7 · Closed-loop research goal

Proposed design, build, measurement, review, and improvement tasks with human authority and safety controls.

The published mechanical cohort evaluates a limited L1-style kit-assembly task. It does not establish the frontier capability of all systems, and later rungs remain proposals until their methods and evidence are published.

Capability × Readiness

The maintainers provide an indicative comparison between the research ladder and TRL, MRL, and IRL concepts. It is a discussion aid—not an equivalence, official mapping, agency assessment, or shortcut to a readiness review.

The CADCLAW capability ladder (L0–L7) mapped to TRL, MRL, and IRL evidence bands

CADCLAW capability ladder ↔ TRL / MRL / IRL, indicative evidence bands, not equivalence.

Two different axes. Formal readiness processes evaluate a technology or system under their own authorities and evidence requirements. MARB evaluates a recorded workflow and exported artifact under a benchmark method. CADCLAW findings may be one input to an engineering review, but they are not readiness evidence by themselves.

Reliance limit. Neither MARB nor CADCLAW assigns TRL, MRL, IRL, compliance, certification, safety, or production status. Those conclusions require the applicable authority, broader evidence, qualified review, and physical validation.

How the benchmark runs

Within a declared cohort, MARB records the task, kit, prompt, tool path, model, assistance conditions, software versions, attempts, and grader. Shared inputs and a shared scorer improve comparability; differences in tools, hints, interfaces, seeds, and execution conditions remain potential confounders.

MARB pipeline: inputs to AI driver to exported STEP to CADCLAW gates to MARB score and readiness to human review

The MARB pipeline, inputs → AI driver → exported STEP → CADCLAW gates → MARB score + readiness → human review, with a read-fix-rerun loop.

The metrics

Each compatible record reports two categories separately: declared artifact metrics under the named scorer version and recorded effort fields such as elapsed time, attempts, and available usage. Neither category proves physical quality or total cost.

Configured artifact checks

MARB v0.9 reports inventory, interference, floating-part, GAP, ORIENT, and POS evidence where applicable. Each check has a defined digital scope and known omissions.

Mechanical sub-grade · 0–100

A method-specific combination of declared placement and interference components. It is a benchmark score, not a physical-buildability, manufacturing, safety, or structural-performance factor.

L0–L7 research roadmap

A maintainer-authored capability map for framing future evidence. It is not a scalar readiness score and is not equivalent to an official TRL, MRL, IRL, certification, or gate review.

Effort fields

Wall-clock, available usage, attempts, retries, corrections, and recorded human interventions are reported separately. Host meters and missing fields can differ, so they are not a normalized cost ranking.

Comparison boundary. A cell is a recorded workflow result, not a pure estimate of model capability. Cross-cell interpretation must account for kit version, authoring interface, prompt/hints, assistance, attempts, seed count, grader version, and missing token data.

Dated digital reference task

The reference is a resolver-built answer key for a two-metre 3D-printer / CNC frame of roughly 100 parts. The named MARB v0.9 metrics compare submitted artifacts with that reference; it is not an AI result or a physical-validation record.

The resolver-built answer key is the benchmark reference, not an AI result or a physical certification. The dated cohort below compares three recorded workflows under the published grader.

The board, 2026-06-11

This is the site snapshot dated 2026-06-11: eleven summarized cells for the mechanical task, ranked by the published primary metric. The repository run registry may contain later or additional records, and the Hugging Face Space is a separate publication surface that must be checked independently for freshness.

MARB v0.9 scoreboard: eleven AI builds ranked by GAP median, from Claude Opus 4.7 at 0.0 mm down to the sighted 32B vision model at 873 mm

Published 2026-06-11 snapshot. Ranks and descriptive labels apply only to this image, scoring version, included cells, and recorded conditions.

Dated scatter chart of recorded billed tokens versus GAP median for six frontier workflow cells; token accounting differs by host and is incomplete

Dated token-versus-GAP view for cells with available or recovered usage. Token accounting differs by host and is incomplete; see the method and limitations in the recap.

Deep dives: the recap article (findings to date + the token ledger), the studies log (the Fable 5 effort sweep and the graded local open-weight anchor), and the first-results write-up (the original three-way head-to-head, preserved below as the founding study), and the new MARB-A architecture lane (three models build a house in Pascal; two fall into the same units trap).

First head-to-head (the founding study, 2026-05-26)

Three publisher-run workflows—Claude Opus 4.7 with Fusion and CadQuery, and OpenAI Codex / GPT-5 with CadQuery—used fresh prompt-only sessions and the same founding kit. Each placed the supplied parts without human placement assistance; none met the benchmark's declared geometric threshold. These are single workflow runs, not independent product tests.

The target M3-CRETE machine; one of four reference renders each AI was given

One of four goal renders provided (the 3/4 overview); the kit also includes front, top, and side views.

GAP median versus build time for all seven frontier builds; Claude Opus 4.7 on CadQuery hit 0.0 mm in 49 minutes, OpenAI Codex was fastest at 13 minutes

Recorded GAP versus wall-clock time for the included frontier cells. The chart is descriptive; it does not establish that model, iteration count, or reasoning effort caused the observed differences.

#Model · toolGAP median ↓ORIENT aligned ↑POS rel median ↓TimeCost est.
1Claude Opus 4.7 · CadQuery0.0 mm51%49.9 mm48.8 min~$68
2Claude Opus 4.7 · Fusion2.0 mm47%47.7 mm33.7 min~$174
3GPT‑5 Codex · CadQuery7.8 mm69%47.2 mm13.0 minnot reported
n/aCADCLAW reference (answer key)0.0 mm100%0.0 mmresolvern/a

MARB v0.9. GAP median = error vs the answer key's intended interface gap (≈0 mm bolted, ≈1 mm motion clearance). ORIENT aligned = % of asymmetric parts in the correct rotation. POS rel median = position error relative to each part's neighbours (frame-invariant). The configured bands (≤5 mm = located, ≤5° = aligned) are benchmark definitions, not machining tolerances or physical-buildability criteria. Cost estimates use the price basis recorded for the dated study; Codex CLI did not report comparable usage. Results 2026‑05‑26 · CadQuery 2.7.0, Autodesk Fusion 2702.1.58 · graded per the MARB v0.9 scoring spec.

Every flagged clip is structural, beams overlapping at splice joints and post/frame junctions, and the centered 2040 inserts overlapping their beams. Both also placed the Z-posts in the wrong rotation versus the reference, an orientation error the current gates don't yet catch, and the next gate we're adding.

Interference is one declared geometric gate. Detected overlaps reduce the L1-style score under the published method. Passing that gate would not establish manufacturability, assembly sequence, tolerances, structural performance, safety, or physical buildability.
Fairness note. CADCLAW and its placement resolver were first built around CadQuery, before the Fusion connection (MCP) existed, so the CadQuery runs may carry a home-field edge in tooling maturity. We flag it so the comparison stays honest; an orientation gate and more reference tasks will tighten it.
Claude-Fusion build progression from 10 to 100 parts

Claude-Fusion, building the frame (10 → 100 parts). The model emits no parametric timeline, so we recovered the build order afterward by driving the live Fusion model through its MCP, revealing the placed parts in order under a fixed isometric camera.

Grid of in-process CAD review renders the Claude-CadQuery driver generated while assembling the machine

Claude-CadQuery's own in-process renders, the orthographic + isometric checks it generated as it built. This is the human-reviewable output that lets a watcher catch a bad run early and stop it to save tokens.

What we asked, and what we didn't

We gave the goal, not the method. The driver got the target, the pictured assembly plus design constraints, and the kit, but not the build sequence. The original human-guided build specified an inside-out order (X axis → Y → Z-posts) and detailed steps; here we deliberately withheld that and let the AI decide how to reach the pictured result.

MARB measures the submitted artifact under the declared task and metrics. It does not directly measure a model's internal understanding. A submission with present but overlapping or mis-oriented parts demonstrates those artifact conditions, no more.

Human-reviewable throughout. Both drivers emitted orthographic + isometric renders as they built, so a person could watch the assembly take shape and abort a bad run early to save tokens. (Fusion didn't persist its in-session views, so we recovered the progression afterward straight from the live model via its MCP, see above.)

Descriptive founding-cohort spread. The three runs took 13–49 minutes. The most iterative workflow recorded the lowest GAP value and the longest time; the one-attempt workflow recorded the shortest time and a higher GAP value. With one run per workflow and different tool paths, MARB does not infer that self-review caused either result. Later local and effort studies are reported separately in the studies log and recap.

Why it matters, in plain English

Repeatable geometric checks can help teams find selected conditions earlier and compare workflow changes under a controlled method. The economic value depends on the assembly, review process, defect class, compute cost, and what the configured checks can actually observe.

CADCLAW can report selected properties of a supplied STEP artifact against declared rules. It cannot prove an assembly is right, prevent an escaped defect, quantify savings, or replace native-CAD review, engineering analysis, inspection, or physical testing.

For engineers

A versioned way to evaluate configured geometric checks on exported artifacts and inspect what each gate did not cover.

For programs

A reproducible benchmark record that may support research review; it is not a substitute for formal readiness evidence.

What makes it different

MARB is not "a CAD checker." Specifics separate it from CAD-vendor tooling and from prior benchmarks:

Open method and open engine. MARB's published method and CADCLAW's MIT-licensed code can be inspected and challenged. Openness supports review; it does not make a score independently validated, unbiased, complete, or suitable for a safety or procurement decision. Qualified humans retain engineering and release authority.

Papers and public benchmark data

The public result records link to the repository's scoring spec, graders, public kits, grades, and run registry. The registry lists available per-run provenance fields such as model, tool, timing, usage, and attempts; fields can be unavailable, and answer keys remain separately gated.

ArtifactWhat it isWhere
The Angry MillimeterRecap article: findings to date + the token ledgermarb.cadclaw.io/recap
MARB-A: the Pascal laneThree-run pilot in the open-source Pascal editor under a separate methodmarb.cadclaw.io/pascal
Studies logPublished study records: effort sweep, local open-weight anchor, and founding resultsmarb.cadclaw.io/studies
Scoring spec (v0.9)The versioned method used for the published mechanical-track gradesMARB/spec
Frontier track write-upsClaude tracks comparison + prompt-framework findingscomparison · findings
Local-anchor study30 runs, six prompt cohorts, graded — the full write-upMARB/results
Grades & registryGraded metrics and the provenance fields available for each registered rungrades · registry
Blind kits & gradersRun your own model against the board (kits, harness, metrics)MARB repo
CADCLAW engineThe open engine used for the published mechanical-track gradingCADCLAW repo
Hugging Face mirrorBenchmark dataset, gated answer key, and the live leaderboard Spacedataset · answer key · leaderboard

Read & review

We are publishing a maintainer-authored draft method and an indicative readiness crosswalk for review. MARB is not a ratified consensus standard, accredited test, or official readiness framework.

Reviewers and collaborators welcome, labs, CAD vendors, and standards bodies especially.

This is a portfolio map, not a claim of automated integration, a shared production dataset, physical validation, certification, or an end-to-end commercial system. Each project states its own evidence and readiness boundaries.