MARB — Mechanical Assembly Readiness Benchmark
How close can an AI workflow come to a declared assembly target? MARB compares exported artifacts under a versioned geometric method: the same recorded task and grader, with GAP, ORIENT, POS, and configured gate results reported separately. It measures the published run—not physical buildability, safety, or certified readiness.
Benchmark, grades, and write-ups are open (papers & data). White paper review copies: [email protected]
The gap
AI-CAD benchmarking includes geometry-generation, part-classification, joint-prediction, and other assembly-oriented methods. MARB focuses on one narrower question: how a submitted multi-part artifact compares with a declared answer key and gate configuration.
The method reports inventory, interference, floating-part, GAP, ORIENT, and POS evidence where the applicable scorer version supports them. It does not prove that a machine is functional, safe, manufacturable, physically buildable, or ready for use.
The ladder
The maintainers use an L0–L7 research roadmap to separate increasingly broad capability claims. A rung can be evaluated only after its task, evidence, and acceptance rules are defined; the ladder is not a certification scheme.
L0 · Component
One part is exactly as specified.
L1 · Assemble the kit Today
Parts placed, aligned, no collisions, nothing floating. The Model T.
L2 · Constraint-robust
Re-solves correctly when parameters change.
L3 · Mechanically valid
Full-travel kinematics + load-holding, measured.
L4 · Engineering change
Re-design to a new requirement, no regressions.
L5 · Design from intent
Pick & arrange parts from functional goals. Tesla.
L6 · Optimization research
Candidate multi-physics comparison against a declared human baseline.
L7 · Closed-loop research goal
Proposed design, build, measurement, review, and improvement tasks with human authority and safety controls.
The published mechanical cohort evaluates a limited L1-style kit-assembly task. It does not establish the frontier capability of all systems, and later rungs remain proposals until their methods and evidence are published.
Capability × Readiness
The maintainers provide an indicative comparison between the research ladder and TRL, MRL, and IRL concepts. It is a discussion aid—not an equivalence, official mapping, agency assessment, or shortcut to a readiness review.
CADCLAW capability ladder ↔ TRL / MRL / IRL, indicative evidence bands, not equivalence.
Two different axes. Formal readiness processes evaluate a technology or system under their own authorities and evidence requirements. MARB evaluates a recorded workflow and exported artifact under a benchmark method. CADCLAW findings may be one input to an engineering review, but they are not readiness evidence by themselves.
How the benchmark runs
Within a declared cohort, MARB records the task, kit, prompt, tool path, model, assistance conditions, software versions, attempts, and grader. Shared inputs and a shared scorer improve comparability; differences in tools, hints, interfaces, seeds, and execution conditions remain potential confounders.
The MARB pipeline, inputs → AI driver → exported STEP → CADCLAW gates → MARB score + readiness → human review, with a read-fix-rerun loop.
The metrics
Each compatible record reports two categories separately: declared artifact metrics under the named scorer version and recorded effort fields such as elapsed time, attempts, and available usage. Neither category proves physical quality or total cost.
Configured artifact checks
MARB v0.9 reports inventory, interference, floating-part, GAP, ORIENT, and POS evidence where applicable. Each check has a defined digital scope and known omissions.
Mechanical sub-grade · 0–100
A method-specific combination of declared placement and interference components. It is a benchmark score, not a physical-buildability, manufacturing, safety, or structural-performance factor.
L0–L7 research roadmap
A maintainer-authored capability map for framing future evidence. It is not a scalar readiness score and is not equivalent to an official TRL, MRL, IRL, certification, or gate review.
Effort fields
Wall-clock, available usage, attempts, retries, corrections, and recorded human interventions are reported separately. Host meters and missing fields can differ, so they are not a normalized cost ranking.
Dated digital reference task
The reference is a resolver-built answer key for a two-metre 3D-printer / CNC frame of roughly 100 parts. The named MARB v0.9 metrics compare submitted artifacts with that reference; it is not an AI result or a physical-validation record.
- The answer key records zero findings under the configured inventory, interference, and floating-part gates. That is a property of this reference and scorer configuration, not an AI result or proof of physical quality.
- Consistently graded by the publisher: the same versioned automated method is applied to the recorded artifacts; this is not third-party independent testing.
- Dated cohort result: one Claude · CadQuery run recorded 0.0 mm median GAP error while still failing other geometric conditions. One metric is not physical buildability.
The board, 2026-06-11
This is the site snapshot dated 2026-06-11: eleven summarized cells for the mechanical task, ranked by the published primary metric. The repository run registry may contain later or additional records, and the Hugging Face Space is a separate publication surface that must be checked independently for freshness.
Published 2026-06-11 snapshot. Ranks and descriptive labels apply only to this image, scoring version, included cells, and recorded conditions.
Dated token-versus-GAP view for cells with available or recovered usage. Token accounting differs by host and is incomplete; see the method and limitations in the recap.
Deep dives: the recap article (findings to date + the token ledger), the studies log (the Fable 5 effort sweep and the graded local open-weight anchor), and the first-results write-up (the original three-way head-to-head, preserved below as the founding study), and the new MARB-A architecture lane (three models build a house in Pascal; two fall into the same units trap).
First head-to-head (the founding study, 2026-05-26)
Three publisher-run workflows—Claude Opus 4.7 with Fusion and CadQuery, and OpenAI Codex / GPT-5 with CadQuery—used fresh prompt-only sessions and the same founding kit. Each placed the supplied parts without human placement assistance; none met the benchmark's declared geometric threshold. These are single workflow runs, not independent product tests.
One of four goal renders provided (the 3/4 overview); the kit also includes front, top, and side views.
Recorded GAP versus wall-clock time for the included frontier cells. The chart is descriptive; it does not establish that model, iteration count, or reasoning effort caused the observed differences.
| # | Model · tool | GAP median ↓ | ORIENT aligned ↑ | POS rel median ↓ | Time | Cost est. |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.7 · CadQuery | 0.0 mm | 51% | 49.9 mm | 48.8 min | ~$68 |
| 2 | Claude Opus 4.7 · Fusion | 2.0 mm | 47% | 47.7 mm | 33.7 min | ~$174 |
| 3 | GPT‑5 Codex · CadQuery | 7.8 mm | 69% | 47.2 mm | 13.0 min | not reported |
| n/a | CADCLAW reference (answer key) | 0.0 mm | 100% | 0.0 mm | resolver | n/a |
MARB v0.9. GAP median = error vs the answer key's intended interface gap (≈0 mm bolted, ≈1 mm motion clearance). ORIENT aligned = % of asymmetric parts in the correct rotation. POS rel median = position error relative to each part's neighbours (frame-invariant). The configured bands (≤5 mm = located, ≤5° = aligned) are benchmark definitions, not machining tolerances or physical-buildability criteria. Cost estimates use the price basis recorded for the dated study; Codex CLI did not report comparable usage. Results 2026‑05‑26 · CadQuery 2.7.0, Autodesk Fusion 2702.1.58 · graded per the MARB v0.9 scoring spec.
Every flagged clip is structural, beams overlapping at splice joints and post/frame junctions, and the centered 2040 inserts overlapping their beams. Both also placed the Z-posts in the wrong rotation versus the reference, an orientation error the current gates don't yet catch, and the next gate we're adding.
Claude-Fusion, building the frame (10 → 100 parts). The model emits no parametric timeline, so we recovered the build order afterward by driving the live Fusion model through its MCP, revealing the placed parts in order under a fixed isometric camera.
Claude-CadQuery's own in-process renders, the orthographic + isometric checks it generated as it built. This is the human-reviewable output that lets a watcher catch a bad run early and stop it to save tokens.
What we asked, and what we didn't
We gave the goal, not the method. The driver got the target, the pictured assembly plus design constraints, and the kit, but not the build sequence. The original human-guided build specified an inside-out order (X axis → Y → Z-posts) and detailed steps; here we deliberately withheld that and let the AI decide how to reach the pictured result.
MARB measures the submitted artifact under the declared task and metrics. It does not directly measure a model's internal understanding. A submission with present but overlapping or mis-oriented parts demonstrates those artifact conditions, no more.
Human-reviewable throughout. Both drivers emitted orthographic + isometric renders as they built, so a person could watch the assembly take shape and abort a bad run early to save tokens. (Fusion didn't persist its in-session views, so we recovered the progression afterward straight from the live model via its MCP, see above.)
Descriptive founding-cohort spread. The three runs took 13–49 minutes. The most iterative workflow recorded the lowest GAP value and the longest time; the one-attempt workflow recorded the shortest time and a higher GAP value. With one run per workflow and different tool paths, MARB does not infer that self-review caused either result. Later local and effort studies are reported separately in the studies log and recap.
Why it matters, in plain English
Repeatable geometric checks can help teams find selected conditions earlier and compare workflow changes under a controlled method. The economic value depends on the assembly, review process, defect class, compute cost, and what the configured checks can actually observe.
CADCLAW can report selected properties of a supplied STEP artifact against declared rules. It cannot prove an assembly is right, prevent an escaped defect, quantify savings, or replace native-CAD review, engineering analysis, inspection, or physical testing.
For engineers
A versioned way to evaluate configured geometric checks on exported artifacts and inspect what each gate did not cover.
For programs
A reproducible benchmark record that may support research review; it is not a substitute for formal readiness evidence.
What makes it different
MARB is not "a CAD checker." Specifics separate it from CAD-vendor tooling and from prior benchmarks:
- Grades the submitted STEP, not a vendor's feature tree. A shared exported-artifact method reduces some vendor-specific differences while retaining export, tool-path, and metadata limitations.
- Scores one submitted multi-part artifact. The configured checks operate across the supplied assembly, while some public joint and mate datasets evaluate part pairs. Coverage still depends on the declared gates and task.
- Separates benchmark capability from readiness. An indicative crosswalk is published for discussion, but MARB scores do not enter or replace an official gate review.
- Uses a versioned reference assembly. The mechanical task is based on the M3-CRETE experimental reference design; the benchmark does not claim physical conformance or production validation.
- Publishes claim checks and limits. CADCLAW includes configurable text-audit tooling, which is a backstop rather than a guarantee of truthful or complete claims.
Papers and public benchmark data
The public result records link to the repository's scoring spec, graders, public kits, grades, and run registry. The registry lists available per-run provenance fields such as model, tool, timing, usage, and attempts; fields can be unavailable, and answer keys remain separately gated.
| Artifact | What it is | Where |
|---|---|---|
| The Angry Millimeter | Recap article: findings to date + the token ledger | marb.cadclaw.io/recap |
| MARB-A: the Pascal lane | Three-run pilot in the open-source Pascal editor under a separate method | marb.cadclaw.io/pascal |
| Studies log | Published study records: effort sweep, local open-weight anchor, and founding results | marb.cadclaw.io/studies |
| Scoring spec (v0.9) | The versioned method used for the published mechanical-track grades | MARB/spec |
| Frontier track write-ups | Claude tracks comparison + prompt-framework findings | comparison · findings |
| Local-anchor study | 30 runs, six prompt cohorts, graded — the full write-up | MARB/results |
| Grades & registry | Graded metrics and the provenance fields available for each registered run | grades · registry |
| Blind kits & graders | Run your own model against the board (kits, harness, metrics) | MARB repo |
| CADCLAW engine | The open engine used for the published mechanical-track grading | CADCLAW repo |
| Hugging Face mirror | Benchmark dataset, gated answer key, and the live leaderboard Space | dataset · answer key · leaderboard |
Request a sponsored benchmark run
Sunnyday can scope a new model/tool cell, repeat-seed cohort, or private evaluation as a manual service. A sponsored public run is identified in the run registry, follows the versioned method, and is published regardless of whether the result is favorable. Sponsorship does not buy a score, rank, conclusion, endorsement, certification, or claim about the sponsor's product.
Frontier run · planning range $100+
One scoped scripted-CAD run under a named kit and grader. Final price and deliverables depend on model access, compute, and rights.
GUI/tool run · planning range $250+
One scoped run through an approved tool interface. Availability, software terms, operator controls, and artifact rights must be confirmed.
Seed cohort · planning range $500+
A proposed repeat-run cohort under one frozen cell definition. Sample size and statistical treatment are agreed before work starts.
Model sweep · custom scope
A multi-cell study with a written protocol, disclosure terms, deliverables, review window, and publication decision defined in advance.
Planning ranges are non-binding and exclude applicable taxes, third-party licenses, unusual compute, travel, and custom engineering. This site does not accept payment or create an order. Work begins only after entity, tax, rights, privacy, independence, cancellation/refund, payment, and deliverable terms are confirmed in a signed order or invoice. Sponsors should obtain their own legal, accounting, and tax advice.
Read & review
We are publishing a maintainer-authored draft method and an indicative readiness crosswalk for review. MARB is not a ratified consensus standard, accredited test, or official readiness framework.
Reviewers and collaborators welcome, labs, CAD vendors, and standards bodies especially.