MARB research update · September 17, 2026 · Five exports

Astra vs Grok Bot: M3-CRETE assembly results

Astra placed its matched components closer to the reference. Grok Bot matched more of the reference assembly. Neither produced a verified, build-ready machine.

Astra's two Medium-effort runs recorded relative-position median errors of 20.6 and 20.3 mm. Grok Bot's three exports recorded 44.3, 53.9 and 79.3 mm. Astra's lower errors apply to smaller matched sets: 94 and 96 reference solids, compared with Grok's 101, 101 and 100.

The assembly task

The M3-CRETE M3-2 task asks an agent to assemble a roughly two-metre gantry from 14 authored STEP components and four reference images. The target has about 100 placed instances, represented by 101 reference solids. Required returns include an assembly STEP, editable source, review images and a run log. The printhead payload and final tool interface are out of scope.

Grok A01

Grok A01 submitted assembly

Grok A02

Grok A02 submitted assembly

Grok A03

Grok A03 submitted assembly

Astra A05 · Medium

Astra A05 submitted assembly

Astra A06 · Medium

Astra A06 submitted assembly
Submitted review images. Image pixels are unchanged; camera angles and scales differ. Select an image for the full-resolution view.

Results

All five exported STEP files were measured against the same reference with frozen scoring code. The review did not execute the submitted scripts. These are Sunnyday Technologies' measurements, not third-party testing.

Artifact metricGrok A01Grok A02Grok A03Astra A05 · MediumAstra A06 · Medium
Reference solids matched / 1011011011009496
Interfaces evaluated / 207207207203169191
Relative POS median error ↓44.3 mm53.9 mm79.3 mm20.6 mm20.3 mm
Absolute POS median error ↓84.2 mm94.7 mm88.5 mm70.9 mm64.3 mm
GAP median error ↓0.0 mm5.0 mm5.0 mm0.0 mm0.0 mm
Evaluated GAP errors ≤1 mm ↑55.1%39.1%42.4%72.8%75.9%
ORIENT aligned, asymmetric solids ↑60.8% (31/51)72.5% (37/51)46.0% (23/50)65.9% (29/44)67.4% (31/46)
Relative position error paired with reference-solid matching coverage for all five exports
Astra's lower position medians cover smaller matched subsets. Unmatched geometry is not automatically penalized in these medians.
Median GAP error paired with evaluated-interface coverage for all five exports
Interface coverage matters: the GAP medians cover 169–207 interfaces. Zero median error does not mean collision-free assembly.

Assembly outcomes

Astra: A06 matched two more reference solids and evaluated 22 more interfaces than A05, with a similar position median. Both exports include additional drive components and assumptions about extrusion lengths. Their agents reported unresolved fits, overlaps and mounting details. A05 leaves seven reference solids unmatched; A06 leaves five.

Grok: A02 improved orientation alignment over A01 but increased position and GAP errors. A03 lost ground on relative placement and orientation. Its absolute-position error improved over A02, while its GAP median stayed at 5.0 mm. The slightly higher share of GAP errors within 1 mm covers fewer interfaces and does not establish an overall improvement.

A03's source replaces the supplied 1,000 mm spreader with a generated 1,920 mm rectangular bar, not the authored V-slot profile. We have not established whether that substitution explains its unmatched solid. Wheel seating, belt routing and plate orientation remain approximate in the agent's account.

Test conditions and evidence

RunsDriver / effortReported CAD runtimeKey limitation
Grok A01–A03Grok Bot; underlying model and effort unavailableCadQuery 2.8.0 / Python 3.13.5Incomplete execution evidence; A02 reports delegation; A03 has inconsistent metadata
Astra A05GPT-6 Astra / MediumCadQuery 2.7.0 / Python 3.11.15Diagnostic/non-blind
Astra A06GPT-6 Astra / MediumCadQuery 2.6.1 / Python 3.11.9Diagnostic/non-blind

Astra's configured Medium effort is supported by local session records, not independent server-side verification. It received historical benchmark context. A06 also received an automated shell suggestion, but no placement guidance.

Grok received blind-kit instructions; isolation was not independently audited. A01's supplied conversation excerpts support CadQuery execution but are not authenticated host transcripts. A02's original assembly and rendering command records remain unavailable. Its reported executor delegation conflicts with the single-agent instructions.

A03's run log identifies the new attempt, but other fields retain earlier attempt labels. Its input-kit hash has not been linked to the retained handoff, and no shell transcript is included. Its return manifest verifies successfully. These limitations remain part of the result; original records were not rewritten.

CadQuery, Open CASCADE and setup

Grok Bot's returned scripts import, place and export parts with CadQuery. CadQuery uses Open CASCADE through OCP bindings: an Open CASCADE STEP header does not indicate a separate CAD workflow. CADCLAW was reported as the review-rendering tool, not the assembly exporter.

Grok reports creating its own Python environment and installing dependencies during each attempt. The successful Astra runs used a separate automated preparation phase. Both products used a runtime; the distinction is how setup was handled. This evidence does not establish a public Grok tool API or a newly introduced capability.

RunPreparationAssembly benchmarkReported total
Grok A01Included; not separatedNot separated6.63 min
Grok A02Included; not separatedNot separated11.40 min
Grok A03Included; not separatedNot separated7.57 min
Astra A055m 40s19m 55s25m 36s
Astra A067m 52s13m 36s21m 29s

Four earlier Astra attempts stopped during setup and produced no scored geometry. Totals include packaging; timing boundaries and rounding differ. Comparable token counts are unavailable. These records do not support a normalized speed or cost-efficiency ranking.

Comparison with earlier MARB results

The three zero GAP medians equal the leading numerical value on the June board. Astra's roughly 20 mm position medians are below that board's 30.4 mm Fable ultra result. These are numerical comparisons only: historical conditions differ, matching coverage varies, and the new exports do not form qualified repeat-run cohorts. The official leaderboard is unchanged.

Method and limits

Grading used a frozen CADCLAW source revision, CadQuery 2.7.0 and OCP 7.8.1.1.post1. A reference-against-itself control returned zero position and GAP errors and full orientation alignment. It checks the scoring setup, not equivalence with every historical environment.

Independent interference, floating-part and buildability checks were not rerun for this comparison. No physical validation is claimed. A reliable model comparison needs repeated runs with fixed model, effort, runtime, kit, prompt and scorer, reporting failures and variation alongside the best result.