MARB HandBench: Some Assembly Required

Testing whether AI can put the pieces together

Sunnyday Technologies · September 19, 2026 · Initial results, unranked

Conceptual illustration: separate mechanical parts become a robotic hand, with a future simulation loop. This is not a benchmark return or an FEA result.

One Grok Bot run returned all 226 required parts of a robot hand. Only one matched the target placement: the part used to align the comparison.

Astra Medium matched 12 placements; Extra high matched 20. The scores are low, which makes this a useful starting point. Sunnyday Technologies' MARB HandBench establishes a baseline for measuring improvement after model and tool updates.

Why a hand?

Holding a mug, picking up a washer and turning a key require different movements. A robot hand must produce them through joints, links, tendons or flexible elements packed into a small space.

Schematic movements: bending at a joint, spreading fingers, and flexing a spring-like element.

Small details matter. Pins must line up with holes. Tight fits can restrict movement; loose fits can introduce play. Manufacturing variations can accumulate across several joints. Cables need clearance, and worn components need to be accessible for replacement. Some contact is intentional: a flexible latch may bend against another part and spring back.

Why the economics matter

Production ambitions make those details worth studying. Hyundai announced plans in January 2026 for capacity of 30,000 robots a year by 2028. That is a target, not delivered production.

A hypothetical program producing 100,000 two-handed robots would need 200,000 hands. Separately, a fleet with 200,000 hands would need 20,000 replacements annually if 10% were replaced each year. That replacement rate is an assumption; actual servicing may replace only a finger, tendon or sensor.

Illustrative quantities, not forecasts: 100,000 new robots require 200,000 hands. A separate 200,000-hand fleet needs 10,000, 20,000 or 40,000 replacements at assumed annual rates of 5%, 10% or 20%.

A single design can reach many manufactured units. Better placement tools could speed design review, compare variants and expose problems before production. Cost savings were not measured.

The first results

The initial case uses Pollen Robotics' four-fingered Amazing Hand. The kit contains 29 component designs used 226 times. Models must position those parts in CAD without altering their shapes or seeing the completed reference assembly.

The target budget was 30 minutes and 100 calls to execution tools; Grok's limits were operator-supervised. Time, tool-call and reasoning budgets were deliberately constrained to limit token use. Astra ran at Low (Lite), Medium and Extra high; Ultra was not tested. These results describe performance within those budgets. No higher-budget comparison establishes how much the limits affected the scores.

The score counts placements matching one reference pose within a nominal 0.1 mm threshold. The known-correct control matched 226/226.

Six attempts, submitted/matched: Astra Low unavailable/ungraded; Medium 98/12; Extra high 226/20; Grok 01 1/1; Grok 02 226/1; Grok 03 139/1. The denominator is 226 for every attempt.

Across all three Grok runs, the sole match was the alignment reference. No other placement matched the reference. Astra Low remains ungraded because its execution evidence was incomplete. Every attempt is retained; none was replaced.

These are baseline observations, not a model ranking. All three Astra attempts timed out, their run-control software differed, and Medium retains a tool-call accounting qualification. Grok's underlying model version was unavailable and its activity records incomplete. A different, mechanically plausible hand pose can also miss this snapshot's target.

The score measures geometry and placement. STEP export checks are separate; future cases can use any format that preserves the required information.

From placement to optimization

Reliable placement could support an automated engineering loop: adjust an allowed position, check movement and clearance, simulate the loads, and compare the result.

Finite element analysis (FEA) divides a digital part into small elements to estimate how it bends and carries force. With suitable materials, loads and contact definitions, it could help compare joint settings or mounting positions while preserving design constraints.

Current work: measure placement. Proposed next stages: check motion and fit, define materials and loads, run FEA, adjust allowed parameters and repeat. Engineering review and physical tests remain necessary.

This loop is future work. The current pilot measures reconstruction, not optimized performance. Later candidates would face the same tasks and loading conditions, followed by engineering review and physical validation.

Keep the task and scoring fixed, and the next comparison can answer a useful question: how much better did the systems get?

The study began with an additional usage-reset credit from OpenAI. Building the benchmark consumed the renewed weekly allowance in roughly a day. That was a substantial investment in time and compute. The aim is to make that work useful to others.