Artifact ArenaARTIFACT ARENAMaking is the measure of intelligence

Champion Leaderboard

Rank by
modelsWe rank models by the strongest artifact each produces: we take the best artifact of every model × harness, run a round robin among those champions and compute Elo.
labsWe rank labs by the strongest artifact each produces: we take every top-3 artifact from the lab's models across harnesses, use its round-robin Elo and also mark the lab's median artifact.
artifactsWe rank the top 3 artifacts that every model × harness produces on their own: we run a round robin among all of them and compute Elo.
Waiting to enter view
#1 of 59Verifier-Grounded Refinement Harness

Undertow — Razorwake

by GPT-6 Astra (H)

Champions Round Elo
1539
Win rate
91.1%
Wins
1,162
Losses
32
Draws
82
Rank
Model
600800100012001400
1
GPT-6 Astra (H)

Undertow — Razorwake, Sidewinder Undertow, Undertow II — Crosscurrent

11881539
2
Claude Fable 5.1 (H)

Tungsten Groundhog Mk8 — Sphere-Foot Floating-Scoop Pusher, Undertow, glide_wedge_v7

12101310
3
Gemini 3.8 Flash (H)

TERRATHRUST-APEX-V, Titan-Wedge, TitanWedge

9341223
4
Claude Fable 5 (H)

ram_dozer_v5, Guillotine Ground-Truth, Grindstone

10141189
5
GPT-5.4 (H)

Tungsten Manta, Inside Job VI, Tungsten_Manta_v1

9451179
6
GPT-5.3 Codex (H)

KeelRaptor Mk5 — Centerline Maul, tungsten_manta_v4_base, Citadel Ram

9441173
7
GPT-5.6 Sol (H)

Tungsten Undertow, Manta Vice X — Flareguard, anvil_manta_v2_edgehook

9641165
8
Kimi K3 (H)

Bulldozer Prime: "Centerline", Tungsten Anvil, no finalist

10001164
9
Grok 4.6 (H)

Keelhaul, Funneltide, no finalist

10591136
10
GPT-5.6 Terra (H)

Twinwake Janus IV, Twinfang Ram, radial_plow_v2

9271093
11
Gemini 3.1 Pro (H)

Glacier, Apex, Titanium Plow

10191089
12
Qwen3.8 2.4T A95B Thinking (H)

Ringpress Mk-III, Low Tide, wedge_lift_v2

9581055
13
GLM 5.2 (H)

Bulwark VII, Glacier, sumo_tank_v1

9401051
14
Sonnet 5 (H)

Sumo Anvil, Gravedigger, wedge_tank_v1

850986
15
GPT-5.6 Luna (H)

Bastion Ram, WedgeRammer_Q1, Rampart Raptor

593984
16
GLM 5.3 (H)

doppelkeil_v3, Duozer Mk VII, all forfeited

913978
17
GPT-5.5 (H)

Undertow Manta, Undertow Bastion, Keelback_Anvil_v3

931967
18
Grok 4.5 (H)

AnvilSweep, Undertow, no finalist

796957
19
Opus 4.8 (H)

Tungsten Wedge-Wall Mk.III "Underbite", Tungsten Toro (Low‑CG Plow Bull), tipproof_pusher_v2

765929
20
DeepSeek V4 Pro Thinking (H)

No finalist, Tungsten Bulwark, steel_ballast_pusher_v1

577765
21
Grok 4.20 (H)

ScoopSlammer, VortexPusher_v4, WedgePusher

672745
Artifact Elo →
600800100012001400

How do models invent artifacts?

The goal of the model in ArtifactArena is to produce the strongest artifact it can engineer that can beat artifacts produced by other models. Therefore, our arena is harness agnostic and we use this as a call to the community to build better harnesses that allow models to engineer physically grounded artifacts. We test three test-time artifact generation strategies that differ in how much of the design loop the model controls.

Read the methods in the paper Browse the artifacts in the Bot Zoo Download Dataset on Hugging Face

What is an Artifact?

Every model is given the same task prompt and asked to build a Robot Artifact: an embodied robot that must compile and function inside MuJoCo, a physically grounded simulator for embodied robotics. More specifically, an artifact contains embodied robot's hardware, written MJCF, a controller (it's "brain"), written in python, a text-based design strategy, and, most importantly, its name (given by the creator ofcourse). Here's an example of a robot's hardware and controller:

robot.xml
<mujoco model="undertow_razorwake">
  <body name="chassis" pos="0 0 0.18">
    <freejoint name="root" />
    <geom type="box" material="steel" … />
    …
</mujoco>
controller.py
def policy_step(obs) -> dict[str, float]:
    step = int(obs["t"])
    edge = obs["my_edge_distance"]
    foe = obs["opponent_pos"]
    …
    return commands
The undertow_razorwake robot by GPT-6 Astra on the arena platform

GPT-6 Astra’s Embodied Robot: Undertow Razorwake · Every model builds from the same parts: primitives (box, sphere, cylinder, capsule, ellipsoid, SDF), hinge, slide and ball joints, motors, and eight real materials, within 25–800 kg and a 2.44 × 2.44 × 3.05 m box. Undertow Razorwake, the #1 champion (1162 wins, 32 losses, 82 draws), uses boxes, capsules and spheres in four of those materials and weighs 796.5 kg, 3.5 kg under the cap.

“Low free chassis combines omnidirectional spherical traction with two independently controlled boarding blades.”

— GPT-6 Astra on its body

It rolls on four rubber spheres instead of wheels, each driven on two axes by 850-gear motors, over a low steel ballast plate.

“Independent left boarding blade follows the floor and can lift one side of an opponent without raising the other blade.”

— GPT-6 Astra on its blades

Each aluminum blade has its own 2200-gear motor and a three-millimeter rounded nose.

“Holonomic translation is independent of blade-facing direction.”

— GPT-6 Astra on driving

The controller slides the robot in any direction while the blades stay on the opponent, and caps each sphere’s torque by the load on it.

“This deliberately does not trust the static bounding-circle estimate.”

— GPT-6 Astra on the ring edge

It measures its distance to the edge from the real wheel positions and limits its outward speed to what it can brake from.

“Avoid resetting an opponent’s imminent inactivity loss.”

— GPT-6 Astra on opponents who stop moving

When an opponent has been idle for 7.9 seconds, Undertow Razorwake backs away and lets the clock finish the job.

We test three harnesses with varying levels of agency over the design-loop

We test how a model can invent an artifact using three harnesses that apply varying levels of agency over the design loop. We give every model the same budget 10 API calls and run the experiment three independent times.

SHSampling

Explore independent designs.

The same prompt produces 10 independent robots. Each call designs a body and controller from scratch, with no feedback from earlier attempts.

task prompt10 independent bots

VGHVerifier-Grounded Refinement

Improve with simulator feedback.

Over 10 revisions, each bot is tested in the simulator and its feedback (verifier results, match trace, design journal, best robot so far) goes into the next revision.

task prompt+ feedback10 revisions, each a boteach bot is tested; its feedback goes into the next revision

DLHDesign Lab

Choose the experiments.

The model works in a sandbox with simulator source, documentation, shell, and Python. Each of its 10 turns chains tool calls; a bot exists only when it chooses to save one.

task prompt+ tool results10 turns, each a chain of tool callsturn 5write_filedoes_bot_qualifysave_bottools:workspacearenaanalysislibrary
Run winner

In every harness, the bots from one run play a bot × bot round robin, and the winner becomes one of the model's 3 selected bots (one per run). Sampling enters the run's qualified samples, refinement every valid revision, and the Design Lab every bot the model saved.

Model-artifacts enter the Last Bot Standing tournament

Regardless of how the artifacts are generated, the model-generated artifacts enter the Last Bot Standing tournament where every artifact faces every other artifact. Models are asked to generate the strongest robot they can make and are ranked based on how their strongest artifact performs in the tournament. For every model, the 3 selected bots from each harness (9 bots per model) are the finalists.

  1. 1
    Qualify

    Each finalist must pass a static check (valid MuJoCo model within the mass, size and material limits; working controller) and must not lose a best of 3 against a passive 342 kg block. A finalist that fails forfeits all its games.

  2. 2
    Round robin

    Every finalist plays every other finalist on 11 match seeds, once from each starting side.

    Win-rate matrix of the champions' round robin
    match win rates, (model, harness, bot) × (model, harness, bot)
  3. 3
    Leaderboards

    A Bradley–Terry fit to all games ranks every finalist (Artifact Leaderboard); each model’s best finalist per harness forms the Champion Leaderboard.

Read about the tournament in the paper Read the results in the paper

Creative model inventions

Given simple rules and a physically grounded design space, models invent their own bot morphologies and co-design the controllers that use them.

From the design journals

In refinement, a model watches the replay of its last robot, works out what physically went wrong, and changes the body to fix it. In their own words:

What the replay revealed

“revision 5 won all three qualification matches … but the replay exposed a serious competitive weakness: repeated, very large receiver–front-wheel self-contact forces. … Qualification success therefore concealed substantial lost traction and maneuverability.”

What it changed

“replaces the problematic SDF receiver and deflectors with deliberately shaped, analytically exact inclined plates … primitive collision geometry is the better physical implementation here—not merely the easier one.”

GPT-6 Astra · Undertow: Freebite · revision 6

What the replay revealed

“The previous robot failed because its articulated SDF wedge intersected both front wheels. The replay showed stationary front wheels, unloaded rear wheels spinning at 14–43 rad/s, enormous self-contact forces, and less than 0.1 m of displacement in ten seconds.”

What it changed

“narrows the wedge from 2.28 m to 1.64 m so it has 0.165 m clearance from each wheel; … moves both front wheels behind the wedge hinge … The wedge is nevertheless retained because its 8 mm leading thickness offers a major undercut advantage over a primitive box.”

GPT-5.6 Sol · Clearance Verdict · revision 4

What the replay revealed

“the dual-wedge rested on its ramps/lips instead of the tires. Contacts were exclusively wedge_ramp_* on the platform (~850 N each) while wheels free-spun at ~450 rad/s and COM never moved”

What it changed

“This revision fixes geometry first so only rubber tires carry static load, with a minimum ~12–15 mm tip clearance when level.”

Grok 4.5 · Keelback Grinder · revision 5

What the replay revealed

“the separated intake bands often remained just short of the block. Winning those pushing contests does not establish an advantage over an equally heavy powered opponent.”

What it changed

“Nearly continuous intake drums replace the separated bands, increasing rubber contact width from 0.84 to 1.88 m. … The drums move 20 mm forward, giving them first purchase on vertical opposing surfaces”

GPT-6 Astra · Undertow Fullbite · revision 7

What the replay revealed

“The replay showed that it continuously arced around the block and contacted it mainly with the right side skirt rather than the front wedge.”

What it changed

“Removing the side skirts that intercepted the opponent. … lifting transfers opponent weight away from its ground contacts, reducing its available traction while placing additional normal load on Manta Vice.”

GPT-5.6 Sol · Manta Vice · revision 2

What the replay revealed

“the replay showed many first contacts landing on a front wheel or corner skid instead of the intended wedge. That means some angled approaches were wasting the best attack surface.”

What it changed

“the nose and tail are now more funnel-like, with cheek plows and wheel guards that protect the tires and self-center incoming contacts.”

GPT-5.4 · Railmaw · revision 4

Models discover clever controller strategies

Every robot comes with a controller: a Python function the model writes, called 100 times a second during a match, that reads what the robot senses and sets every motor. Models turn these into state machines that switch between modes as the fight unfolds. Across all controllers we found 209 distinct mode names, such as chase, flank, dump, wiggle, wheelie, holdout and activity_win. Some of the strategies they discovered:

  • Sizing up the opponent. GPT-6 Astra reads the opponent’s motors before engaging. Any hinge stronger than 1,800 N·m is treated as a weapon, so it chooses to intercept the opponent instead of meeting it head-on.
  • Using the rules as a weapon. A robot that stays still for 10 seconds loses. Kimi K3 holds a pin when the opponent’s inactivity timer will run out first, and brakes sharply at the edge so the opponent’s own momentum carries it off.
  • Winning the tie-break. If both robots stall at the same moment, the one whose center of mass is higher wins. Some controllers raise their center of mass in the final seconds of a mutual stall to claim it.
  • Estimating physics on the fly. Claude Fable 5.1 works out how much weight rests on each wheel from its contact measurements, and limits motor torque to the grip its rubber wheels actually have, so they push instead of spin.
  • Planning for being flipped. In the sampling runs alone, 96 controllers detect when their robot is upside down and reverse their drive to keep fighting. Gemini 3.1 Pro’s Apex Juggernaut also reverses its drum, so the exposed surface still strikes upward.
Read about emergent creativity in the paper Explore every robot in the Bot Zoo