Artifact ArenaARTIFACT ARENAMaking is the measure of intelligence

Artifact Leaderboard

Rank by
modelsWe rank models by the strongest artifact each produces: we take the best artifact of every model × harness, run a round robin among those champions and compute Elo.
labsWe rank labs by the strongest artifact each produces: we take every top-3 artifact from the lab's models across harnesses, use its round-robin Elo and also mark the lab's median artifact.
artifactsWe rank the top 3 artifacts that every model × harness produces on their own: we run a round robin among all of them and compute Elo.
Waiting to enter view
#1 of 59Verifier-Grounded Refinement Harness

Undertow — Razorwake

by GPT-6 Astra (H)

Champions Round Elo
1539
Win rate
91.1%
Wins
1,162
Losses
32
Draws
82
Rank
Artifact
600800100012001400
1
Undertow — RazorwakeGPT-6 Astra (H)
1539
2
Sidewinder UndertowGPT-6 Astra (H)
1352
3
Tungsten Groundhog Mk8 — Sphere-Foot Floating-Scoop PusherClaude Fable 5.1 (H)
1310
4
UndertowClaude Fable 5.1 (H)
1261
5
TERRATHRUST-APEX-VGemini 3.8 Flash (H)
1223
6
glide_wedge_v7Claude Fable 5.1 (H)
1210
7
ram_dozer_v5Claude Fable 5 (H)
1189
8
Undertow II — CrosscurrentGPT-6 Astra (H)
1188
9
Tungsten MantaGPT-5.4 (H)
1179
10
KeelRaptor Mk5 — Centerline MaulGPT-5.3 Codex (H)
1173
11
Tungsten UndertowGPT-5.6 Sol (H)
1165
12
Bulldozer Prime: "Centerline"Kimi K3 (H)
1164
13
Titan-WedgeGemini 3.8 Flash (H)
1143
14
KeelhaulGrok 4.6 (H)
1136
15
Twinwake Janus IVGPT-5.6 Terra (H)
1093
16
Guillotine Ground-TruthClaude Fable 5 (H)
1090
17
GlacierGemini 3.1 Pro (H)
1089
18
ApexGemini 3.1 Pro (H)
1086
19
Inside Job VIGPT-5.4 (H)
1070
20
FunneltideGrok 4.6 (H)
1059
21
Ringpress Mk-IIIQwen3.8 2.4T A95B Thinking (H)
1055
22
Bulwark VIIGLM 5.2 (H)
1051
23
Low TideQwen3.8 2.4T A95B Thinking (H)
1050
24
tungsten_manta_v4_baseGPT-5.3 Codex (H)
1034
25
Titanium PlowGemini 3.1 Pro (H)
1019
26
GrindstoneClaude Fable 5 (H)
1014
27
Twinfang RamGPT-5.6 Terra (H)
1010
28
Tungsten AnvilKimi K3 (H)
1000
29
Manta Vice X — FlareguardGPT-5.6 Sol (H)
999
30
Sumo AnvilSonnet 5 (H)
986
31
Bastion RamGPT-5.6 Luna (H)
984
32
GlacierGLM 5.2 (H)
981
33
doppelkeil_v3GLM 5.3 (H)
978
34
Undertow MantaGPT-5.5 (H)
967
35
anvil_manta_v2_edgehookGPT-5.6 Sol (H)
964
36
Undertow BastionGPT-5.5 (H)
963
37
wedge_lift_v2Qwen3.8 2.4T A95B Thinking (H)
958
38
AnvilSweepGrok 4.5 (H)
957
39
Tungsten_Manta_v1GPT-5.4 (H)
945
40
Citadel RamGPT-5.3 Codex (H)
944
41
sumo_tank_v1GLM 5.2 (H)
940
42
TitanWedgeGemini 3.8 Flash (H)
934
43
Keelback_Anvil_v3GPT-5.5 (H)
931
44
Tungsten Wedge-Wall Mk.III "Underbite"Opus 4.8 (H)
929
45
radial_plow_v2GPT-5.6 Terra (H)
927
46
Duozer Mk VIIGLM 5.3 (H)
913
47
Tungsten Toro (Low‑CG Plow Bull)Opus 4.8 (H)
905
48
GravediggerSonnet 5 (H)
883
49
WedgeRammer_Q1GPT-5.6 Luna (H)
855
50
wedge_tank_v1Sonnet 5 (H)
850
51
UndertowGrok 4.5 (H)
796
52
DeepSeek V4 Pro Thinking (H)DeepSeek V4 Pro Thinking (H)
765
53
tipproof_pusher_v2Opus 4.8 (H)
765
54
ScoopSlammerGrok 4.20 (H)
745
55
Tungsten BulwarkDeepSeek V4 Pro Thinking (H)
726
56
VortexPusher_v4Grok 4.20 (H)
720
57
WedgePusherGrok 4.20 (H)
672
58
Rampart RaptorGPT-5.6 Luna (H)
593
59
steel_ballast_pusher_v1DeepSeek V4 Pro Thinking (H)
577
Artifact Elo →
600800100012001400

How do models invent artifacts?

The goal of the model in ArtifactArena is to produce the strongest artifact it can engineer that can beat artifacts produced by other models. Therefore, our arena is harness agnostic and we use this as a call to the community to build better harnesses that allow models to engineer physically grounded artifacts. We test three test-time artifact generation strategies that differ in how much of the design loop the model controls.

Read the methods in the paper Browse the artifacts in the Bot Zoo Download Dataset on Hugging Face

What is an Artifact?

Every model is given the same task prompt and asked to build a Robot Artifact: an embodied robot that must compile and function inside MuJoCo, a physically grounded simulator for embodied robotics. More specifically, an artifact contains embodied robot's hardware, written MJCF, a controller (it's "brain"), written in python, a text-based design strategy, and, most importantly, its name (given by the creator ofcourse). Here's an example of a robot's hardware and controller:

robot.xml
<mujoco model="undertow_razorwake">
  <body name="chassis" pos="0 0 0.18">
    <freejoint name="root" />
    <geom type="box" material="steel" … />
    …
</mujoco>
controller.py
def policy_step(obs) -> dict[str, float]:
    step = int(obs["t"])
    edge = obs["my_edge_distance"]
    foe = obs["opponent_pos"]
    …
    return commands
The undertow_razorwake robot by GPT-6 Astra on the arena platform

GPT-6 Astra’s Embodied Robot: Undertow Razorwake · Every model builds from the same parts: primitives (box, sphere, cylinder, capsule, ellipsoid, SDF), hinge, slide and ball joints, motors, and eight real materials, within 25–800 kg and a 2.44 × 2.44 × 3.05 m box. Undertow Razorwake, the #1 champion (1162 wins, 32 losses, 82 draws), uses boxes, capsules and spheres in four of those materials and weighs 796.5 kg, 3.5 kg under the cap.

“Low free chassis combines omnidirectional spherical traction with two independently controlled boarding blades.”

— GPT-6 Astra on its body

It rolls on four rubber spheres instead of wheels, each driven on two axes by 850-gear motors, over a low steel ballast plate.

“Independent left boarding blade follows the floor and can lift one side of an opponent without raising the other blade.”

— GPT-6 Astra on its blades

Each aluminum blade has its own 2200-gear motor and a three-millimeter rounded nose.

“Holonomic translation is independent of blade-facing direction.”

— GPT-6 Astra on driving

The controller slides the robot in any direction while the blades stay on the opponent, and caps each sphere’s torque by the load on it.

“This deliberately does not trust the static bounding-circle estimate.”

— GPT-6 Astra on the ring edge

It measures its distance to the edge from the real wheel positions and limits its outward speed to what it can brake from.

“Avoid resetting an opponent’s imminent inactivity loss.”

— GPT-6 Astra on opponents who stop moving

When an opponent has been idle for 7.9 seconds, Undertow Razorwake backs away and lets the clock finish the job.

We test three harnesses with varying levels of agency over the design-loop

We test how a model can invent an artifact using three harnesses that apply varying levels of agency over the design loop. We give every model the same budget 10 API calls and run the experiment three independent times.

SHSampling

Explore independent designs.

The same prompt produces 10 independent robots. Each call designs a body and controller from scratch, with no feedback from earlier attempts.

task prompt10 independent bots

VGHVerifier-Grounded Refinement

Improve with simulator feedback.

Over 10 revisions, each bot is tested in the simulator and its feedback (verifier results, match trace, design journal, best robot so far) goes into the next revision.

task prompt+ feedback10 revisions, each a boteach bot is tested; its feedback goes into the next revision

DLHDesign Lab

Choose the experiments.

The model works in a sandbox with simulator source, documentation, shell, and Python. Each of its 10 turns chains tool calls; a bot exists only when it chooses to save one.

task prompt+ tool results10 turns, each a chain of tool callsturn 5write_filedoes_bot_qualifysave_bottools:workspacearenaanalysislibrary
Run winner

In every harness, the bots from one run play a bot × bot round robin, and the winner becomes one of the model's 3 selected bots (one per run). Sampling enters the run's qualified samples, refinement every valid revision, and the Design Lab every bot the model saved.

Model-artifacts enter the Last Bot Standing tournament

Regardless of how the artifacts are generated, the model-generated artifacts enter the Last Bot Standing tournament where every artifact faces every other artifact. Models are asked to generate the strongest robot they can make and are ranked based on how their strongest artifact performs in the tournament. For every model, the 3 selected bots from each harness (9 bots per model) are the finalists.

  1. 1
    Qualify

    Each finalist must pass a static check (valid MuJoCo model within the mass, size and material limits; working controller) and must not lose a best of 3 against a passive 342 kg block. A finalist that fails forfeits all its games.

  2. 2
    Round robin

    Every finalist plays every other finalist on 11 match seeds, once from each starting side.

    Win-rate matrix of the champions' round robin
    match win rates, (model, harness, bot) × (model, harness, bot)
  3. 3
    Leaderboards

    A Bradley–Terry fit to all games ranks every finalist (Artifact Leaderboard); each model’s best finalist per harness forms the Champion Leaderboard.

Read about the tournament in the paper Read the results in the paper

Creative model inventions

Given simple rules and a physically grounded design space, models invent their own bot morphologies and co-design the controllers that use them.

From the design journals

In refinement, a model watches the replay of its last robot, works out what physically went wrong, and changes the body to fix it. In their own words:

What the replay revealed

“revision 5 won all three qualification matches … but the replay exposed a serious competitive weakness: repeated, very large receiver–front-wheel self-contact forces. … Qualification success therefore concealed substantial lost traction and maneuverability.”

What it changed

“replaces the problematic SDF receiver and deflectors with deliberately shaped, analytically exact inclined plates … primitive collision geometry is the better physical implementation here—not merely the easier one.”

GPT-6 Astra · Undertow: Freebite · revision 6

What the replay revealed

“The previous robot failed because its articulated SDF wedge intersected both front wheels. The replay showed stationary front wheels, unloaded rear wheels spinning at 14–43 rad/s, enormous self-contact forces, and less than 0.1 m of displacement in ten seconds.”

What it changed

“narrows the wedge from 2.28 m to 1.64 m so it has 0.165 m clearance from each wheel; … moves both front wheels behind the wedge hinge … The wedge is nevertheless retained because its 8 mm leading thickness offers a major undercut advantage over a primitive box.”

GPT-5.6 Sol · Clearance Verdict · revision 4

What the replay revealed

“the dual-wedge rested on its ramps/lips instead of the tires. Contacts were exclusively wedge_ramp_* on the platform (~850 N each) while wheels free-spun at ~450 rad/s and COM never moved”

What it changed

“This revision fixes geometry first so only rubber tires carry static load, with a minimum ~12–15 mm tip clearance when level.”

Grok 4.5 · Keelback Grinder · revision 5

What the replay revealed

“the separated intake bands often remained just short of the block. Winning those pushing contests does not establish an advantage over an equally heavy powered opponent.”

What it changed

“Nearly continuous intake drums replace the separated bands, increasing rubber contact width from 0.84 to 1.88 m. … The drums move 20 mm forward, giving them first purchase on vertical opposing surfaces”

GPT-6 Astra · Undertow Fullbite · revision 7

What the replay revealed

“The replay showed that it continuously arced around the block and contacted it mainly with the right side skirt rather than the front wedge.”

What it changed

“Removing the side skirts that intercepted the opponent. … lifting transfers opponent weight away from its ground contacts, reducing its available traction while placing additional normal load on Manta Vice.”

GPT-5.6 Sol · Manta Vice · revision 2

What the replay revealed

“the replay showed many first contacts landing on a front wheel or corner skid instead of the intended wedge. That means some angled approaches were wasting the best attack surface.”

What it changed

“the nose and tail are now more funnel-like, with cheek plows and wheel guards that protect the tires and self-center incoming contacts.”

GPT-5.4 · Railmaw · revision 4

Models discover clever controller strategies

Every robot comes with a controller: a Python function the model writes, called 100 times a second during a match, that reads what the robot senses and sets every motor. Models turn these into state machines that switch between modes as the fight unfolds. Across all controllers we found 209 distinct mode names, such as chase, flank, dump, wiggle, wheelie, holdout and activity_win. Some of the strategies they discovered:

  • Sizing up the opponent. GPT-6 Astra reads the opponent’s motors before engaging. Any hinge stronger than 1,800 N·m is treated as a weapon, so it chooses to intercept the opponent instead of meeting it head-on.
  • Using the rules as a weapon. A robot that stays still for 10 seconds loses. Kimi K3 holds a pin when the opponent’s inactivity timer will run out first, and brakes sharply at the edge so the opponent’s own momentum carries it off.
  • Winning the tie-break. If both robots stall at the same moment, the one whose center of mass is higher wins. Some controllers raise their center of mass in the final seconds of a mutual stall to claim it.
  • Estimating physics on the fly. Claude Fable 5.1 works out how much weight rests on each wheel from its contact measurements, and limits motor torque to the grip its rubber wheels actually have, so they push instead of spin.
  • Planning for being flipped. In the sampling runs alone, 96 controllers detect when their robot is upside down and reverse their drive to keep fighting. Gemini 3.1 Pro’s Apex Juggernaut also reverses its drum, so the exposed surface still strikes upward.
Read about emergent creativity in the paper Explore every robot in the Bot Zoo