Artifact ArenaARTIFACT ARENAMaking is the measure of intelligence
Back to Artifact Arena
Paper · arXiv preprint

ArtifactArena: Evaluating Models by What They Build in the Physical World

Preprint · last updated September 26, 2026

Authors

Kushagra Tiwary*1,2, David Mayo*1, Nikhil Behari2, Xiangzhou Sun1, Abdulrahman Alabdulkareem1, Isaac Galatzer-Levy3, Boris Katz1, Brian Cheung†4

  1. 1InfoLab, MIT CSAIL
  2. 2Camera Culture, MIT Media Lab
  3. 3Department of Psychiatry, NYU Grossman School of Medicine
  4. 4Discovery Lab, UCSF

* Equal contribution·† Corresponding PI

† Correspondence: {ktiwary@mit.edu, dmayo2@mit.edu, bcheung@ucsf.edu}

Email the authors
Download PDFGitHub (Coming Soon!)
Abstract

To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce ArtifactArena, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical design capabilities through three harnesses that refine their bot artifacts based on text descriptions, physics simulator feedback, and gameplay data. We benchmark these capabilities with an Elo ranking of frontier models derived from head-to-head tournaments between their artifacts. By releasing this framework and tournament infrastructure for ongoing community submissions, we establish a living, non-saturating testbed to continuously measure the expanding limits of open-ended intelligence in the physical world.