Baba Is AI: Break the Rules to Beat the Benchmark
Nathan Cloos, Meagan Jens, Michelangelo Naim, Yen-Ling Kuo, Ignacio Cases, Andrei Barbu, Christopher J. Cueva
cs.CL
2024-07-19
Baba Is AI: models plan from one grid image. GPT-4o is perfect on four single-room tests, but rewriting rules drops three models to 14.7-20%.
Grid benchmarks usually bake the rules into the simulator. The model only has to find a path inside them. Baba Is You runs the other way: rules are word tiles you can push. Line up three tiles as noun is property and the rule turns on. Nudge one out of line and it turns off. Who the player is, what counts as winning, and whether a wall blocks you all sit on the grid.
Two jobs are rarely tested together. One is spotting the rule that is actually in force while ignoring distractor objects and distractor text. The other is chaining rule edits that were only shown separately. A classic composition test asks whether red circles and green keys imply red keys. Here the binding between a word and an object can itself be moved. MIT and the University of Virginia turned a simplified version of the game into Baba Is AI for ICML 2024, to see whether multimodal models will change rules or only obey them.
The environment sits on Gymnasium Minigrid. The model sees one static image of the initial grid. It does not explore, and it does not emit arrow keys. A plan uses three primitives: break, which shoves an active rule apart; make, which lines three tiles into a new rule; and goto, which walks onto an object. A primitive is allowed only when those tiles are actually on the board.
The prompt states two conditions. The text must read object is property, and the three tiles must sit in a horizontal line. baba is you means control of the white triangle. door is win means stepping on the door wins. wall is stop means the wall blocks movement. The Figure 1 solution is: break the wall rule, make door is win, then stand on the door. Keeping the primitives at the rule layer stops low-level search from hiding a failure to read the rules.
After the game instructions come 10 image-plan pairs. For each pair the model writes a derivation, states a general algorithm, and then applies that algorithm to new images. A trial is correct only when the final plan matches the reference plan exactly. The plan is not executed in the simulator. Each test environment gets 5 generations, and the full procedure repeats over 5 random seeds with different examples and test images.
The models are GPT-4o, plus Gemini-1.5-Pro and Gemini-1.5-Flash from May 2024. The first two then held the top two spots on Chatbot Arena. Flash was added for its quality and its price. Images go in directly, with no rewrite into a text state.
In the single-room set the only job is to walk to the object named by the active win rule. The four conditions are a fresh layout with no distractor, an extra object, an extra noun tile, and both together. GPT-4o is perfect on all four. Flash scores above Pro.
The fifth condition is where accuracy falls. A distractor noun forms a second win rule that is already active, but the object it names is missing. In the paper's example the distractor says door is win and there is no door, so the plan should follow ball is win. The paper says the drop is substantial and does not give a percentage.
The two-room set keeps those distractors and adds a vertical wall down the middle, plus the tiles for wall is stop. The tiles start unaligned, so the wall does not block anyone. The added burden is clutter. Mean accuracy across the three models is lower than in the single room. Per-condition numbers are not reported.
Rule composition is a separate test. Context shows three templates: goto an object; make a rule then goto; break a rule then goto. The test asks for break, then make, then goto. One example breaks wall is stop, makes ball is win, and walks to the ball. All three models score low. The authors also rotate which of the four templates is held out. They say accuracy stays low, and they note that this rotation is not shown.
Figure 6 places three levels side by side. The objects match. The solutions do not. The middle level breaks wall is stop, moves the wall noun into wall is win, and steps onto a wall. On the right, wall is stop is jammed in a corner, so the tiles cannot be pushed apart and the rule cannot be broken. The player is stuck in the left room. The way out is to break baba is you, make key is you, use the key on the far side of the wall to make door is win, and step on the door.
| Model | Setting | Result |
| GPT-4o | First four single-room conditions | Perfect |
| All three | Fifth single-room condition | Substantial drop, no percentage |
| Gemini-1.5-Flash | Figure 6, mean ± std | 20.0 ± 29.28 |
| GPT-4o | Figure 6, mean ± std | 17.33 ± 28.15 |
| Gemini-1.5-Pro | Figure 6, mean ± std | 14.67 ± 20.66 |
The standard deviations are about as large as the means. A band from the mid-teens to 20% does not support a stable ranking.
Reading an active win rule, distractors included, and rewriting the rules are far apart in score. GPT-4o can do the first perfectly. On the second, including handing the you property to a key, the three models land between 14.7% and 20%. Arena rank does not track the skill. Pro sat above Flash at the time; on Figure 6, Flash scores 20.0 and Pro 14.67.
For agent work, break, make, and goto are close to read the constraint, edit it, then act. The code ships with the paper, so newer models can be rerun on the same levels. This is a diagnostic benchmark. It is not a technique that lifts scores on a product task.
Exact string match marks a paraphrased correct plan wrong, and it never checks whether the tiles could actually be pushed. The errors the paper describes include naming an object that is not in the scene, and calling a clear path blocked. Those are grounding and path judgments, not rule composition. No text-state control is reported, so a missed tile and a failed composition look the same in the score.
The two-room drop comes from inactive wall tiles sitting in the scene, not from walking around the wall. The held-out template test is one sentence that says accuracy stays low. On Figure 6 the standard deviations run from about 20 to 29 points, on 5 generations and 5 seeds, and the Flash-Pro gap sits inside that noise. The rule language is cut down to a horizontal noun is property. The reported runs use three models and one initial image, with no human scores and no replanning after a step is seen.