Automating the Search for Artificial Life with Foundation Models
Akarsh Kumar, Chris Lu, Louis Kirsch, Yujin Tang, Kenneth O. Stanley, Phillip Isola, David Ha
cs.AI, cs.NE
2024-12-24
ASAL scores ALife simulations with CLIP, turning life-discovery into automated search. It finds new Lenia and Boids lifeforms and CAs more open-ended than Conway's Game of Life.
Artificial life (ALife) studies "life as it could be" through simulation: simple update rules iterated until cell-like or flock-like structures emerge. The field's long-standing bottleneck is the rules themselves. Emergent behavior cannot be predicted from a simulation's configuration, so researchers fall back on intuition and trial and error, hand-tuning worlds that feel right and leaving little room for unexpected discovery.
Automated search has been tried before, but it needs a scorer, something that decides whether a frame "looks alive." Classical complexity measures are either uncomputable or fail to capture human notions of interestingness. The bet here: vision-language models pretrained on internet-scale natural data have representations similar to humans, so they can serve as the judge directly.
ASAL parameterizes a substrate (a family of simulations, e.g. all Lenia worlds) as θ, covering the initial state distribution, the update function, and a renderer. Run T steps, render frames, embed them with CLIP, and the pipeline becomes an optimizable objective. Three search modes share it:
Optimizers switch by substrate: Sep-CMA-ES for Lenia, Boids, and Particle Life; backpropagation through time with Adam for Neural Cellular Automata; plain brute force for Life-like CAs, whose rule space holds only 2^18 = 262,144 rules. The central design choice is outsourcing the judgment of "interesting" to the FM representation, replacing hand-crafted complexity metrics with a pretrained model as the proxy observer.
| Finding | Detail |
| Target search | Prompts like "a caterpillar" or "a self-replicating pattern" produced matching simulations across Lenia, Boids, and Particle Life; in NCA, the sequence "one cell" then "two cells" yielded an update rule that self-replicates and generalizes to a new initial state |
| Open-endedness | Scoring all 262,144 Life-like rules puts Conway's Game of Life in the top 5%; several rules (e.g. B0136/S034678) show CLIP-space trajectories more divergent than Conway's |
| Illumination | The Lenia atlas contains many previously unseen lifeforms resembling microscopy images of cells and bacteria; Boids recovers flocking plus snaking, circling, and grouping variants |
| Quantification | A "caterpillar" in Particle Life emerges only with at least 1,000 particles; sweeping parameters one by one ranks the green-yellow interaction strength as most critical; the rate of change of the CLIP embedding plateaus exactly when a Lenia simulation goes static, giving an automatic halting condition |
An ablation swapping the FM shows CLIP slightly ahead of DINOv2 on illumination, with both far ahead of a pixel representation.
This turns ALife research from designing rules into describing desired phenomena; the researcher writes a prompt and the search does the rest. The payoff is twofold: discoveries that used to depend on luck, and a human-aligned measurement layer that turns qualitative observations such as emergence thresholds, parameter sensitivity, and stopping into computable curves. The pipeline is agnostic to both FM and substrate, so stronger future models plug in directly.
To be clear, this is a paradigm-opening paper. CMA-ES, genetic algorithms, and brute force are off-the-shelf parts; the novelty sits in using an FM representation as the search objective.
The paper's own admissions: open-endedness search covers only Life-like CAs, because preliminary experiments suggested novelty is hard to sustain in Lenia and Boids, and NCA is suspected to be too expressive to yield meaningful structure. Failures in supervised target search are attributed to insufficient substrate expressivity. The deeper dispute, whether open-endedness can be quantified at all, remains open; the paper only moves the subjectivity into the choice of representation.
What is not verified: "human-aligned" rests on prior work linking CLIP to human representations, with no human evaluation study in the paper. Claims of new lifeforms rely on visual inspection with no external benchmark. CLIP's biases toward color and texture flow straight into the open-endedness ranking, and stability across FMs gets only a coarse CLIP-versus-DINOv2 comparison.