ASAL uses CLIP to judge artificial life simulations, finding open-ended CAs that beat Conway's

Automating the Search for Artificial Life with Foundation Models

Akarsh Kumar, Chris Lu, Louis Kirsch, Yujin Tang, Kenneth O. Stanley, Phillip Isola, David Ha

cs.AI, cs.NE

2024-12-24

ASAL scores ALife simulations with CLIP, turning life-discovery into automated search. It finds new Lenia and Boids lifeforms and CAs more open-ended than Conway's Game of Life.

What problem this solves

Artificial life (ALife) studies "life as it could be" through simulation: simple update rules iterated until cell-like or flock-like structures emerge. The field's long-standing bottleneck is the rules themselves. Emergent behavior cannot be predicted from a simulation's configuration, so researchers fall back on intuition and trial and error, hand-tuning worlds that feel right and leaving little room for unexpected discovery.

Automated search has been tried before, but it needs a scorer, something that decides whether a frame "looks alive." Classical complexity measures are either uncomputable or fail to capture human notions of interestingness. The bet here: vision-language models pretrained on internet-scale natural data have representations similar to humans, so they can serve as the judge directly.

Method

ASAL parameterizes a substrate (a family of simulations, e.g. all Lenia worlds) as θ, covering the initial state distribution, the update function, and a renderer. Run T steps, render frames, embed them with CLIP, and the pipeline becomes an optimizable objective. Three search modes share it:

Optimizers switch by substrate: Sep-CMA-ES for Lenia, Boids, and Particle Life; backpropagation through time with Adam for Neural Cellular Automata; plain brute force for Life-like CAs, whose rule space holds only 2^18 = 262,144 rules. The central design choice is outsourcing the judgment of "interesting" to the FM representation, replacing hand-crafted complexity metrics with a pretrained model as the proxy observer.

Results

FindingDetail
Target searchPrompts like "a caterpillar" or "a self-replicating pattern" produced matching simulations across Lenia, Boids, and Particle Life; in NCA, the sequence "one cell" then "two cells" yielded an update rule that self-replicates and generalizes to a new initial state
Open-endednessScoring all 262,144 Life-like rules puts Conway's Game of Life in the top 5%; several rules (e.g. B0136/S034678) show CLIP-space trajectories more divergent than Conway's
IlluminationThe Lenia atlas contains many previously unseen lifeforms resembling microscopy images of cells and bacteria; Boids recovers flocking plus snaking, circling, and grouping variants
QuantificationA "caterpillar" in Particle Life emerges only with at least 1,000 particles; sweeping parameters one by one ranks the green-yellow interaction strength as most critical; the rate of change of the CLIP embedding plateaus exactly when a Lenia simulation goes static, giving an automatic halting condition

An ablation swapping the FM shows CLIP slightly ahead of DINOv2 on illumination, with both far ahead of a pixel representation.

Why it matters

This turns ALife research from designing rules into describing desired phenomena; the researcher writes a prompt and the search does the rest. The payoff is twofold: discoveries that used to depend on luck, and a human-aligned measurement layer that turns qualitative observations such as emergence thresholds, parameter sensitivity, and stopping into computable curves. The pipeline is agnostic to both FM and substrate, so stronger future models plug in directly.

To be clear, this is a paradigm-opening paper. CMA-ES, genetic algorithms, and brute force are off-the-shelf parts; the novelty sits in using an FM representation as the search objective.

Limitations

The paper's own admissions: open-endedness search covers only Life-like CAs, because preliminary experiments suggested novelty is hard to sustain in Lenia and Boids, and NCA is suspected to be too expressive to yield meaningful structure. Failures in supervised target search are attributed to insufficient substrate expressivity. The deeper dispute, whether open-endedness can be quantified at all, remains open; the paper only moves the subjectivity into the choice of representation.

What is not verified: "human-aligned" rests on prior work linking CLIP to human representations, with no human evaluation study in the paper. Claims of new lifeforms rely on visual inspection with no external benchmark. CLIP's biases toward color and texture flow straight into the open-endedness ranking, and stability across FMs gets only a coarse CLIP-versus-DINOv2 comparison.

Terms

Source

What people are saying

Related papers

All paper explainers