New paper shows models can learn general patterns from zero real-world data

mark_k · x · 2026-09-25

A new paper describes a training setup that uses no real-world data at all: two models start from scratch, one writing small programs that generate byte sequences and the other learning to predict them. The first model then learns to generate sequences that push the second to improve.

After training only on these generated sequences, the learner got better at predicting real text, images, speech, and DNA as compute scaled up — an early proof of concept that useful general patterns can emerge without any natural examples.

Original post →

More from Research

Research channel →