753B model thinks, 4B model writes: latent-space handoff matches frontier reasoning at 20x speed

TheMoonMidas · x · 2026-09-03

mostikai, a team of 12 PhDs who spent four months building, claims first place on the ARC-AGI leaderboard and argues the wrong question is "will open models catch up?" — the right one is why a frontier model should generate your answer at all.

Their protocol lets models communicate in latent space: a 753B frontier model "thinks," and its hidden states flow directly into a 4B model running on your own infrastructure, which "writes" the answer. No text passes between the models, and neither is fine-tuned. The result: performance near the big model at roughly 20x the speed.

The idea effectively decouples reasoning (the expensive part) from generation (the cheap part), using the small model as a latent-space decoder. Details are scarce since the competition is still running, and the claimed results await independent verification.

Related event: Mostik Bridges Large and Small Models via Latent-Space Communication, Tops ARC-AGI Leaderboard(5 posts)→

Original post →

More from Models

Models channel →