Epoch AI's Furniture Assembly Benchmark: top model score jumped from 28% to 80% in 10 months
rohanpaul_ai · x · 2026-09-25
Epoch AI released the Furniture Assembly Benchmark (FAB), which gives models an instruction manual plus a photo of a half-assembled piece of furniture and asks them to spot the mistake — a practical test of spatial reasoning.
Key points:
- Top score climbed from 28% to 80% in just 10 months
- Spatial reasoning used to be the go-to example of what models couldn't do
- Building the dataset presumably required someone to assemble furniture wrong on purpose
Related event: Epoch AI's Furniture Assembly Benchmark Shows Big Spatial Reasoning Gains(2 posts)→
More from Research
- New Research Examines How AI Agents Enter Markets and How Markets Should Be Designed for Them — sethlazar · 2026-09-25
- Vite core member finds methodology flaws inflating oj's memory benchmarks by 100MB+ — cnakazawa · 2026-09-25
- DeepMind Paper: One Misleading Hint Drops Coding Agent Scores by Up to 46.7% — dair_ai · 2026-09-25
- NeurIPS paper shows self-improving, self-replicating agents evolve cooperation from scratch — maxhkw · 2026-09-25
- Continuous diffusion LM RePlaid accepted at NeurIPS 2026, matches discrete diffusion scaling — ArashVahdat · 2026-09-25
- AAAI 2027 phase 1 reviews are out — CSProfKGD · 2026-09-25