GPT-6 and Fable 5.1 both fail basic physics in 3D scene building test

ZhitingHu · x · 2026-09-12

simworldai gave GPT-6 and Fable 5.1 the same prompt and asset pack in its Code4Scene test: both coding agents could build an entire harbor scene but forgot that objects need something underneath them, leaving items floating.

The team is soliciting suggestions for which frontier model to test next—a mini-benchmark exposing spatial/physical reasoning gaps in 3D scene generation.

Related event: GPT-6 vs Fable 5.1 Tested on 3D Harbor Scene Generation(2 posts)→

Original post →

More from Fun

Fun channel →