The Bitter Lesson breaks down in 3D: procedural descriptions beat learned representations
keenanisalive · x · 2026-09-19
A counterintuitive take on the "Bitter Lesson" in 3D learning: pure scaling of data and compute doesn't automatically yield good shape representations.
- The impressive LLM/agent 3D generation demos all rely on configured procedural languages or software (Three.js, OpenSCAD, Blender), not representations learned in weights.
- Procedural descriptions are vastly more token-efficient than explicit (points, splats, polygons) or implicit (SDF, NeRF) encodings — "unit sphere" compresses enormous geometric information into two words.
- The author frames this as bounding Kolmogorov complexity: finding the domain-specific language that minimizes description length and entropy.
- This doesn't contradict Sutton — input/output encoding always mattered — but "the network will learn the best representation" is a misguided fiction.
- He predicts an era of "code as geometry" versus World Labs' "geometry as code," with DSL design as the key task and computer graphics reviving as the decompression mechanism.
More from Multimodal
- Tencent open-sources WeVisDoc: end-to-end document parsing model turns a page image into Markdown — xiaohu · 2026-09-19
- Awwwards-mcp turns award-winning sites into an agent-readable design library — _insane7 · 2026-09-19
- Story Illustrator: Open-Source Tool Auto-Illustrates Stories via Local LLM + ComfyUI — Natrimo · 2026-09-19
- ComfyUI v0.36.0 Adds FastH3 Video, Marigold V2 Estimation and YuE2 Music Support — Gremlation · 2026-09-19
- Qwen Image 2.1 Open-Source Image Model Teased, Coming Soon — Time-Teaching1926 · 2026-09-19
- Qwen3.8-Omni-Flash undercuts Gemini Flash pricing while matching its multimodal benchmarks — The Decoder · 2026-09-19