Tesla FSD Technical Reveal: Shift to End-to-End Video Model
PTrubey · x · 2026-08-24
Tesla's FSD head revealed a complete rewrite from modular assembly to a single video foundation model, following a 'Video in, Action out' methodology similar to GPT training.
Key Points:
- Failure of Modularity: Accumulated interface errors, unenumerable long tails, and violation of the Bitter Lesson.
- End-to-End Benefits: Enables homogeneous compute and deterministic latency.
- Core Challenge: The curse of dimensionality (massive input from 8 cameras vs. minimal output DOF).
- Training Method: Mimics human learning—starting with conscious focus on 'proxy tasks' (lanes, signs) before transitioning to subconscious processing.
Related event: Tesla FSD Talk Reveals End-to-End Video Model Training(2 posts)→
More from Embodied
- Xiaomi launches local AI host with three custom chips supporting dual models — op7418 · 2026-08-24
- SMPLOlympics: RL Policies Trained for 10+ Humanoid Sports in Simulation — zhengyiluo · 2026-08-24
- Robot Olympics surprisingly fun to watch, says Linus Ekenstam — LinusEkenstam · 2026-08-24
- Robotics funding triples to $8.76B in 2025, production lessons from the 'bubble' — davidyin44 · 2026-08-24
- China vs US Robotics Tech Tree: China Focuses on Locomotion, US on Dexterity — AccBalanced · 2026-08-24
- Tesla FSD talk reveals end-to-end training works like the human brain learning to drive — PTrubey · 2026-08-24