aimotive driving dataset: train labels come from a tracker that sees the future
RexDouglass · x · 2026-10-09
Voxel51 released the aimotive-multimodal autonomous driving dataset on Hugging Face, highlighting a counterintuitive labeling quirk: the train split is auto-labeled by a non-causal tracker that watches the whole 15-second clip before placing boxes, so models train on hindsight while the human-labeled validation split grades them in real time.
Dataset details:
- 176 MCAP episodes, viewable in FiftyOne, filterable by day/night/rain with lidar + camera + radar on one timeline
- Boxes out to 200 meters; a quarter of 425k frames extend beyond 75 meters
- Sensors: lidar + 4 cameras + 2 radars, across day/night/rain in three countries
The train-on-hindsight vs. human-graded asymmetry is a notable bias to keep in mind when interpreting results on this data.
More from Embodied
- Tesla Vision-Only Robotaxi Handles Nighttime Downpours, Silencing Earlier Skeptics — elonmusk · 2026-10-09
- Hark Sees Explosive 48-Hour Growth, Brett Adcock Says Team Scrambling to Fix Infra — adcock_brett · 2026-10-09
- Ruggedize, first agtech conference for field-ready farm robots, draws builders and investors — Scobleizer · 2026-10-09
- Atomic Machines emerges from 6-year stealth with AI-native Matter Compiler that builds micro-machines from code — Scobleizer · 2026-10-09
- Hands-On With Astra Robot Control: Johns Hopkins Deep-Dive Tests Its Motion Understanding — jmin__cho · 2026-10-09
- Pollen's Microduck robot gets custom board: 3h battery life, CPU drops from 114°C to 60°C — huggingface · 2026-10-09