2021's CABiNet beats YOLO26-sem on UAVid: +2.7 mIoU at 3x lower latency
Naive-Explanation940 · reddit · 2026-09-02
The original first author of CABiNet (ICRA 2021) rebuilt the repo and ran a controlled comparison against YOLO26 semantic segmentation variants on the aerial UAVid dataset, standardizing data, class weighting, and single-scale evaluation while keeping each model's own training recipe.
Key results (1024×1024, RTX 4070 SUPER, FP16)
- CABiNet-L: 67.14 mIoU, 9.17M params, 4.44ms latency (225 FPS) — +2.7 mIoU over YOLO26x at 3x lower latency
- CABiNet-S: 65.25 mIoU, 3.09ms; at 44 GFLOPs it beats same-compute YOLO26s by +3.6 mIoU with near-identical latency
- YOLO26n/s are legitimate low-latency Pareto points, but YOLO26m/l/x are strictly dominated — slower and less accurate
The author stresses this is not an architecture-only ablation (pretraining, epochs, optimizer, augmentation differ), but the takeaway is that a 2026 general multi-task model doesn't automatically beat a purpose-built efficient architecture on its home task.
More from Embodied
- LimX's semi-humanoid Tron2 pairs industrial arms with humanoid base for high-altitude work — CyberRobooo · 2026-09-02
- Yacine: Training Cross-Robot Generalist Models Is the Easiest Path to Sim2Real — yacineMTB · 2026-09-02
- DexForce W1 Pro wheeled humanoid shows off popcorn-scooping service skills — CyberRobooo · 2026-09-02
- MIT's CW-Net Explains Self-Driving Decisions Using Human-Readable Concepts — MIT News AI · 2026-09-02
- LimX Dynamics TRON 2 humanoid shows formwork assembly and rebar tying for construction — chris_j_paxton · 2026-09-02
- Dev jokes about Physical AI: at least software only gets your drive reformatted — burhop · 2026-09-02