DepthART Shrinks Monocular Depth to 6M Params, Hits 0.918ms at 224² on TensorRT
机器之心 · wechat · 2026-10-03
DepthART (Depth Anything Rethought for Tiny Models), accepted at ACM Multimedia 2026, explores whether foundation-style monocular depth estimation can survive at roughly 6M parameters while keeping cross-scene generalization and on-device deployability.
Two core problems and fixes
- Data, not architecture, breaks first at small scale: multi-source training data is imbalanced, and a 6M model's limited capacity lets high-frequency data dominate. The authors propose Bias-Resistant Data Sampling (BRDS): starting from 44M multi-source candidate images, they lower the sampling probability of dense, redundant samples in visual feature space while keeping sparse regions, yielding 1.7M more balanced images. Pseudo-depth supervision from DepthAnything V2-L then distills relative-depth priors into a TinyViM+DPT student.
- Metric fine-tuning causes forgetting: full fine-tuning on NYUD/KITTI overwrites general geometry. Camera-conditioned Fine-tuning (CamFT) freezes the distilled encoder, injects intrinsics via a lightweight camera adapter, and recovers scale with a multi-query scale estimation head — protect geometry first, then learn scale.
Performance and deployment
- Three sizes: 6.0M / 11.4M / 32.6M params. DepthART-S reaches δ1=0.964 zero-shot relative depth on NYUDv2; DepthART-L hits δ1=0.958 on ETH3D. DepthAnything V2-S is 24.8M params, about 4x larger than DepthART-S.
- In strict FP32, DepthART-S runs 347 FPS on an RTX A6000 at 224²; the TensorRT FP16 build cuts model-only latency to 0.918ms. PyTorch, ONNX, TensorRT and Jetson Orin NX paths are provided.
Later extensions
- A September 2026 release adds MobileNetV4-S/M and MobileNetV4-M-slim-SPF (pruned encoder plus a lighter Single-Path Pyramid Fusion replacing the DPT decoder): slim-SPF is 6.05M params / 3.78 GMAC, with NYUD δ1 dropping only from 0.953 to 0.952 and KITTI from 0.931 to 0.928 — over half the compute cut with performance largely intact. It exports using only standard ONNX ops, easing porting to older BPU/NPU platforms.
- The authors note lightweight depth is becoming its own track: ZipDepth (Jul 9), DepthART (Jul 19) and Ultralytics YOLO26-Depth (Jul 22) appeared almost simultaneously, shifting competition from scaling foundation models to bringing foundation capability onto devices.
Paper, code and weights are open-sourced.
More from Embodied
- "Optimus is AGI": Tesla humanoid demo sparks hype on X — ns123abc · 2026-10-04
- Dev ports Meta's open-sourced Muse gadget to ESP32 devices, gets it running with Claude — alexandr_wang · 2026-10-04
- Tesla exec: FSD's super-human reaction times beat insurance at keeping you crash-free — aelluswamy · 2026-10-04
- Local voice assistant auto-picks its Ollama model by VRAM, shows status on Arduino — PrimeEmre · 2026-10-04
- Tesla FSD drives chest-pain victim to hospital after son remotely reroutes the car — XFreeze · 2026-10-04
- Musk Endorses Case That Humanoids Supply 8,700 Hours a Year vs. a Human's 2,000 — elonmusk · 2026-10-04