ZipDepth: a 6.1M-parameter model brings real-time zero-shot monocular depth to phones
rsasaki0109 · x · 2026-09-07
Official ECCV 2026 implementation of "ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device." A compact 6.1M-parameter network for zero-shot monocular depth estimation, running in real time from server GPUs to mobile phones via knowledge distillation from foundation models.
It combines reparameterizable convolutions (RepVGG), efficient channel and spatial attention (strip pooling and global context blocks), and a compact FPN decoder, achieving the best accuracy-efficiency trade-off among lightweight methods — approaching transformer-based foundation models at a fraction of the cost.
More from Research
- kalomaze restates compositional skill model of cross-domain generalization — kalomaze · 2026-09-07
- Why LLM judge ensembles amplify bias: correlated models boost systemic errors — IanArawjo · 2026-09-07
- kalomaze: fluid cross-domain generalization hinges on composing skills end-to-end — kalomaze · 2026-09-07
- LLM judge ensemble hit high IRR yet produced a bogus super-significant result — IanArawjo · 2026-09-07
- Dharmamitra lexicon adds 584k Sanskrit headwords and 3.47M cited attestations — SebastianNehrd2 · 2026-09-07
- Simulating a fruit fly's 166,700-neuron brain in the browser with WebGPU — Vjeux · 2026-09-07