TAPe+ML: sub-100K-parameter vision system hits 84.7 mAP50 on COCO detection
Comexp · hf · 2026-09-22
Researchers present TAPe+ML v3, built on the Theory of Active Perception (TAPe): a structured representation encoding relations among perceptual elements before recognition, instead of operating on raw pixel tensors, paired with a modular recognition architecture for classification, detection, and instance segmentation.
- Scale: fewer than 100K parameters total
- COCO detection: 84.7 mAP50 / 65.3 mAP50-95
- COCO instance segmentation: 80.7 mask mAP50 / 58.4 mask mAP50-95
- Classification: 92% validation accuracy on Imagenette under identical training vs a raw-pixel baseline; 89.9% Top-1 on ImageNet-Real
- Extensions: compactness in video scene detection and adaptation under distribution shift in an industrial pilot
Core claim: shifting part of the modeling burden from network parameters to structured input representations can support compact multi-task vision systems with reduced data, memory, and compute.
More from Research
- Berkeley talk: A new conceptualization of language across humans, animals, and machines — begusgasper · 2026-09-22
- Yarin Gal: Most LLM-written papers won't stand the test of time — yaringal · 2026-09-22
- Estimating Unitree G1 actuator heat loss in Isaac Sim with a physics-based model — IsaiahBallah · 2026-09-22
- Virtual Biotech teardown: 6 design decisions behind Science's 37,000-agent drug company — bravo_abad · 2026-09-22
- LinearSolveBench debuts: testing if AI can write fast C solvers for sparse linear systems — hgarud · 2026-09-22
- Navier-Stokes partial regularity machine-checked in Lean by 50-agent swarm in 36 hours — RexDouglass · 2026-09-22