GTR: Softmax-Free Recurrent Vision Backbone Hits 58.9 COCO AP at 1.9ms on RTX 4090
Intellindust · hf · 2026-10-05
Intellindust introduced Gated Token Recurrence (GTR), a softmax-free recurrent vision backbone combining gated linear attention, alternating spatial scan directions, and spatially enhanced SwiGLU blocks to overcome the quadratic cost of global softmax attention at high resolutions.
Key results:
- Distilled from a detection-specialized DINOv3 teacher using only final-layer patch-token alignment — no masked-token prediction or intermediate-layer supervision
- With Objects365 pre-training, GTR-L achieves 58.9 box AP on COCO val2017 at 1.908ms median batch-one latency on RTX 4090 (compiled FP16)
- Transfers to instance segmentation, pose estimation, oriented detection, semantic segmentation, and monocular depth estimation
- Custom chunkwise CUDA operator is 4.0x faster than FLA v0.5.0 at 1.6K tokens
- TensorRT on DRIVE AGX Thor: 2.282–8.769ms latency, targeting edge deployment
More from Infra
- UBS projects AI rack capacity to 4x to 104.5 GW by 2030, with ASICs taking ~43% share — AccBalanced · 2026-10-05
- Geekerwan teardown reveals Huawei Kirin 9050 Pro's dual-die LogicFolding architecture — teortaxesTex · 2026-10-05
- Sovereign compute isn't about owning hardware: the real test is whether you can break dependencies — AccBalanced · 2026-10-05
- Kirin 9050 Pro efficiency tested: matches 3nm flagships despite 7nm process — teortaxesTex · 2026-10-05
- Modern AI data centers with GB200/300 chips cost an estimated $57.2M per MW — AccBalanced · 2026-10-05
- A 1 GW fusion reactor could breed ~2 tonnes of gold a year from mercury, half its revenue — anselm · 2026-10-05