Samsung's reViT: one recurrent Transformer block matches full-depth encoders with ~70% fewer parameters
SamsungResearch · hf · 2026-10-09
Samsung Research introduces reViT, a recurrent vision Transformer where a single block applied repeatedly matches full-depth encoder accuracy at comparable inference FLOPs, with FFNs at each depth represented as mixtures of a small shared expert bank programmed by a normalized-depth coordinate.
Key results:
- reViT-B/16 trained from scratch attains DeiT III accuracy with 70% fewer stored parameters
- Weight-space merging outperformed token-dispatch and output-mixture MoE alternatives at a one-FFN budget
- An 8-expert model distilled from DINOv2 output features alone retains nearly all linear-probe accuracy and transfers across classification, segmentation, and depth prediction
- Elastic-depth training lets one checkpoint run at multiple depths, and fixed-depth deployment can be materialized as a dense graph without online routing
More from Research
- OpenAI's math results released as open source RL environments — SergioPaniego · 2026-10-09
- Mathematicians slam OpenAI's 700+ math papers as 'impossible to read without AI help' — burny_tech · 2026-10-09
- OpenAI releases 372 math proof claims, including 23 Erdős problems spanning 1268 pages — burny_tech · 2026-10-09
- Claim-Locked Reporting: EMNLP paper fixes LLM hallucinated claims over correct numbers — jiqizhixin · 2026-10-09
- NVIDIA NeMo-DCR cuts 1T-model RL weight sync from 87.5 min to 150 sec — IanAndrewsDC · 2026-10-09
- Dex-One2Many: one human video trains a dexterous hand that transfers zero-shot to real robots — furongh · 2026-10-09