Samsung's reViT: one recurrent Transformer block matches full-depth encoders with ~70% fewer parameters

SamsungResearch · hf · 2026-10-09

Samsung Research introduces reViT, a recurrent vision Transformer where a single block applied repeatedly matches full-depth encoder accuracy at comparable inference FLOPs, with FFNs at each depth represented as mixtures of a small shared expert bank programmed by a normalized-depth coordinate.

Key results:

Original post →

More from Research

Research channel →