Apple ML Research proposes RayRoPE for multi-view attention
Apple ML Research · rss · 2026-07-20
RayRoPE: Projective Ray Positional Encoding for Multi-View Attention
Apple ML Research studies positional encoding for multi-view transformers that take a set of posed input images. The paper argues that prior absolute or relative encodings do not satisfy three goals at once: uniquely identifying patches, enabling SE(3)-invariant attention with multi-frequency similarity, and adapting to scene geometry.
To address this, the authors introduce RayRoPE. It represents patch positions using the rays associated with each patch, but instead of relying only on ray direction, it leverages a predicted point along the ray as part of the encoding. This is designed to better fit the geometry of the underlying scene while preserving useful invariances for attention.
More from Research
- OpenAI-style autonomous researchers could become real scientific collaborators — Promptmethus · 2026-07-21
- Soft Clamp cuts tool-call overuse in multi-teacher distillation, from 13.7% to 9.0% — antgroup · 2026-07-21
- ShotPlan adds learnable planning tokens for cinematic multi-shot video generation — Tele-AI · 2026-07-21
- A silicon photonic reservoir chip compensates fiber distortion in real time at 28 Gbps — bravo_abad · 2026-07-21
- A developer maps out six design rules for CLIs that humans and AI agents can both use — yujiezha · 2026-07-21
- GPT 5.6 vs. Claude Fable tested in Dyad AI for Physical AI model tuning — ChrisRackauckas · 2026-07-21