Sony's Syn-Omni: Shared + Expert LoRA Paths Beat Omnimodal Embedding Baselines Across 81 Tasks
_reachsumit · x · 2026-10-09
Sony's paper Syn-Omni [2610.12256] (EMNLP 2026 Findings) proposes a structured omnimodal embedding framework.
- Problem: existing omnimodal embedding methods share a single parameter space over mixed-modality data, limiting separation of universal and modality-specific features
- OME-LoRA (Orthogonal Modality-Expert LoRA): decomposes adaptation into a shared LoRA path for universal semantics plus per-modality expert LoRA paths
- PSR (Progressive Synergy Routing): experts first establish modality-specific priors, then gradually interact across modalities
- Evaluated on 81 tasks spanning image, video, audio, and audiovisual modalities, consistently outperforming omnimodal baselines; code released
More from Multimodal
- You can now spot Opus AI video slop by its sound: synced beats as a fingerprint — hudzah · 2026-10-09
- Monkey King riding a tiger: AI video nails a stunning Chinese-style action scene — lucky-plume · 2026-10-09
- Autoregressive Retriever (ARR) Refines Queries with Retrieved Item Feedback via SFT and RL — _reachsumit · 2026-10-09
- LEGO: lifting-free exocentric-to-egocentric video generation beats depth-lifting SOTA pipelines — 25frms · 2026-10-09
- RISEBench++: 65 reasoning-based visual editing tasks; best model GPT-Image-2.5 hits only 56.6% — VisionXLab · 2026-10-09
- VibeEdit Replaces Text Prompts with Canvas Marks, Scoring 79.9 on Edit Benchmark — Sydney-Uni · 2026-10-09