ECCV seethes as GenCeption shows video models can swallow traditional CV research
CSProfKGD · x · 2026-09-17
- Letian Wang reports deep frustration among senior attendees at CVPR, and even more at ECCV 2026 right after the GPT-6 release, where researchers openly questioned how much traditional CV research will be absorbed by large foundation models.
- Michael Black (attending ECCV since 1992) asks the same questions; his VIGA paper solves inverse graphics (image → 3D Blender scene) with an agentic approach, after an early rejection delayed it for years.
- The catalyst is GenCeption (Google DeepMind + Toronto/Oxford/MIT etc., authors include Kaiming He and Andrew Zisserman): it repurposes a pretrained video generation model into a single unified feed-forward vision model that, steered only by text instructions, matches or beats specialized SOTA like DepthAnything3 across many vision tasks, with striking learning efficiency and emergent behavior.
- The parallel to NLP is explicit: vision is transitioning from task-specific models to general-purpose visual intelligence.
More from AGI Musings
- ValsAI launches MysteryMechanism, a benchmark testing whether AI agents can rediscover sealed math mechanisms — burny_tech · 2026-09-17
- UW professors on the 'grand challenge' of building transparent, beneficial AI — lazowska · 2026-09-17
- Get past the next 20 years and AI could usher in a golden age — Objective_Singer_404 · 2026-09-17
- AI is killing cold email: professor gets 75 PhD inquiries months before Fall 2027 cycle — xuanalogue · 2026-09-17
- 25 Fields Medalists Sign Declaration Warning Against Measuring AI Progress by Solved Problems — krishnan · 2026-09-17
- How Effective Altruism Lost Its Way: Matt Johnson's 2023 Essay Still Holds Up — inductionheads · 2026-09-17