Why Multimodal Input Matters for AGI: DeepSeek & Anthropic's Approach
dotey · x · 2026-08-02
Countering the claim that multimodality is irrelevant to AGI, the author argues that since human-built physical and digital worlds rely heavily on specific modalities like vision, AI must process these inputs to accurately represent the world and execute long-horizon tasks.
Regarding multimodal output, the author notes that both DeepSeek and Anthropic treat it as a vertical engineering problem (e.g., diffusion models) that can eventually be externalized and called upon by general intelligence. Furthermore, DeepSeek hasn't abandoned vision; its non-open-weight versions already process visual signals, and enterprises are actively integrating ViT modules with it.
More from AGI Musings
- Redis vs Keras Authors Debate: Deep Inference or Mechanical Parroting in LLMs? — antirez · 2026-08-03
- AI Boosts Scientific Productivity but May Stifle Radical Breakthroughs — JMateosGarcia · 2026-08-03
- Does AI Diminish Human Glory? The Cure Matters More Than the Discoverer — nptacek · 2026-08-02
- AI in Math Acts Like a Narrow-Minded Obsessive, Missing the Forest for the Trees — neurovium · 2026-08-02
- Ethical Debate: When Will Manual Driving Become Obsolete? — cgarciae88 · 2026-08-02
- MIT's Catalini: Traditional Moats Fail in AI Era, Only Verification-Grade Network Effects Survive — kimmonismus · 2026-08-02