Why Multimodal Input Matters for AGI: DeepSeek & Anthropic's Approach

dotey · x · 2026-08-02

Countering the claim that multimodality is irrelevant to AGI, the author argues that since human-built physical and digital worlds rely heavily on specific modalities like vision, AI must process these inputs to accurately represent the world and execute long-horizon tasks.

Regarding multimodal output, the author notes that both DeepSeek and Anthropic treat it as a vertical engineering problem (e.g., diffusion models) that can eventually be externalized and called upon by general intelligence. Furthermore, DeepSeek hasn't abandoned vision; its non-open-weight versions already process visual signals, and enterprises are actively integrating ViT modules with it.

Original post →

More from AGI Musings

AGI Musings channel →