Why LLMs Can't Count Windows: Elorian CEO on Multimodal Reasoning
bigdata · x · 2026-08-27
Andrew Dai, CEO of Elorian AI, discusses why today's frontier models still struggle with complex visual reasoning, arguing that scaling language-centric architectures won't solve it. He outlines Elorian's approach to building specialized multimodal foundation models, covering data, architecture, reinforcement learning, and pre-training. The conversation also touches on visual chain-of-thought, the difference between generation and understanding, and applications in video, robotics, and CAD.
More from Embodied
- NVIDIA's Jensen Huang envisions a future where every Disney character is a robot — CyberRobooo · 2026-08-27
- QNX partners with Hailo for edge Physical AI: 14x performance consistency — pdamodaran · 2026-08-27
- Hugging Face launches $399 open-source duck robot Microduck — TechCrunch AI · 2026-08-27
- Robotics Moat: Data, Hardware, and the Deployment Layer Flywheel — chris_j_paxton · 2026-08-27
- Pollen Robotics and Hugging Face release Microduck, an open-source bipedal robot — -Cubie- · 2026-08-27
- End-to-End RL Drone Policy Passes Sim2real on Multiple Hardware — yacineMTB · 2026-08-27