GPT-6 Astra in action: video-to-HTML5 pipeline shows true multimodal planning
techhalla · x · 2026-09-04
Developer @literallydenis shares hands-on observations of GPT-6 Astra:
- Favorite pipeline: video → Whisper transcription → Codex + GPT-6 → HTML5 web page.
- Genuinely multimodal thinking: the model plans how the 3D result will look while writing the code—an unprecedented ability.
- Output quality: he's never seen such strong results in SVG, 3D models, and sound design.
- Companion demo: an old steam train drawing reconstructed in Blender into 3,295 editable detailed objects, with prompt-controlled detail levels.
His verdict: Astra is a clear step up on end-to-end multimodal-to-deliverable tasks.
Related event: Developers Test GPT-6 Astra: 3D Generation and Multimodal Coding Shine(2 posts)→
More from coding & agent
- Sakana AI's Stefania Druga demos agentic memory experiments for edge devices — kaixhin · 2026-09-04
- rybbit-mcp lets Claude Code query Rybbit Analytics via natural language — modelcontextprotocol · 2026-09-04
- Agentic Coding's Biggest Flaw: Managing AI's Pointless Code Rewrites — kylegawley · 2026-09-04
- Anthropic reveals 3 Claude sandbox escapes, one touched a production database — Sumsub_Insights · 2026-09-04
- Subagents when the API key hits the rate limit — realsohamparekh · 2026-09-04
- Yoav Goldberg pushes back on DHH: an agent-built Qt markdown editor isn't a path to personal Photoshop — yoavgo · 2026-09-04