Early GLM video-input experiments: screen recordings for UI bugs, voice still weak
DevDminGod · x · 2026-09-16
A developer shares early usage ideas for GLM's video input: recording the screen and having the model spot UI bugs works well, but voice recognition is weak—converting spoken notes to text is a better route, since the model mainly relies on the visual track.
More from coding & agent
- Turn Muse agent into a content studio that writes viral Instagram scripts weekly — thederbiedone · 2026-09-16
- free-claude-code: open-source project with 55k stars runs Claude Code on free models from 50 providers — goyalshaliniuk · 2026-09-16
- Gemini CLI ships v0.62.0 nightly with AgentLoopContext and MCP display fixes — gemini-cli-robot · 2026-09-16
- Agent Flywheel guide: exhaustive markdown plans, beads tasks and agent swarms to cut token waste — doodlestein · 2026-09-16
- A tiered approach to token budgeting when developing with agent swarms — doodlestein · 2026-09-16
- Doctor builds WebMCP-enabled DICOM viewer so Codex can read scans and write reports — iamrobotbear · 2026-09-16