A multimodal “task” prompt for agents bundles voice, screen, marks, and text
omarsar0 · x · 2026-07-23
Karpathy’s idea of long voice prompts is extended into a broader agent workflow: use a task as the unit of prompting.
- A task bundles spoken intent, the current screen, annotations, and exact text into one turn.
- The agent reconstructs intent from the combined signals, which reduces correction loops and makes larger handoffs possible.
- The attached diagram frames this as capturing context, forming a task representation, and reconstructing intent from multimodal evidence.
More from coding & agent
- Non-devs should buy Claude Code or Codex themselves, says a reposted tip — HankYeomans · 2026-07-23
- LangChain says the real agent problem is loop engineering, not just execution — LangChain · 2026-07-23
- Teaser: Orchestrator System with Dynamic Multi-Model Routing and Task Splitting — omarsar0 · 2026-07-23
- Claude Code Autonomously Configures Stream Deck Keyboard and Generates Icons — SimonBalmain · 2026-07-23
- Reddit asks how to manage email identity for AI agents — BlakSavageGaming · 2026-07-23
- Developer Uses GPT to Curate Best Codex and ChatGPT Work Use Cases from X — nickbaumann_ · 2026-07-23