Tweaking prompts and state to make a model act from vision alone
mervenoyann · x · 2026-07-23
The author says they tweaked the prompt and added a few extra states, but still expects the model to complete the task purely from vision.
In a reply, they point Ben Burtenshaw to OpenEnv and suggest building the environment yourself, implying this is about setting up and testing vision-based agent workflows rather than just talking about model capability.
More from coding & agent
- LangChain and Cognition will host a meetup on open memory for agents — LangChain · 2026-07-23
- Factory Co-founder Predicts 90% of Coding Agent Tokens Will Be Fully Autonomous in 12-24 Months — matanSF · 2026-07-23
- Voice assistant tool calls sped up instantly after moving the backend to Europe — ur_piyo_a_hoe · 2026-07-23
- W&B’s Scott Condron wants to push research agents, trace insights, and marimo eval UIs — _ScottCondron · 2026-07-23
- Clean Plate LoRA examples show a practical video-cleanup workflow — nazihater3000 · 2026-07-23
- DeepWiki Tested: Auto-Analyzing Three.js Game Agent Skills System — majidmanzarpour · 2026-07-23