Tweaking prompts and state to make a model act from vision alone

mervenoyann · x · 2026-07-23

The author says they tweaked the prompt and added a few extra states, but still expects the model to complete the task purely from vision.

In a reply, they point Ben Burtenshaw to OpenEnv and suggest building the environment yourself, implying this is about setting up and testing vision-based agent workflows rather than just talking about model capability.

Original post →

More from coding & agent

coding & agent channel →