Same model, three harness fixes: how a local 35B browser agent went from wrong to right on Mind2Web
Stunning-Sherbet1853 · reddit · 2026-10-05
The author ran Online-Mind2Web (300 tasks, 136 live sites) with a local 35B model (Ornith 1.5) on a single 3090 and fixed a failing TED-talk retrieval task without touching the model.
- Run 1: the agent searched "robots" and browsed results; the correct answer was a quadcopter talk with no "robot" in the title, so search could never surface it.
- Run 2: adding a checklist item ("did you use the site's own sort/filter?") got it to the topic page, but it sorted by newest and skipped the duration filter — and the loose grading still passed it.
- Run 3: tightening the item to "applied everything the request needs" and moving the hint to the first tool result made the agent apply most-viewed sort plus the 12–18 minute filter. It even noted "strictly speaking this is a drone talk, not a robot one."
Key lesson: when the easy path looks like it's working, the model won't leave it — a "use filters first" prompt line did nothing, but a checklist item plus a nudge in the tool result did.
On 30 fresh tasks: 6/10 passed without a given start site (excluding 5 bot checks); 5/14 with the official setup, and 6 of 9 misses were UI manipulation issues (autocomplete, select boxes, cookie overlays, search boxes). Research-style tasks mostly work; UI manipulation is where agents lose. Grading is self-built, not WebJudge, and the sample size is small.
More from coding & agent
- Apple quietly shipped MagSafe 3 cable firmware 3.2.0; GLM used to diff the changes — steipete · 2026-10-05
- Databricks-backed Omnigent hits 10.5k GitHub stars: open-source meta-harness to swap agent harnesses without rewrites — bibryam · 2026-10-05
- He rebuilt his personal site with an AI agent in 5 minutes, migrating 83 essays and saving $120/year — thisiskp_ · 2026-10-05
- Agent policy-boundary starter updated with idempotency, audit logging and independent validation — Dapper-Roof2370 · 2026-10-05
- Biggest multi-agent system win was making the agents replaceable, not central — Druss_ · 2026-10-05
- Anthropic exec hits inbox zero for the first time by making it an explicit goal for his Claude agent — BorisMPower · 2026-10-05