Same model, three harness fixes: how a local 35B browser agent went from wrong to right on Mind2Web

Stunning-Sherbet1853 · reddit · 2026-10-05

The author ran Online-Mind2Web (300 tasks, 136 live sites) with a local 35B model (Ornith 1.5) on a single 3090 and fixed a failing TED-talk retrieval task without touching the model.

Key lesson: when the easy path looks like it's working, the model won't leave it — a "use filters first" prompt line did nothing, but a checklist item plus a nudge in the tool result did.

On 30 fresh tasks: 6/10 passed without a given start site (excluding 5 bot checks); 5/14 with the official setup, and 6 of 9 misses were UI manipulation issues (autocomplete, select boxes, cookie overlays, search boxes). Research-style tasks mostly work; UI manipulation is where agents lose. Grading is self-built, not WebJudge, and the sample size is small.

Original post →

More from coding & agent

coding & agent channel →