Astra ran 21 hours optimizing but botched the dataset: labeled too few, never followed up
ivan_bezdomny · x · 2026-09-07
ivanbezdomny let Astra run 21 hours building a grammar scorer from scratch, and is cooling on it after reviewing its dataset work: it labeled too few examples despite 10,000 options; labeled with GPT-5.5-high, saw GPT-5.6-sol was better but didn't follow up; noticed variance between labeling runs but never tried multi-agent voting or improved the labeling prompt. Only after probing questions did it concede its scoring was shaky.
More from coding & agent
- Watcher's agents started talking to each other to avoid stepping on each other's toes — ccerrato147 · 2026-09-07
- Stripe launches Machine Payments so AI agents can pay for APIs programmatically — jeff_weinstein · 2026-09-07
- GEPA isn't just an RL alternative: new TRL + OpenEnv project trains agents — lateinteraction · 2026-09-07
- AI agent dev warns: run agents in a cloud VM, never on your personal computer — blelbach · 2026-09-07
- Eden turns into an MCP server: list your Eve agents without forking or PRs — evilrabbit_ · 2026-09-07
- Reddit essay argues one-shot AI coding demos are huge token waste with poor decisions — rsvallen · 2026-09-07