Astra ran 21 hours optimizing but botched the dataset: labeled too few, never followed up

ivan_bezdomny · x · 2026-09-07

ivanbezdomny let Astra run 21 hours building a grammar scorer from scratch, and is cooling on it after reviewing its dataset work: it labeled too few examples despite 10,000 options; labeled with GPT-5.5-high, saw GPT-5.6-sol was better but didn't follow up; noticed variance between labeling runs but never tried multi-agent voting or improved the labeling prompt. Only after probing questions did it concede its scoring was shaky.

Original post →

More from coding & agent

coding & agent channel →