Auto Research Agent Night Training Recap
richdotca · x · 2026-07-14
This thread shares how the author used autofrontier for real training: an automated agent ran nightly training to chase a record on a visual grounding task.
Key details:
- Task: referring-expression grounding, metric: RefCOCO [email protected]
- An 8B model trying to beat a 72B model's record
- Improved from 0.850 to 0.920, briefly thought 'won'
- Later ran to 0.95 but couldn't replicate under stricter check, author rejected the 'false breakthrough'
Overall: Auto research agents can uncover new directions, but must pass reproducibility checks — one-off scores aren't enough.
More from coding & agent
- Devin adds e2b sandboxes for remote agent execution — badphilosopher · 2026-07-22
- Hermes Agent Refactoring Proposal: Decoupling via Event Bus and Monorepo Slicing — Promptmethus · 2026-07-22
- ty now reads Pydantic config keywords and field metadata — charliermarsh · 2026-07-22
- Pensar Launches AI Security Agent to Autonomously Discover and Patch 0-Days — andriy_mulyar · 2026-07-22
- ty adds first-class Pydantic support, including strict and lax field handling — charliermarsh · 2026-07-22
- Google launches Gemini 3.5 Flash Cyber for CodeMender, with limited access for governments — GoogleAI · 2026-07-22