Amazon runs Karpathy's AutoResearch at production scale for 12 weeks, finds 5 failure modes
amazon · hf · 2026-10-09
Amazon applied Karpathy's AutoResearch paradigm — an LLM iteratively editing a training script and keeping changes that improve a held-out metric — to production embedding optimization for book recommendations, running 220+ experiments over 12 weeks.
Key findings:
- Five recurring failure modes absent from the original setting: infrastructure fragility, agent memory decay, search-direction stagnation, iteration-cost asymmetry, and metric fixation
- A prevent-persist-redirect scaffolding design maps each failure mode to a structural remedy
- Results: 1.82x Recall@6 lift and 2.1x coherence lift over hand-tuned baselines; the agent autonomously designed a text-only fallback expanding catalog coverage 5.8x
- Two systems spanning 3 orders of magnitude in iteration cost showed the same failure modes, suggesting structural properties of production-scale autonomous research
More from coding & agent
- Claude Opus 5 nearly triples Qwen's SWE-bench score in open-source RSIGym auto-research env — rohanpaul_ai · 2026-10-09
- Khan Academy launches MCP server letting AI assistants read courses and transcripts key-free — modelcontextprotocol · 2026-10-09
- BAAI's AREX research agent checks answers requirement-by-requirement, hits 82.5% BrowseComp — DeepLearningAI · 2026-10-09
- His personal AI agent now auto-compares construction bids in his inbox — dkundel · 2026-10-09
- His 16-Month-Old Call: Claude Code Shifts Agent Workflows and API Spend to Anthropic — majidmanzarpour · 2026-10-09
- pydantic-ai-go brings Pydantic AI's typed agent loop to Go with ~20 providers — samuelcolvin · 2026-10-09