kalomaze: user interactions get farmed for general failure classes, not specifics
kalomaze · x · 2026-09-09
kalomaze argues the value of farming real user interactions in RL environments isn't specific user knowledge, but recognizing broader genres of under-covered failure — how attempts at general task families tend to fail — rather than benchmaxxing on individual user problems.
More from Research
- ValsAI Launches Tax Agent Bench: 391 Expert Questions to Test LLMs on Corporate Tax — JenniferHli · 2026-09-09
- Noam Brown: lab results show scaling agent swarm size improves performance — possibly a new scaling law — Afinetheorem · 2026-09-09
- OpenAI Says Its New Model Cracked the Navier-Stokes Millennium Problem in 88 Hours — nordicinst · 2026-09-09
- AI Intern reproduces ICML best paper "How much do LLMs memorize" for under $8 — _akhaliq · 2026-09-09
- Buckmaster & Alpöge publish Euler (112pp), Boussinesq (76pp), porous media papers with Lean proofs — xennygrimmato_ · 2026-09-09
- Boltz model contributor named to MIT Technology Review's 2026 Innovators Under 35 — GabriCorso · 2026-09-09