lateinteraction: 8 hours of frontier-model 'Ultra' work still collapses like autocomplete
lateinteraction · x · 2026-09-20
lateinteraction (Charles Packer) describes the agony of extracting merely OK work from frontier models at "Ultra" settings: he now provides more detailed context and feedback than he has ever given a human — and giving feedback is literally his job. His thesis: if you're particular about what you want in a high-dimensional way, you'll quickly lose your mind over how brittle and autocomplete-like even 8 continuous hours of Astra or Fable work remains. Don't be distracted by sparse unimodal verifiable successes, he says — they predict nothing about the seemingly inherent brittleness/collapse on high-dimensional tasks.
Related event: Researcher's rant on brittle frontier models sparks resonance(5 posts)→
More from coding & agent
- LangChain's Jev evaluator cuts agent eval score variance by up to 913x at 1/80th the cost — LangChain · 2026-09-20
- Open-source NL logic interpreter unifies facts with Jev, queries cost a quarter cent — narphorium · 2026-09-20
- Dev Integrates Jev into Codex to Speed Up Browser Use, Hundreds of Trials Cost $0.00725 — TheZachMueller · 2026-09-20
- Annotating CUA agent data: tasks go stale, so pseudo-annotate the tasks themselves — mervenoyann · 2026-09-20
- Pi coding agent ships 0.86.0 with mid-conversation system messages and dynamic tools — mitsuhiko · 2026-09-20
- Dev ports all his dither shaders into Leafalia and they actually work — eschadiol · 2026-09-20