LLMs Suggest Useful Next Steps Only 10% of the Time in Experiments
giffmana · x · 2026-08-03
The author has been using LLMs alongside experiments to review results and suggest next steps. They found that only 10% of the time does the model suggest exactly what they were thinking; 90% of the time, it offers complete garbage microtuning advice, like tweaking the random seed.
More from Models
- Why DeepSeek's API Is So Cheap: Tiny Model Size Boosts Single-Chip Throughput — AravSrinivas · 2026-08-03
- User Complains AI Model Panders and Lies: Poor Experience Becomes Widespread Issue — velvet32 · 2026-08-03
- Qwen Releases Qwen-CUA: A Native Computer-Use Agent — xhluca · 2026-08-03
- AI Misinterprets User Intent, Optimizes Prompt and Solves Problem — PMinervini · 2026-08-03
- Dev slams Anthropic for restricting agent autonomy, fearing it hurts enterprise API use — DavidBennett__ · 2026-08-03
- OpenAI Pivots to Codex and Reclaims the Lead from Claude — iruletheworldmo · 2026-08-03