Dev calls for a benchmark measuring the influence of AGENTS.md across different models
johnlindquist · x · 2026-09-20
X user johnlindquist observes that more than ever, each model needs to be handled differently, and he's starting to feel he's losing control. He wishes someone would create a benchmark for "the influence of AGENTS.md" on each model—quantifying how the same agent config file affects different models' behavior differently.
More from coding & agent
- Swapping a 13s pipeline step for a 200ms call saves thousands per month — hardimanjames · 2026-09-20
- Opus 5-built Three.js WebGPU waves deliver stunning real-time shoreline in browser — majidmanzarpour · 2026-09-20
- Jev, a 'System One' model that only makes decisions, questions how many LLM calls agents really need — ThePromptIndex · 2026-09-20
- Bend 2 called a 'software factory language': formal verification ends code review — airesearch12 · 2026-09-20
- Claude authored all 3D models, animations and SFX in a web game after 50 iterations — Sheru7000 · 2026-09-20
- Flask creator Armin Ronacher asks: what struggles most with AI in software engineering? — mitsuhiko · 2026-09-20