Benchmark: AI models underperform simple greedy algorithms in retail simulation
ycombinator · x · 2026-08-28
ShelfLife E-Sim is a long-horizon resource management benchmark based on real retail data. Experiments had frontier AI models run a retail business for 120 days; most earned less money than a simple greedy algorithm, highlighting limitations in complex long-term planning.
More from Models
- ChatGPT fails to generate anatomically correct human motion diagrams — kaljakin · 2026-08-28
- OpenAI reportedly running many pre-trains; far-future model codenamed 'Bel' — haider1 · 2026-08-28
- Yutori n2 launches on Crusoe Cloud for low-cost computer use — DhruvBatra_ · 2026-08-28
- Epoch Had GPT-5.6 Sol Play Slay the Spire: Tactics, Not Computer Use, Is the Bottleneck — Jsevillamol · 2026-08-28
- GPT-5.6 Sol reverses engineers 32-bit iOS games in an afternoon — gpt2chatbot · 2026-08-28
- Qwen Flash outperforms DeepSeek and GLM 5.3 Flash — bindureddy · 2026-08-28