STS2-Bench: Testing Long-Range Decision Making
MrBoxChen · reddit · 2026-07-13
This post introduces STS2-Bench, a benchmark evaluating the long-range decision-making capabilities of LLMs using the game Slay the Spire 2.
Evaluation Design
- Models must make decisions in constantly changing game states: choosing cards, selecting routes, allocating resources, and balancing short-term gains against long-term survival.
- The benchmark emphasizes irreversible decisions in real matches, with no save scumming (nosl) or future previews allowed.
- The author used the same random seeds for 7 model configurations to ensure consistent scoring across full runs.
Key Results
- Across 3 seeds × 7 models, all 21 runs ended in failure; no model managed to beat the game.
- The best performance reached Act 3; overall, Floor 17 (the Act 1 boss) acted as a collective "wall of death."
- Sol demonstrated the most stability in terms of median score, average ranking, and raw game score.
- Massive cost differences were observed: DeepSeek cost only $0.33 for three runs, whereas Opus cost around $91; the total 21 runs consumed roughly 470 million tokens and $336.
Takeaway
Such long-range, stochastic, and irreversible environments can expose capability gaps that single-turn benchmarks miss, particularly regarding a model's stability in continuous decision-making.
Related event: STS2-Bench Uses Slay the Spire 2 to Test Long-Horizon Decisions(2 posts)→
More from Models
- Gary Marcus says LLM math skills are like knowing only a car’s engine size — GaryMarcus · 2026-07-22
- OpenAI’s Codex + GPT-5.6 Sol hits 99% recall in Project APE verification tests — soumitrashukla9 · 2026-07-22
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22
- Nanbeige4.2-3B launches as a 3B Looped Transformer model that beats larger baselines — Wooden-Deer-1276 · 2026-07-22
- A post says six companies now beat Google’s best LLM, including two open-source models — soham_btw · 2026-07-22