STS2-Bench: Testing Long-Range Decision Making
MrBoxChen · reddit · 2026-07-13
This post introduces STS2-Bench, a benchmark evaluating the long-range decision-making capabilities of LLMs using the game Slay the Spire 2.
Evaluation Design
- Models must make decisions in constantly changing game states: choosing cards, selecting routes, allocating resources, and balancing short-term gains against long-term survival.
- The benchmark emphasizes irreversible decisions in real matches, with no save scumming (nosl) or future previews allowed.
- The author used the same random seeds for 7 model configurations to ensure consistent scoring across full runs.
Key Results
- Across 3 seeds × 7 models, all 21 runs ended in failure; no model managed to beat the game.
- The best performance reached Act 3; overall, Floor 17 (the Act 1 boss) acted as a collective "wall of death."
- Sol demonstrated the most stability in terms of median score, average ranking, and raw game score.
- Massive cost differences were observed: DeepSeek cost only $0.33 for three runs, whereas Opus cost around $91; the total 21 runs consumed roughly 470 million tokens and $336.
Takeaway
Such long-range, stochastic, and irreversible environments can expose capability gaps that single-turn benchmarks miss, particularly regarding a model's stability in continuous decision-making.
Related event: STS2-Bench Uses Slay the Spire 2 to Test Long-Horizon Decisions(2 posts)→
More from Models
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11