STS2-Bench: A Benchmark for Long-Term Decisions
MrBoxChen · reddit · 2026-07-12
The author proposed a new benchmark, **STS2-Bench**, using *Slay the Spire 2* to test models' long-term decision-making capabilities. ### Core Concept Instead of answering isolated questions, the model must continuously make decisions throughout an entire game: - Read constantly changing states - Balance short-term gains with long-term survival - Choose cards, routes, and resources under uncertainty - Adjust strategies based on outcomes after each round ### Experiments and Findings - The author evaluated 7 model configurations within the same framework - **5.6Sol** performed surprisingly well on this sequential decision-making task - The author explicitly stated that this shouldn't be viewed as a general intelligence ranking, but it might reveal signals hidden from traditional single-turn benchmarks ### Feedback Requested - Whether this game task is suitable as a proxy metric for planning and long-range reasoning - What control variables or baselines are needed to make the evaluation more credible - What other games or environments would be suitable for this type of benchmark
Related event: STS2-Bench Uses Slay the Spire 2 to Test Long-Horizon Decisions(2 posts)→
More from Research
- LTX 2.3 LoRA demo changes a video’s camera angle — CQDSN · 2026-07-21
- OpenForecaster uses daily news to improve language-model forecasting — Cohere_Labs · 2026-07-21
- SenseTime unveils U1 Pro and open-sources a 50M-sample vision dataset at WAIC 2026 — 机器之心 · 2026-07-21
- Baseten study finds new facts in LLM weights are fragile unless trained from many restatements — alex_verem · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21
- uv-scripts/ocr returns to the top of Hugging Face datasets with a JSON model picker — vanstriendaniel · 2026-07-21