Meituan Open-Sources CAST: Using Game Solvers as Turn-Level Teachers for LLMs
meituan-longcat · hf · 2026-07-30
Meituan has open-sourced CAST (Credit Assignment from Solver Teachers), a framework designed to address the sparse reward problem in Reinforcement Learning with Verifiable Rewards (RLVR) for Large Language Models (LLMs) in long-horizon decision-making tasks.
- Core Insight: By tracking changes in a game solver's state value, CAST determines whether an action advances toward success, converting these value changes into turn-level reward signals. This is equivalent to on-policy distillation from the solver while requiring only scalar values.
- Performance: Across Sokoban, Minesweeper, and Rush Hour, CAST outperforms all trained baselines under both in-domain and unseen-difficulty evaluations. It also achieves the highest average zero-shot performance on ALFWorld and WebShop.
- Availability: The code is available on GitHub.
More from coding & agent
- OpenAI Responses API to Support Auto-Compaction, Reshaping Long Context Workflows — GregKamradt · 2026-07-30
- Developer Envisions Multi-Agent Pair Programming with Sol and Fable — shakoistsLog · 2026-07-30
- ScalableRAG: High-Quality Retrieval at Zero Ingestion Cost Without Vector DBs — _reachsumit · 2026-07-30
- Volcengine Open-Sources OpenViking: Unified Agent Memory & Skills — sujingshen · 2026-07-30
- Build a Pixel Art Poker Game on iOS in Minutes with Grok Build — DeryaTR_ · 2026-07-30
- 360 Launches 'Nano Work': An Enterprise-Grade Agentic Work Platform — 新智元 · 2026-07-30