CrEST framework solves credit assignment in multi-turn agent training
burny_tech · x · 2026-08-17
A tweet highlights a new paper proposing the CrEST framework for multi-turn AI agent training.
- Problem: In multi-turn training, pure RLVR is accurate but struggles with credit assignment across complex turns; self-distillation offers dense feedback but caps potential and wastes compute on formatting tokens.
- CrEST Solution:
- Allows a verifier to dictate update direction while a self-teacher strictly controls magnitude.
- Uses Turn-Segmented Rewards to isolate the exact turn where errors occur.
- Introduces an Entropy Gate to force updates onto critical decision tokens rather than predictable syntax tokens.
- Results: Shows significant gains on benchmarks like WildToolBench and BFCL V3 for long-trajectory agent sessions.
More from Research
- Cohere Labs session probes scaling law reliability and offers a research checklist — Cohere_Labs · 2026-09-23
- SoL-Pi paper: auto-research loops save $4-13 per hour on coding agents — alex_verem · 2026-09-23
- NVIDIA's SoL-Pi lets AI rewrite agent harnesses, cutting tokens 44.7-49% — alex_verem · 2026-09-23
- Slingshot RL framework jailbreaks Qwen2.5-32B at 67% success, transfers zero-shot to Gemini 2.5 Flash — j_foerst · 2026-09-23
- TLAPS-Bench launches in alpha to test whether AI agents can formally prove system correctness — tianyin_xu · 2026-09-23
- ImIR replaces text prompts with image instructions for all-in-one restoration — Süleyman Aslan · 2026-09-23