TokenCast forecasts agent token spend mid-execution, cutting budget use by 21.3%
SoutheastU · hf · 2026-09-29
Southeast University's TokenCast predicts total token consumption of LLM agent tasks, which can vary by over an order of magnitude across runs of the same task.
- Learns a composable cost representation per execution segment covering both its own consumption and context growth; composing segments captures re-reading costs of earlier context in later calls
- Forecasts refresh as execution unfolds with no extra LLM calls, at 32.8 ms mean cumulative prediction time on SWE-bench Verified
- Reduces MAE vs the strongest comparator by 14.5% on average across 96 combinations (4 suites × 6 agent models)
- In offline budget-control replay, uses 21.3% fewer tokens than fixed-budget policy at matched completion
Code: github.com/DEFENSE-SEU/TokenCast.
More from coding & agent
- Noah Shunn, 23: Reflexion Author Who Beat GPT-4 on HumanEval, Now Agent Pioneer at Sierra — vaibhavbetter · 2026-09-29
- OpenAI Codex CLI 0.159.0 ships instant interrupt, Mermaid rendering and Windows fixes — github-actions[bot] · 2026-09-29
- Dev switches from Cursor+Grok 4.7 to Claude Code+Opus 5.5 for a full day — jonathan_wilke · 2026-09-29
- Vercel migrates Mercedes-AMG F1 site: ~70% faster builds, ~75% faster paints in under a week — shuding · 2026-09-29
- "We used to use TAB to autocomplete AI-generated code" — a meme about how far coding agents have come — tlakomy · 2026-09-29
- IGSD Uses Environment-Verified Hindsight Self-Distillation to Train Search Agents — _reachsumit · 2026-09-29