Fixing Lost Agent Context Skyrockets Model Evaluation Scores
rajistics · x · 2026-08-05
Both OpenAI and Fireworks AI recently discovered and fixed critical bugs in their respective agent testing harnesses that were dropping the models' reasoning context, leading to massive score improvements.
- OpenAI: Their ARC-AGI-3 harness was throwing away the model's reasoning between turns. Fixing this increased the score from 13.3% to 38.3%.
- Fireworks AI: While running CyberGym via OpenHands, they found issues with non-native tool calling, dropped reasoning, and un-forwarded parameters. Fixing these harness issues boosted the score from 40.7% to 70.0%.
More from coding & agent
- LangSmith Launches LLM Gateway for Production-Grade Agent Runtime Controls — LangChain · 2026-08-06
- Driving Codex with ChatGPT Voice: A Practical Workflow for Real-Time Coding — dfinke · 2026-08-05
- When a senior SWE finds four vibecoders stuck on localhost — venturetwins · 2026-08-05
- Tutorial: Building an Autonomous Content Engine with Hermes Multi-Agents — VibeMarketer_ · 2026-08-05
- Stanford Hazy Research: AI Agents Are Driving Traditional CUDA Abstractions Toward Retirement — HazyResearch · 2026-08-05
- Greptile v5 Released: Ground-Up Rewrite of Coding Agent Boosts Speed and Precision — garrytan · 2026-08-05