40k-token tool outputs vs small agent contexts: Reddit weighs truncation, summarization and handles
No-Age-3362 · reddit · 2026-10-09
A Reddit practitioner asks how to handle tools that return far more than an agent can use — a search or DB read returning 40k tokens when the agent needs a few lines. Passing it all burns context and triggers lost-in-the-middle; cutting risks dropping the one crucial line. Three options are weighed: (1) hard truncation with a note so the agent can request more — cheap but blind; (2) pre-summarizing with a smaller model — but its mistakes are invisible to the agent; (3) storing the full result externally with a handle plus paging/search tools — cleanest but adds tools and turns. The thread invites real-world war stories on where each approach breaks.
More from coding & agent
- vLLM on K8s: GPU Faults Can Render a Node Unusable — Inference Is Stateful — tianyin_xu · 2026-10-09
- Addy Osmani: Agents Erode Engineers' Joy of Knowing — If You Only Pick, Never Conjure — addyosmani · 2026-10-09
- Relevance and instruction following: the third automated AI quality check — goyalshaliniuk · 2026-10-09
- 7 Quality Checks to Automate Before Your AI App Ships to Production — goyalshaliniuk · 2026-10-09
- Dev calls for a truly intelligent model selector in Codex/ChatGPT after GPT-5's router flops — flowersslop · 2026-10-09
- Claude Code's Web Search Queries Get Increasingly Unhinged as Projects Progress — jwt0625 · 2026-10-09