40k-token tool outputs vs small agent contexts: Reddit weighs truncation, summarization and handles

No-Age-3362 · reddit · 2026-10-09

A Reddit practitioner asks how to handle tools that return far more than an agent can use — a search or DB read returning 40k tokens when the agent needs a few lines. Passing it all burns context and triggers lost-in-the-middle; cutting risks dropping the one crucial line. Three options are weighed: (1) hard truncation with a note so the agent can request more — cheap but blind; (2) pre-summarizing with a smaller model — but its mistakes are invisible to the agent; (3) storing the full result externally with a handle plus paging/search tools — cleanest but adds tools and turns. The thread invites real-world war stories on where each approach breaks.

Original post →

More from coding & agent

coding & agent channel →