Open-source TrueForge harness matches Claude managed agents' accuracy with 63% fewer tokens

Background-Job-862 · reddit · 2026-09-03

TrueFoundry open-sourced TrueForge, a model-neutral general agent harness (MIT, OpenAI-compatible endpoints, MCP, subagents, approvals, sandbox integration), and benchmarked it on 14 DevRev Enterprise-Bench tasks with a blind judge:

The gains come from the agent loop itself: less carried context, compaction, fewer tool calls, keeping large outputs out of context. Caveats: no first-class tracing/eval, bring-your-own sandbox, lossy compaction. Benchmark methodology is published for reproduction.

Original post →

More from coding & agent

coding & agent channel →