GPT-OSS 20B Ran My Personal Agent for a Week: 312 Tasks, 97.4% First-Shot Tool Calls, Zero Frontier APIs
rasheed106 · reddit · 2026-08-25
A developer ran a bold experiment: taking an open-source assistant harness, stripping out every cloud API call, and making GPT-OSS 20B responsible for the entire agent loop, running 7 days on an M5 MacBook Pro from a compiled 12GB binary.
Results:
- 312 real tasks, 1,847 tool calls
- 93.2% completed without human takeover
- 97.4% first-attempt schema-valid tool calls
- 71 multi-step workflows, 4.1% retry rate
He started the week trying to find where GPT-OSS 20B would fail — and ended it cancelling his Perplexity Computer subscription.
More from coding & agent
- Red-team study finds AI agents deleting entire inboxes and leaking sensitive data to protect secrets — alex_verem · 2026-08-25
- Dev uses agents to rewrite Terraria in C++ as a multi-agent RL environment — jsuarez · 2026-08-25
- Veteran engineer: AI-written storage code gets basic S3 semantics wrong — andersonbcdefg · 2026-08-25
- OnceMesh Open Source System Safely Reuses Exact LLM Agent Work — Critical_Molasses844 · 2026-08-25
- Agent Tether: Mining Personalized Content with Preference Data — tokenbender · 2026-08-25
- Alibaba treats agent context management as a programming problem — omarsar0 · 2026-08-25