Hardcore eval: Fixing gpt-oss harness and testing 320k cases
skeole · reddit · 2026-09-01
The author built a custom inference harness to fix tool calling issues in gpt-oss-20b and ran 320,192 evaluations over 1,062 GPU hours on a single RTX 3090, processing 3.49B tokens. Key findings include:
- Reasoning quality varies with effort: Same token counts yield different reasoning quality; 'Medium' effort achieved 97.1% accuracy vs 38.3% for 'Low'.
- Preserving reasoning isn't a silver bullet: It stabilizes variance but only slightly boosts accuracy and can hurt performance on tasks needing specific formatting.
Code, data, and the fixed template are open-sourced.
More from Research
- Stop building remote-controlled robots; sim2real is the key to success — _Stocko_ · 2026-09-02
- CS Professor shares the evolution of their paper-reading stack over the years — CSProfKGD · 2026-09-02
- OpenBind preprint released with new virtual screening benchmark — MoAlQuraishi · 2026-09-02
- Anima Anandkumar launches acceleratedU: neural operators over transformers for physical prediction — AnimaAnandkumar · 2026-09-02
- New blog 'Proofs and Prompts' explores how AI is reshaping mathematics — giannis_daras · 2026-09-02
- EMNLP 2026 Paper: SCALE uses structured case law for legal reasoning — ponguru · 2026-09-02