No training needed: strong model's harness lifts GPT-5.4-mini from 0.488 to 0.912
jiqizhixin · x · 2026-09-08
UIUC and Salesforce AI present AI4AI at test-time: instead of training or adding data, a strong model designs a harness—a work system—that lets a small model perform far above its native level.
The paper reframes capability transfer as organization, not distillation. The model itself never changes; tasks are auto-classified, error-prone parts go to a calculator, long-term memory lives in tables, fixed rules become programs, and checklists prevent format errors.
In the main experiment, GPT-5.4-mini jumps from 0.488 average accuracy to 0.912 under the strong-model-designed harness—nearly doubling without touching its parameters.
More from Research
- FactoSR: RL with Factorized 4D Objectives Boosts VLM Spatial Reasoning — HKUST-GZ2 · 2026-09-08
- Stanford paper: general coding agents beat hand-built data agents by up to 37 points — CShorten30 · 2026-09-08
- Dhenu Vision 1.0 hits 94.4% precision, beating every frontier model tested — DevDminGod · 2026-09-08
- NBER paper: automation erodes the meaning of work before it eliminates jobs — ArtificialOther · 2026-09-08
- Training a 441M-Param Text-to-Image Diffusion Model From Scratch on a Single Local GPU — ostrisai · 2026-09-08
- 417k-param RNN generates all 6,573 frames of Bad Apple from a single initial state — SEBADA321 · 2026-09-08