No training needed: strong model's harness lifts GPT-5.4-mini from 0.488 to 0.912

jiqizhixin · x · 2026-09-08

UIUC and Salesforce AI present AI4AI at test-time: instead of training or adding data, a strong model designs a harness—a work system—that lets a small model perform far above its native level.

The paper reframes capability transfer as organization, not distillation. The model itself never changes; tasks are auto-classified, error-prone parts go to a calculator, long-term memory lives in tables, fixed rules become programs, and checklists prevent format errors.

In the main experiment, GPT-5.4-mini jumps from 0.488 average accuracy to 0.912 under the strong-model-designed harness—nearly doubling without touching its parameters.

Original post →

More from Research

Research channel →