Strong Models Scaffolding Weak Ones: GPT-5.4-mini Jumps From 0.488 to 0.912

青稞AI · wechat · 2026-08-31

A WeChat account analyzes a Salesforce AI / UIUC paper (arXiv:2608.12307) on Strong-to-Weak Scaffolding: instead of answering for the weak model, a stronger Builder AI designs an external Harness in a programmable environment—prompts, Python code, rules, state tracking and verification—optimizing the solving path rather than model weights.

On four Theory-of-Mind benchmarks, GPT-5.4-mini's bare average of 0.488 rises to 0.763 (best 0.912) with an auto-built Harness, beating the stronger GPT-5.4 without one (0.619); all 57 automated scaffold runs beat baseline, though still below a hand-crafted UserHarness (0.939).

Key findings: the best Harness doesn't make the weak model think longer—it offloads deterministic work to code and explicit state, transferring cognitive structure rather than knowledge. Compilability has limits (94% of BigToM vs 36% of MuMA-ToM), stronger targets gain less (sometimes regressing), and Builder quality correlates with reasoning effort, not iteration count.

Related event: Strong-to-Weak Scaffolding Nearly Doubles Small Model Accuracy(2 posts)→

Original post →

More from coding & agent

coding & agent channel →