Schema Harness Sparks ARC-AGI-3 Debate

From July 16 to 18, discussion clustered around an agent harness called Schema. Across posts by naval, TFenrir, we_are_mammals, and 机器之心, it was described as a system that leaves model weights untouched but rewrites the outer loop of observation, hypothesis formation, experimentation, revision, and execution. The reason it drew so much attention is the claimed near-saturated ARC-AGI-3 Public set score, which quickly shifted discussion toward whether the main lever is now harness design rather than one-shot model ability.

Reported results

The most repeated number in the cluster is a claimed 99% on the ARC-AGI-3 Public set. TFenrir said Schema reached 99% with Fable+4.8 and 95.35% with GPT 5.6 Sol. Some reposts relayed a different high-scoring setup as Opus 4.8 + Fable 5 reaching 99%, so the exact model pairing is not fully consistent across posts. Separately, Greg Kamradt highlighted a community result of 78.4% across 25 public games, completing 160/183 levels, while saying the current best frontier LLM in a single session scores only 4.5%; in the materials here, that serves mainly as context for how much outer-loop methods may matter.

Method ideas being emphasized

Several posts converged on a similar interpretation of Schema's approach. we_are_mammals summarized it as changing observation modeling, historical verification, and planning/execution-revision flows without retraining the base model. naval relayed the framing that the harness can play games, write code, and reason like a physicist. Gary Marcus passed along the view that the key move is to turn experience into a programmatic world model and then validate against full history. xiuyu_l's reposts further stressed making the latent world representation a program rather than a vector, injecting analysis-by-synthesis, and encoding representation finding directly into the harness. Greg Kamradt also relayed a practical policy: on unseen tasks, the agent should first spend actions learning rules and constraints instead of rushing the objective.

What remains unclear

Most posts repeated the high scores, but they did not spell out the full evaluation setup, replication details, or error bounds. The reposts also disagree on some model/version combinations attached to the top result. Based on the posts alone, the clearest conclusion is not that the benchmark has been definitively solved, but that ARC-AGI-3 research is rapidly shifting toward custom harnesses, full-history verification, representation discovery, and action efficiency as key system-design variables.

2026-07-16 ~ 2026-07-18 · 14 related posts