Schema Harness Claims High Scores on ARC-AGI-3
TFenrir · reddit · 2026-07-17
This post introduces a harness called Schema, which, when paired with Fable+4.8 or GPT 5.6 Sol, reportedly achieved scores of 99% and 95.35% on ARC-AGI-3, respectively.
A link to the project homepage is included, clarifying that this is an LLM-centric evaluation/execution framework rather than just a benchmark result for a single model. Key takeaways include:
- Schema is a harness designed for LLMs.
- It claims exceptionally high scores on ARC-AGI-3.
- The results stem from comparisons of different model combinations.
This serves better as research/evaluation material for those interested in benchmarks and execution frameworks.
Related event: Schema Harness Sparks ARC-AGI-3 Debate(14 posts)→
More from coding & agent
- Chaining dependent MCP tool calls: no rollback, duplicate risk — agentrsdg · 2026-09-11
- DeepMind-led paper makes design docs the source of truth, code disposable — SMART regenerates in 1.5-3h for ~$100 — Roger_M_Taylor · 2026-09-11
- Agent-built classifier labels 192k docs for $0.70 vs $13-26 with frontier LLMs — vanstriendaniel · 2026-09-11
- MathModelAgent gains traction: auto-solves math modeling and writes a submission-ready paper — jihe520 · 2026-09-11
- alphaXiv open-sources OpenResearch to run parallel research agents with any model — alphaXiv · 2026-09-11
- DeskcommCRM: open-source AI sales CRM with native agents and WhatsApp hits 1k stars — melgarafael · 2026-09-11