Researchers debate whether a code harness, not the base model, is the real unit of comparison
srush_nlp · x · 2026-07-23
The discussion asks whether the real novelty is not just that harnesses help generalization, but that a harness treated like a model changes the comparison itself.
The reply argues:
- The base Transformer does not get tools, so the relevant baseline is harness-like programs vs. the base model.
- On some tasks, a code tool makes an obvious difference.
- On others, the benefit is less clear because chunking and acting on chunks is not unique to LM programs; in principle, model weights could be constructed to do something similar.
- The point being made is therefore about the role of the code/tool harness as part of the system, not just the underlying LM.
Related event: Researchers Debate Standardizing Code Tools for Base Transformers(2 posts)→
More from coding & agent
- Anthropic opens beta for Claude Code security plugin that scans code before commit — testingcatalog · 2026-07-23
- A deep dive on building agents that can work for days, not minutes — baseten · 2026-07-23
- Anthropic launches Claude Security beta for Claude Code — claudeai · 2026-07-23
- Claude Code gets a beta security-scanning plugin — claudeai · 2026-07-23
- Relay adds open-source multi-model routing for coding: plan, execute, review — MatthewBerman · 2026-07-23
- ASC CLI now covers the full App Store publishing flow, Game Center included — rudrank · 2026-07-23