After visiting Anthropic: harness matters more than the model—until models absorb it
KerenKoshman · x · 2026-10-04
Keren Koshman shares takeaways from visiting Anthropic in SF. She had increasingly believed the harness layer matters more than the model—Berkeley research shows harness affects quality and cost up to 70% more. But Anthropic argues today's harness engineering resembles prompt engineering 18 months ago: a short-lived profession before models started writing the best prompts themselves. Memory, context, and tool layers will likewise be absorbed into models. She left rethinking her view, noting Anthropic sees at least six months ahead.
More from coding & agent
- Meta's RankEvolve doubles down on agent reliability, lifting auto-research accuracy 45.8% to 62.5% — dair_ai · 2026-10-04
- CMU paper: RL-trained 4B proposer edits agent harness code, beats its 35B teacher — omarsar0 · 2026-10-04
- Dev praises pi-durable for running long-lived agents behind any interface — tokumin · 2026-10-04
- Claude Code 2.1.289 prompt tokens jump 17.8%, system prompts now over half — ClaudeCodeLog · 2026-10-04
- Claude Code 2.1.289 patches @-mention symlink bypass of file read deny rules — ClaudeCodeLog · 2026-10-04
- Claude Code 2.1.289 ships agent.spawn and fixes @-mention permission bypass — ClaudeCodeLog · 2026-10-04