RL inside the harnesses: lifting LFM2.5 from 42% to 54% across four agent harnesses
_lewtun · x · 2026-10-02
A new HF technical deep dive tackles why the same model behaves wildly differently across agent harnesses — LFM2.5-2.6B solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code before training. Since the harness controls the agent loop, trainers never see the model's exact tokens; the team solved this with a capture proxy in OpenEnv (recording tokens + logprobs), Harbor for tasks/sandboxes, and TRL's async GRPO trainer. Training across four harnesses lifted held-out performance from 42% to 54%, with gains in every harness and 31% fewer tool calls.
Related event: Multi-Harness RL Boosts LFM2.5 from 42% to 54%(2 posts)→
More from coding & agent
- Jev pitches 'decision primitives': models plugging into logic without text — hardimanjames · 2026-10-02
- LangChain's Sproul: agent core pattern unchanged for a year, "we've been at AGI for four months" — BraceSproul · 2026-10-02
- n8n integrates typesafeai's Jev model as a smarter If/Switch for workflows — hardimanjames · 2026-10-02
- Firecrawl launches People Enrichment Pack so agents can find buyers, talent and company data — devdigest · 2026-10-02
- Stanford launches CS 224V, an Agentic AI course tackling agent reliability with RAG and formal methods — stanfordnlp · 2026-10-02
- Early Hands-On: OpenAI's Dots Agent Impresses With Speed and First-Try Accuracy — billyjhowell · 2026-10-02