Same model scores 62% vs 33% across agent harnesses; Hugging Face open-sources multi-harness RL guide
huggingface · x · 2026-10-02
Hugging Face 团队发现:同一模型、同一权重,在一个 agent harness 中得分 62%,在另一个只有 33%。他们发布了年度最实用的多 harness RL 指南,完全开源。
核心技巧:不动 harness 本身,把模型指向一个代理(proxy)。代理支持编码 agent 用的全部四种 API 格式(OpenAI Chat Completions、OpenAI Responses、Anthropic Messages、Gemini),记录 vLLM 采样出的精确 token id 和 logprobs,并在其上训练。Claude Code、Codex、OpenCode 一行代码都不用改。
结果:
- 同时在 4 个 harness 上训练,Liquid AI 的 LFM2.5-2.6B 从 42% 提升到 54%
- 因给「更少步数完成任务」小奖励,工具调用减少 31%
- 只在 OpenCode 上训练把其成绩从 34% 提到 58%,但多 harness 训练的模型提升更均衡
Related event: Hugging Face Releases Open Multi-Harness RL Guide(5 posts)→
More from coding & agent
- Anthropic engineer: future models will get much better at code deletion and simplification — simpsoka · 2026-10-03
- The Best AI Workflows Keep Friction Exactly Where Mistakes Matter — alifcoder · 2026-10-03
- Dev builds browser 3D game from scratch with Opus 5.5, Blender and Three.js — jason_mayes · 2026-10-03
- Cloudflare Durable Objects now survive client disconnects for long-running agents — threepointone · 2026-10-03
- Dev tired of nbviewer crashing builds serverless browser Jupyter notebook renderer, MIT-licensed — cneuralnetwork · 2026-10-03
- Monitoring Can't Keep Up: What to Auto-Detect When Your AI System Scales — goyalshaliniuk · 2026-10-03