Floatboat Harness Beats Flagship Models Using Low-Cost DeepSeek
机器之心 · wechat · 2026-08-08
Floatboat released evaluation data for its Agent Harness across five third-party benchmarks. The results show that using the highly cost-effective DeepSeek-V4-Flash model paired with the Floatboat Harness outperformed the significantly more expensive Claude Opus 4.8 flagship model combination across all five tests.
Key Findings
- Variable Verification: In isolated sandboxes, the same DeepSeek model performed averagely under its official Harness but saw massive improvements when integrated with Floatboat, proving the decisive impact of the underlying system on Agent capabilities.
- Long-Horizon Advantage: The longer and more complex the task chain, the greater the performance gain from the Floatboat Harness (up to 23.6%), highlighting its strength in maintaining cross-step states and self-correction.
New Metric: HLR
The team proposed the Harness Leverage Ratio (HLR) to measure the ROI of upgrading the execution system versus upgrading the base model. Data shows that for long-horizon tasks, optimizing the Harness yields far greater returns than simply upgrading to a pricier model.
More from coding & agent
- Qwen 35B-A3B MoE is 4x Faster Than 27B Dense in Local Coding Tests — WSTangoDelta · 2026-08-08
- Codex SubAgents Inherit Top-Tier Models, Burning Through Tokens Unnoticed — huangyun_122 · 2026-08-08
- AI Recovers 14-Year-Old Abandoned Minecraft World Death Site — gabriel1 · 2026-08-08
- Design MCP Servers with Frontier LLMs, Test with Low-End Models — tristanbob · 2026-08-08
- Developer launches human-review: visual editor for HTML/Markdown with AI feedback, 500+ GitHub stars — petergyang · 2026-08-08
- Warning: ComfyUI v0.31.0 Causes Random Crashes and CUDA OOM Errors — ShutUpYoureWrong_ · 2026-08-08