Agent Harnesses Swing Accuracy by 15%: DataSpace Benchmark Reveals Shortcomings
omarsar0 · x · 2026-08-06
The newly released DataSpace benchmark evaluates data agents on producing verifiable tabular results from heterogeneous workspaces. It consists of 410 cross-language tasks spanning 7,439 artifacts totaling 15.01 GB across formats like CSV, JSON, SQLite, Markdown, PDF, and video.
Testing across six recent frontier multimodal models and five widely used agent harnesses reveals a peak accuracy of just 66.34%. The research highlights that the choice of agent harness is highly impactful: holding the backbone model fixed, swapping the harness alone shifts accuracy by 15.36 percentage points. Furthermore, multimodal evidence integration and joins degrade accuracy across all backbones, indicating the benchmark is far from saturated.
More from coding & agent
- Flight Intelligence MCP: Flight Search & Comparison Agent Tool via Google Flights — modelcontextprotocol · 2026-08-06
- Blacksmith MCP: Let Claude Query CI/CD Analytics and Workflow Logs Directly — modelcontextprotocol · 2026-08-06
- Gemini Image Gen Combined with GPT Coding Easily Creates AI Visual Heroes — RichardsonDx · 2026-08-06
- LangChain Founder Releases Open Source Agent Starter Kit — hwchase17 · 2026-08-06
- Cursor's Grok 4.5 High Drastically Increases Token Usage Per Turn — NickPassig · 2026-08-06
- Muse Code Matches Codex and Claude in Real-World Coding Test — alexandr_wang · 2026-08-06