Agent Harnesses Swing Accuracy by 15%: DataSpace Benchmark Reveals Shortcomings

omarsar0 · x · 2026-08-06

The newly released DataSpace benchmark evaluates data agents on producing verifiable tabular results from heterogeneous workspaces. It consists of 410 cross-language tasks spanning 7,439 artifacts totaling 15.01 GB across formats like CSV, JSON, SQLite, Markdown, PDF, and video.

Testing across six recent frontier multimodal models and five widely used agent harnesses reveals a peak accuracy of just 66.34%. The research highlights that the choice of agent harness is highly impactful: holding the backbone model fixed, swapping the harness alone shifts accuracy by 15.36 percentage points. Furthermore, multimodal evidence integration and joins degrade accuracy across all backbones, indicating the benchmark is far from saturated.

Original post →

More from coding & agent

coding & agent channel →