HumanEvals: Open-Source Library Adds Real Human Judgment to Multimodal Evals
_akhaliq · x · 2026-08-19
Datapoint has open-sourced HumanEvals, a library for adding real human judgment to multimodal model evals: send image, video, or audio outputs and get pairwise preferences, ratings, or rankings from real people in seconds.
- Uses the same interface as automated eval libraries (autoevals-compatible Score objects).
- Human responses come via the Datapoint annotation API — the same pool leading image, audio, and video model labs use to evaluate checkpoints — at 5,000+ annotations per minute.
- Supports pairwise comparison, fixed-scale rating, multiple choice, and ranking; commonly used for checkpoint evaluation during training and competitive benchmarking.
More from coding & agent
- Open-Source RL Framework Miles v0.1 Ships: 1,326 Commits, 85 GPU E2E CI Tests, Firecracker Sandbox Rollouts — ying11231 · 2026-08-19
- Eigent Uses Miles to Train RL Agents for Terminal Coding and ML Engineering — ying11231 · 2026-08-19
- App Store Connect CLI 4.4.4 adds `asc optimize search plan`, turning Apple Ads data into a keyword plan — rudrank · 2026-08-19
- Open-source mocap pipeline turns any video into Mixamo animation, fully operated by an AI agent — andrew_n_carr · 2026-08-19
- DataSmith agent beats Claude Code in data research tasks — soumitrashukla9 · 2026-08-19
- AWS AI League returns: fine-tune a foundation model in 72 hours, deploy it as an agent, win a trip to re:Invent — ruchi798 · 2026-08-19