Humans-plus-AI need their own evals, not just model benchmarks

paraschopra · x · 2026-07-29

We should build evals for the human+AI combination, not just for AI alone.

The post argues that measuring only model performance misses the real target: what people can accomplish when AI is used as an assistant. If the goal is to improve the future of humanity, benchmarks should optimize the combined system, not the model in isolation.

Original post →

More from AGI Musings

AGI Musings channel →