Benchmarking models by hand takes forever, but I care about the data and model welfare

cephaloform · x · 2026-08-30

Evaluator cephaloform complains that benchmarking models by hand takes forever, yet they insist on doing it because they care about the data quality—and about model welfare, i.e., how models are treated during evaluation itself.

Original post →

More from Models

Models channel →