TypesafeAI precommits to de-emphasizing public evals: "they reward bad actors"

hardimanjames · x · 2026-09-17

A person behind TypesafeAI says public evals are a poor signal of model intelligence because they reward bad actors, and the team has precommitted to de-emphasizing them even when results look good — opting to build trust "the long and hard way" through real-world reliability. The framing post is promotional, but the quoted statement reveals a small lab's stance on eval culture.

Original post →

More from Models

Models channel →