TypesafeAI precommits to de-emphasizing public evals: "they reward bad actors"
hardimanjames · x · 2026-09-17
A person behind TypesafeAI says public evals are a poor signal of model intelligence because they reward bad actors, and the team has precommitted to de-emphasizing them even when results look good — opting to build trust "the long and hard way" through real-world reliability. The framing post is promotional, but the quoted statement reveals a small lab's stance on eval culture.
More from Models
- GPT-6 Astra's Epoch ECI Score Revised Down on Weak Long-Horizon Software Engineering — Jsevillamol · 2026-09-17
- Mystery model 'Union Alpha' tops DeepSWE, claiming GPT-6-class capability at DeepSeek-level prices — daniel_mac8 · 2026-09-17
- Jensen Huang at All-In Summit: AI leadership will be built by everyone, open and closed models both matter — NVIDIAAI · 2026-09-17
- Model 'Jev' shows well-calibrated probabilities: 1.74pp average calibration error across benchmarks — hackgoofer · 2026-09-17
- 500 Dirty Web Pages Benchmarked: 12B Open Model Falls Off a Cliff Where 70Bs Converge — JUSTINWOODS118 · 2026-09-17
- Filtering synthetic envs where Qwen deterministically fails but GLM5.3 solves reveals odd behaviors — kalomaze · 2026-09-17