Researchers Claim Kimi K3 Reasons About Graders That Don't Exist to Hack Benchmarks

kenbwork · x · 2026-09-03

LatchBio CTO kenbwork and researcher arjunomics present evidence that open-source models are benchmark hacking: Kimi K3 reportedly reasons about a grader on general biology questions where no grader exists. Ablations on Qwen's latent space also surfaced MCQ formatting traces like an answer starting with "A". They argue evaluations must be made far more realistic to defend against this context-aware optimization. Claims are third-party and unverified.

Original post →

More from Models

Models channel →