"Terminal-Bench is cooked": coders mock another saturated AI coding benchmark
rickasaurus · x · 2026-10-08
@rickasaurus declares Terminal-Bench "cooked"—saturated by frontier models—adding a self-aware jab that "clearly these are meaningless benchmarks." The thread captures the recurring debate over coding benchmarks losing discriminative power as models race past them.
Related event: Benchmarks Deemed 'Cooked' as AI Models Rapidly Top Leaderboards(2 posts)→
More from Models
- François Fleuret: AI math proofs will go from crude to splendid — francoisfleuret · 2026-10-09
- TensorFold joins NVIDIA Inception, gets early access to next Nemotron for 0-day support — HankYeomans · 2026-10-09
- Open-weight medical decision model MedDecider: 9B beats 3x larger Perplexity model — BraydonDymm · 2026-10-09
- The biggest edge AI misconception: "run anywhere" is not "run everywhere" — ritakozlov · 2026-10-09
- Open-Source Models Power Home Robot Tidying for Toddlers in Weeks, Not Years — chris_j_paxton · 2026-10-09
- Google open-sources EmbeddingGemma 2: 740M multimodal embeddings that run on a Pixel — Prompt Engineering · 2026-10-09