Andrew Ng: We need harder evals — have frontier models chat with me
andrewgwils · x · 2026-09-04
Andrew Ng posted that current benchmarks may no longer be sufficient, jokingly challenging people to have their frontier models chat with him so he can ask questions that would "make their heads explode." The underlying point: frontier models still struggle with sufficiently hard open-ended questions, and the community needs harder evals.
More from Models
- GPT-6 Astra skips CoT yet still answers correctly, sparking debate on CoT's privileged status — Dan_Jeffries1 · 2026-09-04
- Astra solves Excel World Championship cases ~4x faster than human champions using pure computer use — sandersted · 2026-09-04
- Researcher argues harness and MCP will be absorbed into models — data is the only wall — A_K_Nain · 2026-09-04
- Polymarket opens betting on next Grok model (4.7+) release, odds point to mid-September — Polymarket · 2026-09-04
- Ex-Google DeepMind researcher denny_zhou reveals move to Meta, worked on Muse Spark 1.1-1.3 — infoxiao · 2026-09-04
- The Real AGI Benchmark: Models Still Can't Do Data Science — max_paperclips · 2026-09-04