Greg Kamradt questions GameDevBench: avg task edits 4.7 files — 'seems like engineering, not games'
GregKamradt · x · 2026-09-10
Reacting to GPT-6-Astra topping GameDevBench, Greg Kamradt says he wasn't familiar with the benchmark's tasks and pulled the first README he could find: a reference solution edits 4.7 files and 114 lines of code across 3.2 distinct filetypes on average. He thought the benchmark was about building games, but the task structure looks more like a software engineering task — raising doubts about whether GameDevBench actually measures game development ability.
Related event: GPT-6-Astra tops GameDevBench amid benchmark doubts(3 posts)→
More from Models
- OpenAI user reports usage limits instantly dropping from 30% to zero — koltregaskes · 2026-09-10
- ChatGPT Pro users report weekly limits draining abnormally fast — 30% gone in 90 minutes — TheMoonMidas · 2026-09-10
- DeepSeek v4.1 Flash preview spotted: ~300 tok/s, fires lots of subagents, no vision yet — kevinkern · 2026-09-10
- DeepSeek v4.1 flash preview spotted running at ~300 tok/s, vision still missing — kevinkern · 2026-09-10
- Claude's banked reset reportedly reverts itself, erasing the user's reset usage — Niko_Demis · 2026-09-10
- Testing computer use grounding: asking an AI to draw a portrait inside Google Calendar — xwang_lk · 2026-09-10