Cognit scores LLMs on famous trick questions like the bat-and-ball problem, with a human-mode twist
UnemployedTechie2021 · reddit · 2026-10-04
开发者推出了 Cognit,一个让 AI 模型做人类经典「陷阱题」的排行榜:如球棒和球共 1.10 美元、球棒比球贵 1 美元(直觉答案是 10 美分,正确是 5 美分)这类 System 1 直觉陷阱。
每个模型做两遍:第一遍正常作答,第二遍被要求「像人一样回答」。100% 表示给出了仔细的正确答案,0% 表示掉进了人类常犯的坑。「frame drop」指标衡量第二次变差的程度——为正说明「扮演人类」确实让模型连错误都更像人了。站点已上线排行榜,作者开放渠道接受添加模型与反馈。
More from Models
- RLM explained: MIT's recursive REPL approach lets LLMs handle contexts beyond the window — AnuranBuilds · 2026-10-04
- LocalMaxxing: community-run speed tests for local LLM rigs, 8.6K runs across 592 hardware setups — lxfater · 2026-10-04
- KOL pans OpenAI's 'dots' keynote release, says it's an easy win for Anthropic — kimmonismus · 2026-10-04
- Claude is too pessimistic: user says it shoots down ideas instead of exploring them — Jumpy-Cobbler1020 · 2026-10-04
- Claude is now asking regular users for interviews — even non-coders — imfamilyfriendlysd · 2026-10-04
- GPT image generation wows blogger in page-34 test — aiamblichus · 2026-10-04