Testing Jev's calibration: AI's grasp of probability wording matches human intuition chart
burny_tech · x · 2026-09-23
@hammermt tested how well Jev by @typesafeai aligns with the famous chart of what probabilities people actually mean by common words — and found it pretty well calibrated, e.g. treating "impossible" as roughly 10%.
More from Models
- Claude Opus 5.5 lands: beats Fable 5.1 on coding benchmarks, cuts API prices 60% — xiaohu · 2026-09-23
- Devs migrate from GPT-6 Astra to Opus 5.5 as coding model race churns — rudrank · 2026-09-23
- Rumor: Muse Spark 1.3 Max beats GPT-6-Sol on quality, price and speed—and Alexandr Wang hints it's real — alexandr_wang · 2026-09-23
- Opus 5.5 generates realistic Minecraft + StarWars scenes in one shot — ChrisGPT · 2026-09-23
- Third party 'cracks' 5.95GB ternary-compressed Bonsai 2 at the weight level, refusal rate 93.4% to 0% — solyarisoftware · 2026-09-23
- Astra becomes first AI to beat Zork, finishing in 12 minutes 54 seconds — almostsweet · 2026-09-23