LLM shows surprisingly usable calibration classifying abstracts on human subjects
RexDouglass · x · 2026-09-29
Researchers had Jev (an LLM) classify whether abstracts analyze humans, finding its probability calibration is not totally made up — surprisingly useful given low cost, speed, and minimal tuning effort. A practical observation on using LLMs for literature screening tasks.
More from Models
- Sonnet 5.5 clones open-source editor Proof at low effort, joining elite group of just four models — every · 2026-09-29
- UsageBench launches to track Claude and Codex usage limits over time — alejandroll10 · 2026-09-29
- User finds bug in top-reasoning GPT's theoretical proof, urges manual verification — MvsCerezo · 2026-09-29
- Opus 5.5 explains its own animation work in a 2-minute first-principles video — jnack · 2026-09-29
- Anthropic's Thariq Shihipar on Claude Code Mods, mutable software, and agent security — Latent Space · 2026-09-29
- Signull: Gemini team's mistake with scrapping Gemini 3.5 Pro was admitting it — signulll · 2026-09-29