Co-Scientist evaluation: Severe hallucinations drop to 4%, fabrication to 0%
SRSchmidgall · x · 2026-08-29
Double-blind reviews of 150 AI-generated papers showed that Co-Scientist reduced severe data hallucination from 90% to 4%, extreme data fabrication to 0%, and plagiarism from 60% to 16%, while significantly improving proper attribution.
Related event: Co-Scientist Review: Hallucinations Plummet but Limits Remain(3 posts)→
More from Models
- Is Qwen 3.8 27B at Q2 quantization still usable? A 16GB owner asks — Effective_Head_5020 · 2026-08-29
- User Reports Opus 5.1 Fixes Response Style, Ditches Technobabble — daniel_mac8 · 2026-08-29
- Users Report Opus 5 Struggles with Instruction Following, Ignores Negative Constraints — TheOnlyVibemaster · 2026-08-29
- Is it normal to spend $100 in a few hours on GLM 5.3 API? — BLUECOW009 · 2026-08-29
- MiniMax H3 Raises Shape Mismatch Error with Audio Reference Input — Ok-Flatworm5070 · 2026-08-29
- TRACES Leaderboard Evaluates AI Discovery Capabilities; Kimi and GLM Top Charts — SimonShaoleiDu · 2026-08-29