Qwen, Llama oss and Gemma all fall short on academic abstract null-finding task
jon_mellon · x · 2026-09-19
Responding to RexDouglass's question, jonmellon confirms that Qwen did not perform well on their academic text classification task, and neither did Llama oss or Gemma. In context, the task is detecting null findings in academic abstracts, where most cheap models they tested were middling and only a new model looks promising.
Related event: Scholars Find Cheap LLMs Fail at Detecting Null Findings in Abstracts(3 posts)→
More from Models
- Dev runs 16-model eval: Jev and Haiku tie at 0.121/0.122 on calibration error — AlexKim · 2026-09-19
- 16-model eval finds Sonnet 5 hedges most, landing in ambiguous zone 41.3% of the time — AlexKim · 2026-09-19
- Dev tests 16 models to evaluate TypeSafe's Jev — it ranked 10th on accuracy — AlexKim · 2026-09-19
- Mystery Model Jev Launches Claiming 200x Speed and 400x Cost Cuts, Devs Impressed — multiply_matrix · 2026-09-19
- GLM 5.3 Flash leads quality, Qwen 3.8 Flash Next wins speed in open small-model comparison — HankYeomans · 2026-09-19
- Ternary Bonsai 2 27B quantized to a 7GB single file, runs on an 8GB GPU — cephaloform · 2026-09-19