Qwen, Llama oss and Gemma all fall short on academic abstract null-finding task

jon_mellon · x · 2026-09-19

Responding to RexDouglass's question, jonmellon confirms that Qwen did not perform well on their academic text classification task, and neither did Llama oss or Gemma. In context, the task is detecting null findings in academic abstracts, where most cheap models they tested were middling and only a new model looks promising.

Related event: Scholars Find Cheap LLMs Fail at Detecting Null Findings in Abstracts(3 posts)→

Original post →

More from Models

Models channel →