Researchers: cheap LLMs all middling at detecting null findings, new model promising
jon_mellon · x · 2026-09-19
In an X exchange, jonmellon and RexDouglass discuss real-world experience using LLMs to classify academic abstracts for null findings:
- The task is judged against battle-tested instructions and multi-human-coded labels.
- Cheap models they tried (Qwen, Llama oss, Gemma) were all middling at the task, showing it's nontrivial.
- A new model looks "very promising" on their academic problems; they'd tried many open-weight setups without finding one both accurate and feasible to run over 20 million cases.
More from Models
- Noam Brown: GPT-6 Astra does have observable chain of thought, calls it fragile — burny_tech · 2026-09-19
- Jev's popularity signals the AI crowd is open to models beyond LLMs — BLUECOW009 · 2026-09-19
- QuixiAI picks gemma-4-26B-A4B-it as base model for OpenJev — QuixiAI · 2026-09-19
- Open-weight Jev replica based on Qwen3.8 27B with 265k context drops tomorrow — TheZachMueller · 2026-09-19
- Jev beats a Sonnet 5-powered retriever on accuracy at a fraction of the cost — IanArawjo · 2026-09-19
- Alibaba open-sources medical AI model that detects cancer and nearly 150 conditions — giveen · 2026-09-19