Task is detecting null-finding claims, tested against multi-human-coded labels
RexDouglass · x · 2026-09-19
RexDouglass clarifies the task: whether an abstract claims a null finding. Their evaluation uses battle-tested instructions and multi-human-coded labels, giving a reliable ground truth.
More from Models
- Noam Brown: GPT-6 Astra does have observable chain of thought, calls it fragile — burny_tech · 2026-09-19
- Jev's popularity signals the AI crowd is open to models beyond LLMs — BLUECOW009 · 2026-09-19
- QuixiAI picks gemma-4-26B-A4B-it as base model for OpenJev — QuixiAI · 2026-09-19
- Open-weight Jev replica based on Qwen3.8 27B with 265k context drops tomorrow — TheZachMueller · 2026-09-19
- Jev beats a Sonnet 5-powered retriever on accuracy at a fraction of the cost — IanArawjo · 2026-09-19
- Alibaba open-sources medical AI model that detects cancer and nearly 150 conditions — giveen · 2026-09-19