Testing 32 Local Models: Why F1 Scores Hide Dangerous Hallucination Traps

KitchenAmoeba4438 · reddit · 2026-08-06

The author benchmarked 32 local LLMs on a fact-extraction task using 1,001 notes to evaluate their ability to stay silent when no facts are present.

Key Findings

The author advises evaluators to track abstention rates and invented-triple counts alongside F1, and to ensure test corpora include cases where the correct answer is silence.

Original post →

More from Models

Models channel →