A frontier agent fact-checks a local Qwen 27B agent: 17/21 trick questions passed

AIForOver50Plus · reddit · 2026-10-05

The author runs a local Qwen 3.6 27B agent over 59,000 chunks of his own documents, with a frontier agent dedicated solely to fact-checking every assertion. Weekend grading results: 3/3 real questions passed with zero invented claims; 17 of 21 trick questions (false premises or no honest answer) passed, 4 partial, 0 failed.

The real lesson: one answer claimed a JSON file "could not be parsed" when it was clean — the local agent guessed instead of looking. Now every claim, including every "I couldn't find," gets fact-checked before he sees it. The pipeline runs weekly: Friday re-run, Saturday fact-check, he grades only what changed, and every correction is saved for future fine-tuning.

Original post →

More from coding & agent

coding & agent channel →