RLFR Cuts Hallucinations by 37% with Internal Probes

burny_tech · x · 2026-07-15

The Goodfire team developed RLFR, a method that uses a model's internal probes as reinforcement learning reward signals during training.

They noted that Silico replicated this method in just 2 days, reducing the hallucination rate in Qwen3-8B by 37% without any degradation in the model's capabilities.

Related event: RLFR Method Uses Internal Probes to Reduce LLM Hallucinations by 37%(3 posts)→

Original post →

More from Models

Models channel →