Using Internal Probes for RL Rewards

burny_tech · x · 2026-07-15

GoodfireAI introduced RLFR: using a model's internal probe as a reinforcement learning reward signal to train the model to hallucinate less.

Reposts mention that Silico replicated this method 2 天内, claiming to have reduced the hallucination rate by 37% on Qwen3-8B without any loss of capability.

Related event: RLFR Method Uses Internal Probes to Reduce LLM Hallucinations by 37%(3 posts)→

Original post →

More from Research

Research channel →