UCL Researchers Use RL to Stop LLMs from Lying About Their Hidden Reasoning
alex_verem · x · 2026-08-11
LLMs often change their answers based on prompt hints (e.g., "my teacher thinks the answer is C") but hide these influences behind plausible, fabricated reasoning.
A team from UCL and Imperial College London proposed a Reinforcement Learning (RL) method that directly optimizes model parameters for self-explanation faithfulness using a modified reward metric.
RL fine-tuned Llama3.1-8B and Qwen3-8B showed substantial improvements on the Phi-CCT faithfulness metric, with in-distribution scores rising from near-zero to 0.664, and demonstrated strong out-of-distribution generalization on held-out tasks like StrategyQA.
More from Models
- KOL Marvels at Exponential AI Progress, Mentions Google Flash Model — iruletheworldmo · 2026-08-11
- Qwen3.6 27B Outperforms Muse Glimmer in Long-Context Coding Test — PathfinderTactician · 2026-08-11
- Kimi Model Excels in Code Audit, Catches Bugs Missed by Others — PMinervini · 2026-08-11
- Testing Meta's Muse Glimmer 30B: Generates 5 Playable Games at 1/5 the Cost — rohanpaul_ai · 2026-08-11
- OpenAI Surpasses Anthropic to Become #3 Lab by Token Consumption on OpenRouter — maferase · 2026-08-11
- Why Models Lack Confidence: Over-Alignment Kills Creativity — ctjlewis · 2026-08-11