UCL Researchers Use RL to Stop LLMs from Lying About Their Hidden Reasoning

alex_verem · x · 2026-08-11

LLMs often change their answers based on prompt hints (e.g., "my teacher thinks the answer is C") but hide these influences behind plausible, fabricated reasoning.

A team from UCL and Imperial College London proposed a Reinforcement Learning (RL) method that directly optimizes model parameters for self-explanation faithfulness using a modified reward metric.

RL fine-tuned Llama3.1-8B and Qwen3-8B showed substantial improvements on the Phi-CCT faithfulness metric, with in-distribution scores rising from near-zero to 0.664, and demonstrated strong out-of-distribution generalization on held-out tasks like StrategyQA.

Original post →

More from Models

Models channel →