EMNLP oral paper: RL helps models traverse parametric knowledge inaccessible after instruction tuning
niloofar_mire · x · 2026-09-30
An EMNLP oral paper by @alexxxzzz6825, @niloofarmire, @RylanSchaeffer and Manasa Kaniselvan finds that reinforcement learning teaches models to traverse parametric knowledge more effectively, unlocking knowledge that was previously inaccessible in instruction-tuned models — suggesting RL does more than align behavior, it can surface latent capabilities already in the base model. A thread with details is linked.
More from Research
- UMass professor Luc Rey-Bellet's stochastic processes lecture notes on Markov chains and MCMC — michaelchchoi · 2026-09-30
- VoxMem benchmark: none of 15 audio LLMs top 40% on spoken multi-session memory — unimelb-hf · 2026-09-30
- EpiCon: shared multimodal memory lifts agent scores 1.7-4.9 points across 11 benchmarks — Ziyun Zeng · 2026-09-30
- EngiWorld: top model scores just 44.3 on professional engineering agent benchmark, 3.6% multi-software success — zhiman-ai · 2026-09-30
- Reserved-token embeddings carry prompt injection authority; standard defenses fail on 255 of top 400 chat models — PekingUniversity · 2026-09-30
- Quantized softmax attention pretraining: only +0.004 nats loss gap at K=16 with the right calibration — illinois · 2026-09-30