BOLD Lab presents RLConf works: PPO scaling, Offline RL, LLM diagnostics
j_foerst · x · 2026-08-18
BOLD Lab presented several research works at the RL Conference, focusing on the intersection of reinforcement learning and LLMs:
- Preventing Learning Stagnation in PPO by scaling to 1M parallel environments.
- Fully Offline Reinforcement Learning led by Mattie Fellows and Clarisse Wibault.
- Hierarchical Behaviour Spaces led by mitrma.
- When Do We Need LLMs? A diagnostic for Language-Driven Bandits.
More from Research
- Prof. teaches AI architectures via hand calculation: Transformer to Mamba — techNmak · 2026-08-18
- Video Model Evaluation Guide: Quantifying Generation Quality — Majumdar_Ani · 2026-08-18
- Study: LLMs as synthetic survey respondents are plausible but not valid — Mantas Lukauskas · 2026-08-18
- Adaption AI Launches Custom Evals for Pro Users — sarahookr · 2026-08-18
- Interactive Diagram: Understanding Autoencoders by Hand — ProfTomYeh · 2026-08-18
- PyLate Merges Back Into Sentence Transformers — mrdrozdov · 2026-08-18