New RL Framework Learns Transferable Control Policies from Action-Free Neural Recordings
wgilpin0 · x · 2026-10-09
A new arXiv paper by Emonds and Koppe presents a hierarchical model-based RL framework that learns control policies purely from action-free time series, such as neural activity and behavior recordings.
Key ideas:
- A hierarchical dynamical reconstruction model captures shared dynamics and individual variation via low-dimensional embeddings;
- Embeddings parameterize shared policy/value networks, linking dynamical differences to control differences;
- Policies train entirely in simulation with an explicit intervention model; piecewise-linear RNNs make interventions interpretable and modality-selective.
On Lorenz-63 and double-pendulum systems, hierarchical policies transfer better than independently trained ones, approach methods trained with controlled interactions, and generalize to unseen systems. The authors demonstrate training models on neural activity to learn movement-suppressing interventions without recorded interventions.
More from Research
- Blind humanoid walks, plays soccer and lifts suitcases with joint encoders only — accepted at Humanoids 2026 — Jan_R_Peters · 2026-10-09
- Delete object info from observations and PPO learns to search anyway — TU Darmstadt on its Humanoids 2026 paper — Jan_R_Peters · 2026-10-09
- U-Space finds an interpretable subspace for LLM uncertainty, no training needed — Tobias Braun · 2026-10-09
- CARE certifies VLA inference speedups up to 10.8x with statistical guarantees — UMCP · 2026-10-09
- SOL: a sample-based distributional metric proposed for evaluating text diffusion LMs — NandoDF · 2026-10-09
- MaRN: PyTorch library cuts MNIST CNN params 57.7x via low-dim parameter mappings — Less_Dream_6331 · 2026-10-09