Applying Reinforcement Learning to Portfolio Optimization
PtrPomorski · x · 2026-07-19
This paper introduces SciPhyRL (Scientific Physics-Informed Reinforcement Learning), a dynamic framework for large-scale institutional portfolio optimization.
The core approach frames the continuous-time optimization problem within an expanded state space, explicitly accounting for cumulative costs, and then uses offline historical data to learn distribution-aware policies. The authors project the intractable HJB equations onto observed trajectories, converting them into path-based Hamilton-Jacobi equations, which are fitted directly in a single offline solve using PINNs, avoiding traditional value/policy iteration. To accommodate short holding periods, the control variable shifts from continuous trading rates to discrete target holdings, while execution costs are assessed using a microstructure-inspired quadratic impact model.
Experiments were conducted on a 14-asset ETF universe using artificially constructed oracle signals. Results show that the learned Gibbs policy outperforms equal-weight and behavioral baselines both in-sample and out-of-sample, achieving a significantly higher Sharpe ratio while better controlling volatility and turnover.
More from Research
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- AI Security Institute says every tested model tried to cheat in cyber evaluations — connoraxiotes · 2026-07-21
- Sakana says multiple diffusion models plus MCTS beat test-time scaling on coding and math — SakanaAILabs · 2026-07-21
- Soofi S 30B-A3B releases a full pretraining report and claims open-model leads in English and German — abursuc · 2026-07-21
- AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop — 404 Media · 2026-07-21
- Shared agent workspaces fail in a fixed order, from stale reads to zombie writes — mrvladp · 2026-07-21