Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction
Georg Trede, Charlotte Ricarda Doll, Elias Weber, Daniel Durstewitz
cs.LG, math.DS, nlin.CD
2026-06-22
Three structural flaws explain why hierarchical DSR models fail OOD; feature splitting plus sparsity gives zero-shot prediction of chaos and limit cycles across bifurcations.
Dynamical systems reconstruction (DSR) learns a generative surrogate from time series that reproduces a system's attractor geometry and long-term statistics. The current recipe is hierarchical training: a family of systems shares one set of "physical law" parameters, while each system gets a low-dimensional latent feature that generates its specific parameters through an affine map. A recurring empirical finding is that these features end up nearly linearly related to the true control parameters, which should in principle let you extend the feature and predict parameters never seen in training.
It doesn't. Past a tipping point, where the topology of the state space changes, these models break down and need fine-tuning or retraining on data from the new regime. This paper pins the failure on model structure: three structural assumptions conflict with typical properties of physical systems, and for each one a theorem shows extrapolation failure is guaranteed, no matter how well you train.
To predict at unseen parameters, each feature is fitted with one of three candidate functions (power law, signed power law, parabola), selected by 5-fold cross-validation on the training domain, then extended outward. Experiments cover discrete-time RNNs (shPLRNN) and Neural ODEs, with consistent conclusions.
Evaluation uses the Wasserstein-1 distance between true and reconstructed bifurcation diagrams, medians over 10 training runs.
| Experiment | Training domain | OOD result |
| Lorenz-63, Neural ODE | ρ∈[5, 22.5], two-fixed-point regime | Unregularized model collapses at the transition to chaos (ρ≈24.74); with low-rank and sparsity constraints it tracks the true bifurcation diagram across the whole ρ∈[22.5, 35] test range, with equal in-domain performance |
| Lorenz-63, shPLRNN | same | The regularized model reconstructs a 10,000-step chaotic trajectory at ρ=29; the unconstrained one collapses to a fixed point |
| Selkov, both architectures | b∈[0.8, 1.1], fixed-point regime | Feature-splitting models predict both the onset of the limit cycle and the second Hopf bifurcation back to fixed points; single-feature models fail, and so does raising the feature to 2 dimensions |
The two fixes target different defects. Selkov's b is a pure inflow that never enters the Jacobian; splitting does the work and regularization adds nothing consistent. Lorenz-63's ρ is a Jacobian parameter; splitting alone cannot stop the full-rank model from diverging, and the combination gives the lowest OOD error. Compute is trivial: about 10 minutes and 2GB per Selkov model, 30 minutes and 4GB for the Lorenz-63 Neural ODE.
Tipping points are exactly what scientific ML most needs to predict, in climate, epilepsy, and ecosystem collapse, and exactly where current models stop working. This paper turns a vague empirical failure into three provable structural theorems, not tied to one architecture: RNNs and Neural ODEs are both covered. The fixes are small changes to the parameterization and training objective; no extra data, no knowledge of true control parameters during training, minutes of compute. The lesson travels beyond DSR: any conditional model that generates weights from a latent feature through an affine map (hypernetworks, context-conditioned dynamics models) carries dense coupling that is invisible in-domain and collects interest the moment you extrapolate.
Honest framing: this is a capability unlocked on low-dimensional textbook systems, still distant from real high-dimensional observations.
The authors state three: control parameters are assumed to enter the equations affinely (or in a single functional form), which need not hold; the sparsity assumption likewise; and the discretization mismatch is only bounded, not solved. Two more from reading the paper. All results are on Lorenz-63 (3D) and Selkov (2D); there are no high-dimensional or real-world systems. And although the conclusion claims OODG without explicit knowledge of the true control parameters, the extrapolation pipeline (Appendix F) fits each feature as a function of p and evaluates it at the new parameter value, so the numeric value of the control parameter is needed at prediction time; the accurate reading is that training doesn't require it. The bifurcation diagrams displayed in Fig. 3B are the best runs per setting, while median error bands remain visible far outside the training range.