EP-Flow predicts disordered crystals from the formula alone, at a 60.8% match rate

EP-Flow: Disordered Crystal Structure Prediction without Site-Level Annotations

Qiuliang Liu, Liming Wu, Qi Li, Zhonglong Peng, Chang Chen, Xiaolong Chen, Wenbing Huang, Shifeng Jin

cs.LG, cond-mat.dis-nn

2026-10-01

EP-Flow generates occupancies, coordinates, and lattice parameters from a formula alone. On MPDS Sub-20, structure match rate is 60.8%, 15.7 points above adapted DMFlow.

What problem this solves

Crystal generators such as CDVAE, DiffCSP, FlowMM, and MatterGen put one element on each site. Occupancy is 0 or 1, the pattern inherited from 0 K DFT structures. Nearly 47% of verified ICSD entries contain disorder. Substitution, vacancies, and interstitials change ionic transport, magnetism, and superconductivity.

Dis-Gen and DMFlow do model disorder, but they need site-level labels: which sites are doped, and at what occupancy. The usual input is only a chemical formula. That formula fixes how much of each element is present, not how the mass is split across sites. The split is coupled. Mass used on one site cannot be used again on another. For vacancies and interstitials, the reported composition does not even fix how many crystallographic sites exist.

Method

EP-Flow writes a crystal as an occupancy distribution matrix (ODM), together with fractional coordinates and a lattice. There are N sites and M species, with vacancy counted as a pseudo-species. Each row is one site. Row sums equal 1, column sums equal the formula counts c, and every entry is non-negative. Those three constraints form a transportation polytope whose shape depends on c. An ordered crystal is a vertex of that polytope: one species at 1, the rest at 0. The network sees only c.

Flow matching has no shared Euclidean space when every formula brings a different polytope. EP-Flow double-centers the log-occupancy, subtracting row and column means, and lands every formula in one zero-marginal subspace. The removed terms act like site and species chemical potentials. What stays is the part of the pattern that does not depend on the particular stoichiometry. The flow is a straight line in that subspace. Predicted velocities are projected back by double centering, so the marginals remain zero. A Sinkhorn map then exponentiates the matrix and rescales rows and columns until it is the unique valid occupancy for that formula. For fixed c, the interior of the polytope and this subspace are in one-to-one correspondence.

Fractional coordinates move on a torus. The target velocity is the shortest periodic displacement, so 0 and 1 are neighbors rather than opposite ends of an interval. Lattice parameters follow FlowMM: lengths in log space, angles mapped from logits into the Niggli range [60°, 120°]. A formula-conditioned message-passing net on a fully connected graph is built to be rotation- and translation-invariant. Edges carry periodic displacements and site-to-site occupancy differences. The loss is a weighted MSE on four velocities: occupancy, coordinates, cell lengths, and cell angles. A formula mask blocks site-species pairs that cannot occur.

The adapted baselines are constrained too. FlowMM and OMatG predict an unconstrained occupancy, pass it through a sigmoid, and run the same Sinkhorn projection. DMFlow diffuses on a sphere so each site already sums to 1. At inference every model shares the mask and the final projection. The comparison is where the flow is trained.

Results

The sets contain 17,239 COD structures and 110,966 MPDS structures, split about 8:1:1 into substitutional disorder and vacancy-inclusive disorder, with at most 20 or 50 sites. Metrics are structure match rate, RMSD on matched samples, occupancy MAE (OccMAE) on matched samples, and physical validity. RMSD and OccMAE include only hits, so a model that matches harder cases can look worse on those two averages.

SubsetModelMatch rateOccMAEValidity
MPDS Sub-20Ada-DMFlow45.17%0.009692.10%
MPDS Sub-20EP-Flow60.82%0.007597.80%
MPDS Disorder-20Ada-DMFlow42.59%0.021194.10%
MPDS Disorder-20EP-Flow57.74%0.009294.90%
MPDS Sub-50Ada-DMFlow25.00%0.032081.50%
MPDS Sub-50EP-Flow39.33%0.007691.90%
MPDS Disorder-50Ada-DMFlow23.84%0.025084.10%
MPDS Disorder-50EP-Flow32.60%0.006989.60%
COD Sub-20Ada-DMFlow21.61%0.018990.60%
COD Sub-20EP-Flow45.39%0.011195.10%

Across eight subsets, EP-Flow's match rate runs from 17.10% to 60.82%, which is 7.36 to 23.78 percentage points above Ada-DMFlow. Adapted FlowMM stays under 1%. Adapted OMatG spans 3.28% to 11.78%. The weakest COD slice is Sub-50: EP-Flow 17.10%, DMFlow 9.74%.

Validity is not recovery. On MPDS Disorder-20, FlowMM is 95.90% valid and matches 0.24% of structures. On MPDS, EP-Flow's OccMAE stays between 0.0069 and 0.0092, while DMFlow reaches 0.0320 and 0.0250 on 50-site cells. RMSD is often not the best entry in the row. On MPDS Sub-20 it is 0.0483 for EP-Flow, 0.0416 for DMFlow, and 0.0331 for OMatG.

Ablations use MPDS Sub-20. Removing double centering and Sinkhorn drops the match rate to 0. Removing only the projection raises OccMAE from 0.0075 to 0.1277 and cuts the match rate to 15.78%. Removing only double centering leaves 46.06%. Removing formula conditioning drops 60.82% to 20.01%. Removing the occupancy-difference edge feature drops the match rate to 35.23% and raises OccMAE to 0.0159.

Reconstructions are about local patterns, not just the average composition. In Ba2SmAl0.15Cu2.85O6.98, oxygen vacancies sit outside the CuO2 planes (RMSD 0.0424, OccMAE 0.0016). In YTi0.332Fe11.668N0.5, both the roughly 8.3% Fe/Ti substitution and the half-occupied interstitial nitrogen are recovered (RMSD 0.0208, OccMAE 0.0006). In Sr1.8La0.2FeCoO4.242F1.758, A-site, B-site, and O/F mixing come back together (RMSD 0.0879, OccMAE 0.002, largest site error 0.0065).

Why it matters

Ordered CSP misses the nearly half of experimental crystals whose properties sit on defects. Electrolytes, cuprates, permanent magnets, and mixed-anion perovskites are in that group, and the file in hand is often just a formula. EP-Flow takes the formula and the site count. It does not need a doping pattern specified in advance.

An extra projection does not explain the gain. The baselines run Sinkhorn as well. On MPDS Sub-20, removing double centering and the projection together zeros the match rate. Training on the raw polytope pushes OccMAE to 0.1277. The step from 45.17% to 60.82% comes from learning the flow in one shared zero-marginal space.

The scale is easy to read. On the larger MPDS set, substitutional cells with at most 20 sites reach 60.8%. At 50 sites with vacancies, MPDS falls to 32.6%, and the weakest COD slice is 17.1%. This is structure recovery for a given composition. No energy or property is computed, so it is not a materials-discovery result.

Limitations

The matcher is built for this benchmark, and its tolerances are in the appendix. FlowMM under 1% shows that an ordered generator plus continuous occupancies is not enough. On the best MPDS slice, nearly four in ten formulas still fail the structure match.

DMFlow in this comparison was moved onto the no-annotation protocol and given the same formula mask and the same Sinkhorn step. That is not the site-supervised setting of the original DMFlow paper. EP-Flow leads inside this protocol. That lead is not a direct win over the original model.

RMSD and OccMAE average only over hits. On COD Sub-20, FlowMM reports OccMAE 0.0002 with a 0.65% match rate. Validity tilts the same way. On MPDS Sub-50, FlowMM is 98.10% valid and EP-Flow is 91.90%. On COD Sub-50, OMatG is 79.70% valid and EP-Flow is 77.70%. The abstract bundles validity in with the other gains. The tables do not support that on every split. There is also no separate formula-consistency column. Composition is enforced by the inference projection.

The case studies stop at reconstruction. Vacancies placed off the CuO2 plane are read as a preference for the charge-reservoir chains, with no energy or transport number attached. The site count N is still an input. The hard part of vacancy and interstitial disorder is that composition does not determine how many sites exist. The model is not asked to answer that.

Terms

Source

What people are saying

Related papers

All paper explainers