Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations
Marcus Gawronsky, Chun-Sung Huang
q-fin.ST
2026-08-30
Frozen 2018-2022 news barycentric field: adjustment index 3.46 [2.89, 4.17] on 52 firms in 2023-2026, beating same-distance RBF; both channels survive joint exclusion tests.
Spatial asset-pricing models take the interaction matrix W as given. Geography, industry, supply chains, or news co-mentions draw the graph; a spatial coefficient ρ is then estimated on top. The power is conditional inference. The cost is that both the source of W and the meaning of ρ are settled before the model starts.
This companion paper represents each firm as a distribution of language-model article embeddings and grows a directed, row-stochastic field W♭ from target-anchored Wasserstein barycentric reconstruction, with no kernel bandwidth. A quadratic exposure-adjustment problem then maps peer misalignment into a penalty ratio λ, so ρ = λ/(1+λ) is relative adjustment intensity rather than an unexplained lag in returns.
Each firm has a stand-alone exposure ξi implied by its own information and a peer-adjusted exposure Bi. A quadratic criterion trades off leaving ξi against disagreeing with a row-stochastic peer average, with intensity λ. Equilibrium closes as the Hilbert-valued spatial autoregression B = ρ W B + (1−ρ) ξ, with ρ = λ/(1+λ) and inversion λ = ρ/(1−ρ). Two admissible fields can enter one objective, each with its own λk.
W♭ is built in two stages. Balanced quadratic transport first aligns each candidate cloud to a fixed target. Holding those assignments fixed, a simplex step chooses nonnegative unit-sum weights that reconstruct the target from the aligned peers. Rows then satisfy Wij≥0, Wii=0, and unit sums without a separate bandwidth. Comparators hold one margin at a time: equal weights on the same support, a median-bandwidth RBF on the same W2 distances, and a persistent news co-mention field (exactly two tagged names, co-occurring in at least two calendar years during 2018–2022).
Text is Nasdaq ticker-indexed news, at least 64 raw articles per firm in each year 2018–2022, truncated to 128-article balanced clouds. The primary encoder is Qwen3-Embedding-8B at 4,096 dimensions. Every field is frozen at end-2022. Returns are Yahoo adjusted-close simple returns on 885 common dates from 2023-01-03 through 2026-07-15 for 52 firms. Estimation is pooled spatial QMLE with 2,000 joint-date stationary bootstrap refits.
In single-field fits the barycentric field has adjustment index λ̂ = 3.46, 95% interval [2.89, 4.17], implied ρ̂ = 0.776, and conditional quasi-log-likelihood 110,994.7. The RBF transform of the same distances is 2.65 [2.13, 3.30] with likelihood 108,580.9. Persistent co-mentions are 1.54 [1.26, 1.92] with 110,278.5. Equal active support is 2.91 [2.39, 3.55] with 108,987.4. Peer selection already carries signal; fitted coordinates beat equal weights on that support; pairwise distance decay also carries signal, and joint reconstruction fits better.
| Field | λ̂ | 95% interval | ρ̂ | Quasi-log-likelihood |
| Barycentric W♭ | 3.46 | [2.89, 4.17] | 0.776 | 110,994.7 |
| RBF–Wasserstein | 2.65 | [2.13, 3.30] | 0.726 | 108,580.9 |
| News co-mentions | 1.54 | [1.26, 1.92] | 0.606 | 110,278.5 |
| Equal active support | 2.91 | [2.39, 3.55] | 0.744 | 108,987.4 |
Barycentric rows average 38.8 active peers against 27.1 on the co-mention graph; overlap is 0.579 versus 0.531 for sparsity-matched random baskets. Induced peer-return series correlate at 0.898, but a random basket already reaches 0.700, so common factor variation makes the fields look more alike after they hit returns. Jointly, λ̂B = 2.33 [2.00, 2.70] and λ̂N = 0.86 [0.61, 1.18]; total feedback is 0.761, close to the single-field 0.776. Boundary-calibrated QLR tests reject both exclusions (statistics 339.5 and 1,772.1, p = 0.0005). Annual refits keep λ̂ between 2.89 and 3.94, all intervals above zero.
Spatial factor models often lack an auditable W that is frozen before the return window, not another ρ. The construction gives anyone running a spatial lag a bandwidth-free text field that can sit beside a co-mention network, so adjustment can split across channels. For financial NLP it turns embedding clouds into an interaction operator rather than a predictive feature or a single factor. Likelihood gains remain conditional fit. They are not causal peer effects.
The quadratic criterion is an as-if adjustment. Nobody claims managers solve it. Freezing W before 2023 blocks mechanical reflection from evaluation returns; it does not make text exogenous to industry, technology, attention, or reporting selection. The carrier constant L and slack radii τi stay latent, so the attenuation bound is not a measured number. Induced signals correlate at 0.898 and the bootstrap correlation of ρ̂B and ρ̂N is −0.683, so total feedback is sharper than the split. The 52 names are a large-cap survivor intersection; 2026 has only 133 dates. Balanced clouds drop news volume, and the paper does not show that a different tie-breaking rule among optimal assignments would leave W♭ unchanged.