Tsallis Fq turns FST into a q-spectrum that reweights rare versus common alleles

A Tsallis-Entropy Lens on Genetic Variation

Margarita Geleta, Daniel Mas Montserrat, Alexander G. Ioannidis

cs.IT, cs.CE

2025-11-05

Fq nests FST at q=2 and Shannon differentiation at q=1. On 865 Oceanian genomes and 17-generation simulations, OVR tracks differentiation and LOO names the driving deme.

What problem this solves

The workhorse statistic for genetic differentiation is still Wright's fixation index FST: (HT − HS) / HT, the share of heterozygosity lost to subdivision. It is a second-order functional of allele frequency, so common variants dominate. Whole-genome spectra are skewed. Recent drift and founder effects sitting on rare alleles often barely move FST. Jost's D tries to detach differentiation from within-group diversity, but intermediate values lack a clean coalescent or migration reading.

The same group also builds genotype simulators (autoencoders, VAEs, generative moment matching networks). A simulator that only matches variance-based FST can still be wrong on the rare end of the spectrum. They want a ruler whose weights can be tuned along that spectrum.

Method

Tsallis entropy of order q generalizes Shannon entropy. For a biallelic site, Sq = [1 − p^q − (1−p)^q] / (q−1) when q ≠ 1, and it recovers Shannon as q → 1. Let Sq^total be the entropy of the pooled allele frequency and Sq^within the weighted mean of subpopulation entropies. Absolute differentiation Δq is the Jensen-Tsallis gap between them. The relative statistic is Fq = Δq / Sq^total, bounded in [0, 1].

Two anchors pin the family down. At q=2, S2 equals expected heterozygosity 2p(1−p), so F2 is classical FST. At q=1, Δ1 equals the mutual information I(X; Y) between the allele and the population label, so F1 is Shannon differentiation I(X; Y)/H(X).

Low q up-weights rare alleles (recent drift, founder events). High q up-weights common alleles (older structure). Two slicing modes make this usable:

Results

The empirical panel is 865 Oceanian genomes, 1,823,000 biallelic SNPs, split into 1,730 haplotypes and grouped as Polynesia, Micronesia, Melanesia, and Southeast Asia. Sample sizes are uneven, so they bootstrap 100 times inside each population with a cap of 40 haplotypes and report equal-weight estimates.

The text does not give point estimates of Fq. Claims are tied to Figure 2 curves and 95% bootstrap ribbons. In Polynesia, the Cook Islands have the largest positive LOO across q, read as the most differentiating unit, with continental admixture as the favored explanation. French Polynesia separates from Samoa and Tonga at q=1 (rare markers), in line with extra founder effects. In Micronesia, Guam, Kiribati, and Palau sit high on OVR with positive LOO; Nauru is lowest, with a small negative LOO, written as a homogenizer. In Melanesia, Near Oceania (Papua New Guinea, Solomon Islands) has elevated OVR; Remote Oceania (Fiji, New Caledonia) is lowest, matching prior gene-flow narratives. In Southeast Asia, the Andaman Islands are high at both q=1 and q=2; mainland groups such as Myanmar, Laos, and Vietnam contribute less on LOO.

The simulation side is closer to an audit tool. 1,432 unrelated African founders (HGDP plus 1000 Genomes, 322,216 sites) seed three demes (West, East, and Central-Southern-Northern Africa) for 17 generations, offspring Poisson(λ=3), between-deme mating controlled by a panmixia parameter ρ. After a ρ change at generation 8, OVR tracks the level of differentiation and LOO names which deme is carrying it. An isolation-reconnection pulse is cleaner: ρ drops to 0.05 at generation 8 and both OVR and LOO rise; ρ jumps to 0.9 at generation 14 and both curves fall. Isolation onset and contact resumption show up on the time axis.

Why it matters

Genomic simulators, neural ones included, get a second axis that is not collinear with FST. Matching q=2 is not enough; missing q=1 means the rare-allele behavior is wrong. For structure summaries, OVR gives the level, LOO gives attribution, and the slope is a coarse "recent drift versus old structure" diagnostic.

This patches FST rather than replacing it. q=2 is FST. The new piece is putting Shannon differentiation and variance-based fixation on one q-axis, then adding OVR and LOO cuts.

Limitations

The headline claims live in figures, so effect sizes are hard to check without a reimplementation. The locus model is biallelic SNPs only. Multiallelic sites, structural variants, and sequencing error never enter the formula. OVR forces equal weights and down-samples large groups, so small-island confidence bands will be wide. The simulator is a 17-generation monogamous pedigree with hand-written kinship bans; continuous migration in real populations is a different object. There is no head-to-head power study against FST, GST, or Jost's D, so "finer resolution" is still a spectral intuition plus case-by-case agreement, not a power proof. Linkage and multiple testing are not discussed.

Terms

Source

What people are saying

Related papers

All paper explainers