18 market encoders: similar rank IC, opposite latent geometry

Towards Financial World Modeling

Humzah Merchant, Alec Guthrie, Simon Mahns, Randall Balestriero, Bradford Levy

cs.LG

2026-10-07

On Market-1T, 18 encoders span 32 months of 1 Hz U.S. equities. Multihead return rank IC is 0.031; SSL loses to random ViT on spreads. Forecast-structure correlation is -0.19.

What problem this solves

A market world model needs a state that still works for decisions that were not specified at training time: the day's regime, a name's expected return, liquidity, volatility, and how names move together. Financial representation work usually scores one forecast, on a few months or a frozen list of stocks. Rank IC for a fixed encoder moves sharply across 2008-2025. A one-month leaderboard mostly measures whether that month was easy to forecast.

DINO-WM learns dynamics on a pretrained encoder, V-JEPA 2 adds actions, and LeWM can train the representation together with the dynamics. All three need a state that has been checked across regimes. This paper builds that layer only. Dynamics and planning are not trained.

Method

Market-1T, built by Chicago Booth with Massive, stores U.S.-listed equities from 2008 through 2025 at one-second resolution, nearly a trillion rows. 2008 is the first full year after Reg NMS. The source is the CTA and UTP securities information processors: the national best bid and offer plus consolidated trades. Each second has 11 fields: bid and ask price and size, open, high, low, close, VWAP, volume, and trade count. Those fields aggregate to coarser bars by standard market rules, with no interpolation.

The benchmark universe is rebuilt every month from the previous month only, which blocks survivorship and look-ahead. A name must be a common stock, with prior-month average VWAP above $5, average daily dollar volume of at least $20 million, and activity in more than 20% of regular-session seconds. About 700 names qualify on a typical day. Prices share one mean and one standard deviation so the spread is not normalized away. Sizes and volumes pass through log(1+x). Normalization statistics come only from the training window.

Evaluation uses 32 months across 2008-2025. Each model trains on the prior six months. Hyperparameters are chosen on five other months, up to about 125 tuning runs per method. The backbone is a ViT-Small. A ridge probe reads the final token; latent tests use mean pooling. Probe IC rises sharply with sample size, so the headline numbers use the full probe pool. An information token carries normalization statistics, start and end times, and aggregation scale. Frozen time-series foundation models do not receive it. Supervised training uses a pairwise learning-to-rank loss. A typical run takes 1 to 2 H100 hours.

Self-supervised runs match the supervised view count. The objectives are DINO, BYOL, CPC, I-JEPA, MAE, TS2Vec, CoST, TF-C, and TimeMAE. Five augmentations are compared mainly under LeJEPA, the most stable setup; DINO and BYOL default to time warping. Different crops of the same stock push the model to recognize the firm from bid and ask sizes. Time warping changes the aggregation scale. Cross-stock views at the same moment reward the common move. Same-industry pairs use the Fama-French 49 industries. Gaussian noise is the control.

Results

Rank IC is the Spearman correlation between forecast and outcome at a 15-minute horizon, averaged over the 32 months. The specialist row is three separate supervised heads, one per task.

MethodReturnVol. changeSpread change
Supervised multihead0.03110.09410.2455
Task specialist0.02720.08830.2517
TS2Vec0.02050.07190.1500
LeJEPA, time warp0.02040.07210.1691
Random ViT0.01700.06940.1692
DINO0.00970.03100.0626
TimesFM 3.0, frozen0.02320.09470.1362
Chronos-2, frozen0.02300.09280.1289

Multihead return IC is 0.0311, against 0.0272 for the return specialist. Volatility change is 0.0941, against 0.0883. The spread specialist leads, 0.2517 to the multihead's 0.2455. Return and volatility share structure; the shared head is slightly worse on spread. Each supervised head matches its own ridge probe.

SSL clears a randomly initialized ViT by a small margin on return and volatility, and loses on spread. The best SSL return IC is TS2Vec at 0.0205, against 0.0170. Time warping reaches 0.0721 on volatility change, against 0.0694. No SSL method beats the random ViT's 0.1692 on spread; time warping prints 0.1691. DINO's spread IC is 0.0626. BYOL and I-JEPA fall under the same floor.

Frozen TimesFM 3.0 scores 0.0947 on volatility, slightly above the multihead, and 0.0232 on return. Its spread IC is 0.1362. Chronos-2 scores 0.1289. Both trail the random ViT and trail a ridge ARDL on the raw series, which reaches 0.1857 on spread. A ridge on the last 24 tokens already reaches 0.0196 on return, in line with the best SSL. The supervised gap is largest on spread: 0.2517 for the specialist, 0.1820 for a 64-token ridge.

The evaluation month accounts for 97.2%, 98.9%, and 99.9% of variance on the three tasks. At a fixed FLOP budget, ViT-Tiny (5.4M) beats a ViT-Base (86M) trained with 10 times the FLOPs, on every task. Relative to a copy trained closer to the test month, volatility IC loses about 20% of its base level over five years, and spread about 10%. Return is more persistent.

Average forecast rank and average organization rank, across 19 encoders including the random ViT, have Spearman correlation -0.19. LeJEPA on two views of the same stock retrieves the firm at 5.36 times chance (mean percentile 2.5) and the trading day at 1.31 times chance (percentile 46). Firm centroids and day centroids are strongly negatively correlated. RankMe, the number of directions that carry the embedding's variance, is 2.8 to 3.7 for supervised models out of width 384, about 11 to 15 for LeJEPA, and 43.2 for a random ViT. The multihead is the only trained encoder that beats the random ViT on every latent probe. Among frozen foundation models, TimesFM 3.0 does too.

Across 2,448 portfolio implementations, universe, weighting, rebalance, and execution move Sharpe more than the encoder does. Execution is EFQ, effective half-spread over quoted half-spread: 0% is near a midpoint fill, 100% pays the full quoted spread. Once implementations are averaged, correlation of IC with Sharpe falls from +0.92 at 0% EFQ to +0.46 at 100%. Inside one fixed backtest it is only +0.52 to +0.29. At 100% EFQ, rebalance alone opens a 24.7 point range in annualized Sharpe.

Why it matters

What carries over is the protocol and the controls. Month explains more than 97% of the variance, so a recent-year board ranks the regime. A random ViT plus a ridge probe fit on the full pool is already a serious floor. DINO, BYOL, and I-JEPA fall through it.

Augmentation chooses the geometry. Same-asset views memorize the firm from quoted sizes and keep little of the day. Cross-stock views keep the day's common move. The multihead is the only trained model above the random floor on every structure probe, and its effective rank is about 3.7, closer to a low-dimensional forecast than to a state built for planning. At fixed compute, ViT-Tiny is the better fit. A run costs 1 to 2 H100 hours, so retraining is operationally realistic.

Compare representations with rank IC. Sharpe belongs first to the backtest. Measured EFQ is about 55% for direct market access and up to 88% on NYSE routes. Price and volume alone need not clear the spread.

Limitations

No dynamics model is trained, and DINO-WM, V-JEPA 2, and LeWM are not run on these states. The split between forecast rank and organization rank says a state should not be accepted on next-period return alone. It does not show that the organization probes improve planning.

The 18 strategies mix four supervised heads, five LeJEPA augmentations, and nine objectives. The augmentation sweep sits mostly on LeJEPA. Cross-stock views paired with DINO or MAE are missing, so some of the loss to a random ViT may be a mismatched augmentation.

The table covers liquid common stocks. Names under $5 or under $20 million of daily dollar volume are excluded, and so is depth past the best quote. Spread is already on the current quote. A random ViT reads it at 0.1692 and a 64-token ridge at 0.1820. An invariance objective can treat that as nuisance and drop it.

Other encoders are scored at the last layer only, so the foundation-model layer sweep is not a matched comparison. Trained encoders also receive an information token the foundation models do not. At publication, centralized hosting of the complete dataset was still in preparation; pretrained encoders were released separately. All 32 months still score roughly 15-minute price and volume.

Terms

Source

What people are saying

Related papers

All paper explainers