FactorBench backtests 5,000 AI-mined factors across 5 markets: no paradigm consistently wins

FactorBench: A Portfolio-Aware Benchmark for Automated Factor Mining

Zhuohan Wang, Carmine Ventre

q-fin.PM, cs.LG

2026-10-03

A shared 5-market evaluation of ~5,000 factors from 9 mining methods finds no paradigm dominating; hand-written Alpha101 from 2016 stays competitive at the portfolio level.

What problem this solves

Automated factor mining searches financial data for signals that predict future stock returns, and the field has cycled through four generations of methods: genetic programming evolving symbolic expressions, reinforcement learning and generative models learning to construct them, and most recently LLM agents that formulate hypotheses, write code, and iterate on backtest feedback. Every paper reports under its own data, objective, and backtesting pipeline, so numbers are not comparable across studies. The deeper problem: the score optimized during discovery is an intermediate proxy. A mining system delivers a pool of signals, and what matters is whether those signals survive out of sample and make money when combined. FactorBench (King's College London) puts that full chain under one contract.

Method

A four-block pipeline:

Two design choices carry the weight. Factor sign is fixed on training IC and never touched again, so out-of-sample sign reversals show up raw. And residual IC regresses each factor against beta, volatility, momentum, short-term reversal, and liquidity before re-measuring, separating new information from rediscovered known styles.

Results

Of 4,914 candidates, 98.5% execute and 95.6% are scorable (at least 100 valid dates per split). All 76 execution failures sit in the LLM pipelines (19 R&D-Agent, 24 AlphaAgent, 33 QuantaAlpha); symbolic methods lose output mostly to numerical degeneracy. Runtime spans 0.14 hours per run for GP to 27.9 for QuantaAlpha.

On temporal generalization, GP, RL, and generative methods post much higher training IC that largely evaporates by validation and does not return in testing. LLM methods show smaller train-to-test gaps because they had less in-sample performance to lose; their held-out IC is not systematically higher.

Portfolio level (test period, ridge composite, 20-day rebalancing, 10 basis points per dollar traded):

MethodComposite ICResidual ICLong-only CAGR / SharpeLong-short CAGR / Sharpe
Alpha101 (reference)0.02380.023726.5% / 1.242.6% / 0.32
GP0.0165-0.012830.4% / 1.173.7% / 0.34
AlphaQCM0.01910.017529.6% / 1.313.7% / 0.67
R&D-Agent0.02530.001131.2% / 1.196.5% / 0.68
AlphaAgent0.0030-0.002522.4% / 1.030.3% / -0.09

Every long-only book is positive (22.4%-31.2% CAGR), but long-short returns collapse to 0.3%-6.5%, so most of the long-only profit is market beta. R&D-Agent has the highest raw composite IC and sees it fall from 0.0253 to 0.0011 after neutralization; GP, AlphaGen, AlphaForge, AlphaSAGE, and AlphaAgent all flip to negative residual IC. AlphaQCM keeps most of its raw IC, and Alpha101 barely moves. In the long-short setting R&D-Agent and AlphaQCM take the best Sharpe ratios while AlphaAgent goes slightly negative. Alpha101, hand-written in 2016, stays competitive at every level.

At the pool level, LLM-generated factor pools sit closer to published Alpha101 formulas than non-LLM search does, most strongly for R&D-Agent; AlphaGen and AlphaQCM carry the most internal redundancy.

Why it matters

For AI4Finance practitioners this is a calibration paper. LLM agents mining factors look like a shortcut to automated research, but under unified evaluation their out-of-sample edge is absent, a sizable share of their raw predictiveness is measured style exposure, and their output tracks public formulas. Any factor-mining result reported only on its search objective can be re-audited with this three-level protocol. The benchmark is open-sourced with shared fields and splits across five markets, which makes it a usable reproduction base. No new method here, but the negative result is the information.

Limitations

The authors' own list is thorough: conclusions cover the evaluated implementations under the shared task, not the fundamental capabilities of the paradigms; five large-cap markets, one 20-day target, one chronological split, three seeds; budgets follow each method's original paper, so wall-clock comparisons are not controlled; residual IC controls only five exposures and is not pure alpha; no industry classifications; static universes introduce survivorship bias; cost and borrow assumptions are simplified.

Further caveats from a close read: the data is daily Yahoo Finance, reproducible but far from institutional-grade; all three LLM methods share one Qwen3.8-27B endpoint, so the finding says nothing about larger or closed models; the long-only profits concentrate in the 2024-2025 test window and may not survive a bear-market split; R&D-Agent searches over Python programs while the others search expressions, an asymmetry the contract smooths only at evaluation; the Alpha101 reference covers 71 of 101 formulas.

Terms

Source

What people are saying

All paper explainers