The Factor Zoo and Finance's Replication Crisis — why most published factors are false discoveries

verified · provenanceused 1× by assistantsconcept

Academic finance has published 300-450+ candidate return "factors," and systematic replication tests find that most of them fail to hold up once the multiple-testing problem is accounted for. This page gives the scale of the problem, the numbers behind the failure rates, and the one dissenting large-scale study that argues the picture is less bleak than it looks — with what remains unresolved between the two camps.

Scale of the problem: 300-450+ published factors

- 316 distinct factors were documented across published academic finance and accounting papers from the 1960s through 2012, identified from 313 published papers — Harvey, Liu, Zhu (2016), *…and the Cross-Section of Expected Returns*. - By 2019 the census reached 382 documented factors in top journals — Harvey & Liu (2019), *A Census of the Factor Zoo*. Growth in factor-related publications has been nearly exponential since 2008, driven largely by publication demand. - 452 anomalies/factors were tested in the largest systematic replication effort to date, spanning U.S. equity data from 1967 to 2016 — Hou, Xue, Zhang (2020), *Replicating Anomalies*.

Replication results: 65-82% fail t-stat thresholds

- 65% of the 452 anomalies in Hou-Xue-Zhang (2020) fail the standard single-test hurdle of |t| ≥ 1.96 (the conventional 5% significance level) when replicated out-of-sample. - 82.1% fail the stricter multiple-testing-corrected hurdle of |t| ≥ 2.78. Correcting for the size of the factor library sharply raises the failure rate. - Failure is not uniform: 95 of 102 "trading frictions" anomalies (93%) fail the 1.96 threshold — category matters as much as the aggregate number. - Harvey, Liu, Zhu (2016) estimate that once multiple-testing correction is applied across the 316 factors they catalogued, roughly 71% are likely false discoveries. Their illustrative point: testing 316 factors would produce about 16 "significant" results at the 5% level by chance alone, even if no true factor existed. - Feng, Giglio, Xiu (2020), *Taming the Factor Zoo*, used a model-selection test of marginal contribution: of 99 new factors tested against existing high-dimensional factor sets, only 14 added significant explanatory power beyond what was already captured; the other 85 were redundant or spurious.

When these numbers apply

they are U.S. cross-sectional equity anomalies tested against the historical record already in the literature (Hou-Xue-Zhang) or catalogued from published papers (Harvey-Liu-Zhu). They describe the base rate for *already-published* claims, not a general prior for any new backtest — see Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance for how to size the correction for a specific number of trials.

Is There a Replication Crisis in Finance? (Jensen, Kelly, Pedersen 2023) — the dissent

- Jensen, Kelly, Pedersen (2023) built a Bayesian replication model using 153 factors across 93 countries, clustered into 13 themes — the largest global replication test to date. - Replication rates exceeded 75% in 10 of the 13 themes. The three weaker themes were "seasonality," "leverage," and "size," but the paper's headline is that most themes replicate successfully internationally. - Their conclusion runs opposite to Harvey-Liu-Zhu: they argue the sheer number of observed factors in the zoo *strengthens* the evidence for factors overall, not weakens it, because international out-of-sample testing across 93 countries confirms most factors work beyond the U.S. sample where they were discovered. - UNVERIFIED (not independently confirmed in sources reviewed): Jensen et al. also claim most factors are "significant parts of the tangency portfolio" — an economic-significance claim beyond statistical significance that this page does not verify further.

The unresolved tension — what it means for reading any single number

- Hou-Xue-Zhang (2020, frequentist, 65-82% failure) and Jensen-Kelly-Pedersen (2023, Bayesian, >75% success in 10/13 themes) reach opposite headline conclusions from overlapping literatures. UNVERIFIED: the literature has not fully adjudicated whether the gap is driven by methodology (Bayesian vs. frequentist), sample period, geographic scope (U.S.-only vs. 93 countries), or how factors are defined/clustered. Treat any single number from either paper as regime-dependent, not as a settled verdict on "factors in general." - Practical implication for this wiki's objective: a factor's raw t-statistic or "X% of factors replicate" headline is not by itself gradeable evidence — see What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique for the full bar (effect size, proof regime, cost survival, out-of-sample persistence) this wiki applies before marking anything verified. - Harvey, Liu, Zhu (2016) recommend a t-statistic threshold of 3.0 (not the conventional 2.0) for any newly proposed factor to clear the multiple-testing hurdle, corresponding to roughly p < 0.003 instead of p < 0.05 — this is the concrete number to apply when reading a new anomaly paper; the derivation and worked application are detailed in Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance. - A factor paper that does not explicitly address transaction costs, out-of-sample testing, or post-publication decay should be read skeptically regardless of its reported t-stat — see Out-of-Sample vs Post-Publication Decay: The Two Numbers That Tell You If a Premium Is Real for the McLean-Pontiff (2016) numbers on how much of a factor's return typically evaporates after publication, and Transaction Cost Accounting — the arithmetic that separates a real edge from a paper one for the cost side. - Any anomaly tested only in-sample (same market/period where it was discovered) will show inflated effect sizes and t-statistics relative to genuine out-of-sample tests; this is implicit in the entire replication methodology above and is the single most common red flag when What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique is applied to a new paper.

Why this page matters for the wiki's objective

The factor zoo is the base rate this wiki's grading bar exists to counter: with 300-450+ published candidates and only a minority surviving multiple-testing-corrected replication, "a paper found significance" is not evidence of a durable edge on its own. Every reference page in this wiki for an individual factor or anomaly should be read against these two studies' numbers — did the specific factor clear a corrected threshold, and does an independent replication (ideally out-of-sample, ideally international per Jensen et al.) exist for it.

Verified against

30 claims checked against these sources · 1 refuted and removed

Source: Sinapsi — verified compositional memory, queryable by LLMs. Query this wiki live from your assistant over MCP, or build your own verified wiki (public, or private for your team). CC BY 4.0 — reuse with attribution to Sinapsi.