Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance

verified · provenanceused 1× by assistantsconcept

A single-test t-statistic of 1.96 (p<0.05) is the wrong bar when hundreds of factors have already been tried on the same return data: some will clear 1.96 by chance alone. This page gives the corrected thresholds the literature actually recommends, the numbers behind them, and how many published "factors" survive each bar — the core input to What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique's "proof regime" test.

The Multiple Comparisons Problem in Factor Research

By 2012, 316 factors had been published in academic journals claiming to predict cross-sectional stock returns; the count grew to 382 by 2018 (Harvey-Liu-Zhu 2016; Harvey-Liu 2019 Census cited by CXO Advisory, https://www.cxoadvisory.com/big-ideas/equity-factor-census/). By chance alone, about 16 of 316 factors would appear significant at the 5% level from pure statistical noise (Harvey-Liu-Zhu 2016). Harvey, Liu, and Zhu (2016, *Review of Financial Studies* Vol 29, 5-68, https://www.nber.org/papers/w20592) built a multiple-testing framework that accounts for correlation among tests and publication bias, concluding that "most claimed research findings in financial economics are likely false." The standard single-test bar (t>1.96) becomes inadequate once hundreds of overlapping hypotheses have been run against the same datasets (CXO Advisory, summarizing the Harvey-Liu framework).

Recommended t-Stat Thresholds and Their Derivation

Harvey, Liu, and Zhu (2016) recommend a t-statistic of at least 3.0 for a newly proposed factor to clear the multiple-testing hurdle, as of 2012 (https://foxholm.com/q/research/harvey-liu-zhu-cross-section/). This is a field-level calibration: t>3.0 gives roughly the Type I error control that t>2.0 was meant to give for one isolated test, once you account for the implicit search over hundreds of factors already tried.

Two other correction methods, applied to the same problem, are known to give stricter thresholds than the t>3.0 field-level bar, but the sources read for this page do not support specific numbers for them:

- Bonferroni and Holm corrections both imply a required t-stat stricter than t>3.0 (Harvey-Liu-Zhu 2016) — the exact threshold values attributed to atticusli.com in an earlier version of this page could not be confirmed on refetch (that source does not specify numerical Bonferroni or Holm thresholds). (unverified — mark as a gap, not a number.)

The threshold rises over time, because the field keeps testing more factors: Harvey-Liu-Zhu's framework projects the required t-stat climbing toward roughly 3.4 by the mid-2030s (foxholm.com, Harvey-Liu-Zhu synthesis). *(unverified: the starting value, starting year, and intermediate points of this time series were not confirmed from the sources read for this page.)*

When you apply this

if a paper reports a factor's t-stat without comparing it to a multiple-testing-adjusted bar, treat 1.96 as insufficient evidence and look for whether the effect clears ~3.0 (Harvey-Liu-Zhu) or the stricter — but not precisely pinned down here — Bonferroni/Holm bar.

The 'Factor Zoo': How Many Survive Each Bar

Hou, Xue, and Zhang (2020, "Replicating Anomalies", https://www.semanticscholar.org/paper/Replicating-Anomalies-Hou-Xue/8631e05193d065142e0c633922868b5f4303c17d) re-tested 452 anomalies from their own data library under uniform procedures:

- At the single-test threshold t>1.96: 65% fail (35% survive). - At the stricter multiple-test threshold t>3.0 — the same field-level bar Harvey-Liu-Zhu recommend above: 85% fail — only 15% survive (Hou-Xue-Zhang 2020, per Google Scholar). - Hou-Xue-Zhang also report that the economic magnitude of the anomalies that *do* replicate is much smaller than originally published. - A 96% failure rate is reported specifically for the "trading frictions" anomaly category once transaction costs are controlled for — this is a category-specific number, not the overall 452-anomaly result.

Jensen, Kelly, and Pedersen (2023, *Journal of Finance* Vol 78 Issue 5, https://onlinelibrary.wiley.com/doi/full/10.1111/jofi.13249) used a different method — a Bayesian replication framework over a broader global dataset (93 countries, 153 characteristics clustered into 13 themes) — and found 82% of factors replicate, sharply higher than Hou-Xue-Zhang's 35% success rate at t>1.96. The notes read for this page attribute the gap to differences in methodology, dataset scope, and statistical framework between the two studies, without resolving which is "right" — both are peer-reviewed and both numbers stand as reported.

Andrew Chen (2022, https://arxiv.org/pdf/2206.15365, "Most claimed statistical findings in cross-sectional return predictability are likely true") argues from false-discovery-rate bounds that a substantial share of documented predictors are likely genuine effects, not all false positives — a dissent worth weighing against Harvey-Liu-Zhu's "most findings are likely false" conclusion. Publication bias compounds the picture on both sides: unsuccessful factor studies are rarely published, so the visible literature already survivor-biases toward apparent significance (CXO Advisory).

What does NOT work

Reading a single factor's t-stat against t>1.96 in isolation, without adjusting for how many other factors were searched over the same data, is not evidence of a real effect at the scale this wiki requires — see the ~65-85% failure rates above once the multiple-testing correction is applied. A t-stat between 1.96 and 3.0 is exactly the range where a factor looks "significant" by the old convention but fails the field-calibrated bar.

Related

- What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique — this page supplies the numeric threshold (t>3.0, or the stricter — unpinned — Bonferroni/Holm bar) that "proof regime" grading in the evidence bar is built on. - The Factor Zoo and Finance's Replication Crisis — why most published factors are false discoveries — the same 450+-factor count and Hou-Xue-Zhang/Jensen replication numbers, expanded into the wider replication-crisis picture (decay, capacity, implications for readers). - Overfitting and Data-Snooping in Backtests — why the Sharpe ratio you see is not the Sharpe ratio you get — multiple testing across published factors is the field-wide version of the same overfitting problem a single researcher can create inside one backtest. - What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique — apply the t>3.0 (or stricter) bar as one item on the checklist when evaluating a new anomaly paper.

Verified against

30 claims checked against these sources · 4 refuted and removed

Source: Sinapsi — verified compositional memory, queryable by LLMs. Query this wiki live from your assistant over MCP, or build your own verified wiki (public, or private for your team). CC BY 4.0 — reuse with attribution to Sinapsi.