Out-of-Sample vs Post-Publication Decay: The Two Numbers That Tell You If a Premium Is Real
Two numbers separate a discovered artifact from a discovered, then arbitraged, real premium: how much a factor's return falls the moment you leave the sample it was mined on, and how much more it falls once the paper is public. McLean & Pontiff (2016) measured both, on 97 predictors, and the gap between them is the closest thing finance has to a decay signature you can read off a table.
McLean-Pontiff 2016: The Quantified Decay
97 cross-sectional return predictors from the academic literature were reconstructed and tracked in three windows: in-sample, out-of-sample-but-pre-publication, and post-publication (https://onlinelibrary.wiley.com/doi/abs/10.1111/jofi.12365).
- Portfolio returns are 26% lower out-of-sample than in-sample. - Portfolio returns are 58% lower post-publication than in-sample. - The gap attributable to publication itself: 58% − 26% = 32 percentage points.
The out-of-sample decline (26%) is the upper-bound estimate of pure data-mining/overfitting effects — it happens with zero public knowledge involved, so it cannot be arbitrage (https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2156623). The additional 32-point drop after publication is the effect of the paper existing: capital learns and trades against it. Roughly 50% of anomaly alpha disappears post-publication overall (https://onlinelibrary.wiley.com/doi/abs/10.1111/jofi.12365).
A secondary signature of genuine crowding: correlations among published-predictor portfolios rise after publication, consistent with coordinated market response rather than independent overfitting artifacts fading out (https://onlinelibrary.wiley.com/doi/abs/10.1111/jofi.12365).
Mechanisms: Overfitting vs Crowding/Arbitrage
Two decay mechanisms, two timings — use the timing to diagnose which one you're looking at:
- Data-mining artifact: the premium collapses immediately at the out-of-sample boundary, before anyone outside the research team has read the paper. This is the diagnostic signature of a statistical fluke that never had real predictive power (https://microalphas.com/factor-decay/). - Arbitrage of a genuine effect: the premium survives the out-of-sample period intact, then fades specifically after publication. This is the diagnostic signature of a real but exploitable mispricing being competed away — when capital learns a real premium exists, it floods toward it, prices adjust, and forward returns fall (https://microalphas.com/factor-decay/).
Caveat from the same source: the majority of documented anomalies fail to replicate under careful, uniform re-testing at all, meaning many never had genuine predictive power in the first place — they are statistical artifacts of excessive testing, not premiums that later got arbitraged (https://microalphas.com/factor-decay/). See Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance and The Factor Zoo and Finance's Replication Crisis — why most published factors are false discoveries for the scale of that testing problem.
Which Factors Decay Less and Why (Arbitrage Cost Link)
McLean & Pontiff's core cross-sectional finding: portfolios concentrated in high idiosyncratic risk and low liquidity stocks show less post-publication decay (https://onlinelibrary.wiley.com/doi/abs/10.1111/jofi.12365). The mechanism is direct — high trading costs and high idiosyncratic risk limit how much arbitrage capital can flow in to close the gap, so the mispricing survives longer even after it's public (https://onlinelibrary.wiley.com/doi/abs/10.1111/jofi.12365). This is the same logic as Transaction Cost Accounting — the arithmetic that separates a real edge from a paper one: cost isn't just a drag on your net return, it's the barrier that determines whether a real premium can survive publication at all.
Jacobs & Müller (2020) extend this: barriers to arbitrage trading may create segmented markets, where the same friction that protects a premium domestically also explains why it behaves differently abroad (https://www.sciencedirect.com/science/article/abs/pii/S0304405X19301618) — see next section.
International Evidence: The US Exception
Jacobs & Müller (2020), testing 241 anomalies across 39 stock markets and 2+ million anomaly country-months (https://www.sciencedirect.com/science/article/abs/pii/S0304405X19301618):
- The United States is the only country in the sample with a reliable post-publication decline in long-short returns. - Outside the US, international return predictability stays the same or increases after the original sample period ends. - Interpretation offered: anomalies elsewhere more often represent genuine, persistent mispricing rather than data-mining artifacts — publication does not appear to cause decay globally the way it does in the US.
Practical implication
a decay pattern that looks convincing in a US-only study does not automatically generalize. If a claim about a factor is US-only and hasn't been checked internationally, that is a real gap, not a formality — checking whether decay is geography-specific is one of the questions in the next section.
Red Flags in Backtests
From the literature on questionable research practices (https://www.psychiatrist.com/jcp/harking-cherry-picking-p-hacking-fishing-expeditions-and-data-dredging-and-mining-as-questionable-research-practices/):
- Sharpe ratio above 3.0 on daily data — statistically implausible. - Maximum drawdown implausibly small relative to average trade size — overfitting symptom. - Near-zero losing months or weeks — a data-mining signal. - Performance degrades sharply when the test window is shifted by a few weeks — timing cherry-picking. - Narrow time windows chosen without a stated reason — lookback bias.
Publication and Methodology Questions
- Does the paper connect the empirical result to existing economic theory, or is it purely empirical pattern-matching? (https://www.psychiatrist.com/jcp/harking-cherry-picking-p-hacking-fishing-expeditions-and-data-dredging-and-mining-as-questionable-research-practices/) - Is there genuine out-of-sample testing spanning a meaningful period, not just an in-sample fit? (same source) - Timing check: does decay occur years *before* publication in the underlying literature? If so, publication is not the cause — the effect was already dying (https://www.sciencedirect.com/science/article/abs/pii/S0304405X19301618). - Geography check: does the factor work in multiple markets/asset classes? Jacobs & Müller show the US post-publication decay pattern is highly geography-specific, so a US-only "it decayed" or "it didn't decay" claim needs a cross-market check before you trust it (https://www.sciencedirect.com/science/article/abs/pii/S0304405X19301618).
The Momentum Debate: A Case Study
Momentum is the clearest real-world case where the "did it decay" question stays open and contested:
- Momentum returned approximately 12% annualized in the 1990s, following Jegadeesh & Titman's original 1993 study (https://blankcapitalresearch.com/learn/jegadeesh-titman-momentum). - Modern momentum returns have declined to approximately 2% annualized today (https://blankcapitalresearch.com/learn/jegadeesh-titman-momentum). - Yet Jegadeesh & Titman's own 2022 replication reports momentum profits remain "large and significant" across 30 years post-publication and 40+ countries (https://link.springer.com/article/10.1007/s11408-022-00417-8) — the authors' own follow-up directly contradicts a simple decay narrative. - (unverified) Whether the documented 10%→2% decline reflects genuine publication-driven decay or a regime change (growth-stock dominance, changed short-squeeze mechanics) remains debated in the notes reviewed; no source here resolves it.
This is the concrete illustration of the whole page's method: two credible sources disagree on the same factor because they're measuring different things (aggregate return level vs statistical significance of the long-short spread across decades/countries) — you have to read what was actually measured, not just the verdict. See Cross-Sectional Momentum in Equities — the strongest documented anomaly, and how much of it survives costs for the full numeric record on momentum specifically.
Replication Crisis Numbers: Baseline for Skepticism
Before trusting any single decay number, know the scale of the underlying testing problem:
- By 2012, 300+ variables had been published as return predictors, each individually tested at a t>1.96 threshold (Harvey, Liu & Zhu 2016, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2249314). - At t>1.96, roughly 15 of 300 factors are expected to appear "significant" by chance alone — the multiple-testing / Bonferroni intuition applied to the factor zoo. - The recommended corrected threshold is t>3.0, not t>1.96 (same source). - Hou, Xue & Zhang (2020) re-tested 452 anomalies with a uniform methodology (https://global-q.org/uploads/1/2/2/6/122679606/houxuezhang2020rfs.pdf): - 65% fail to clear t>1.96 (only 35% replicate). - Failure rises to 82% at the stricter t>2.78 threshold (the 5% multiple-testing-corrected level). - Replication success by category: momentum 87.7%, value 75.4%, investment 94.7%, profitability 6%, intangibles 44.7%, trading frictions 41.5%.
Full detail on this problem lives in Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance and The Factor Zoo and Finance's Replication Crisis — why most published factors are false discoveries; the numbers are repeated here because a "post-publication decay" claim is meaningless without knowing what fraction of published factors were never real in the first place.
What does NOT work
Treating a post-publication return drop as automatic proof of "arbitrage killed a real premium" does not work: the same drop is equally consistent with the factor never having genuine predictive power (a data-mining artifact whose apparent decay is coincidentally timed near publication). The McLean-Pontiff framework only distinguishes the two mechanisms cleanly when you also have the pre-publication, out-of-sample number — decay observed only post-publication, with no matching out-of-sample decline, is the stronger case for genuine crowding; decay observed already out-of-sample is the signature of overfitting (https://microalphas.com/factor-decay/). A single post-publication return series, without the out-of-sample comparison point, is not enough evidence either way.
Trusting a "still works internationally" or "still works in the US" claim without checking geography also does not work, given Jacobs & Müller's finding that the US is the outlier, not the norm, for post-publication decline (https://www.sciencedirect.com/science/article/abs/pii/S0304405X19301618).
Transaction Cost Survival: The Final Filter
Even a factor that survives both out-of-sample and post-publication tests on paper still has to clear net-of-cost profitability, which is the test that eliminates most strategies in practice (https://alphaarchitect.com/the-factors-that-plague-factor-investing/). Related numbers on the decay-cost link:
- Publication year alone explains 30% of the variance of Sharpe ratio decay across factors (https://microalphas.com/factor-decay/). - Average post-publication Sharpe decay of newly published factors increases roughly 5 percentage points per year, cumulatively (https://microalphas.com/factor-decay/).
Full method for computing the exact cost threshold that kills a given strategy is in Transaction Cost Accounting — the arithmetic that separates a real edge from a paper one and Computing a Strategy's Transaction-Cost Threshold — the four-step check that decides whether a documented edge is tradable.
Why this matters for the wiki's objective
This page is the wiki's core diagnostic for telling apart a real, surviving edge from a mined artifact: the objective requires knowing, with numbers, whether an effect held out of sample and after publication, and this page supplies the exact quantified baseline (26% / 58% / 32-point gap) plus the timing-based method (immediate collapse = overfitting; post-publication-only fade = arbitrage) that every factor page in Cross-Sectional Momentum in Equities — the strongest documented anomaly, and how much of it survives costs, The Value Factor (Book-to-Market / HML) — a founding premium whose post-1991 numbers can't rule out zero, The Size Factor (Small-Minus-Big) — a textbook case of post-publication decay this wiki uses as a yardstick and the rest should be checked against before being graded verified, partially verified, or refuted per What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique.
Related
- What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique — the four-requirement evidence bar this page's decay test feeds directly into: out-of-sample/post-publication persistence is one of the four required numbers for any page in this wiki. - Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance — the replication-crisis numbers summarized here (t-stat thresholds, factor-zoo failure rates) are detailed fully on that page; read it before trusting any single factor's decay claim. - Transaction Cost Accounting — the arithmetic that separates a real edge from a paper one — the arbitrage-cost mechanism that explains *why* some factors decay less is the same friction detailed there as the final filter on net profitability. - Cross-Sectional Momentum in Equities — the strongest documented anomaly, and how much of it survives costs — the momentum case study above is a summary; the full numeric record (12% original, formation/holding windows, crash magnitude) lives on that page. - Overfitting and Data-Snooping in Backtests — why the Sharpe ratio you see is not the Sharpe ratio you get — the immediate-collapse diagnostic described here (data-mining artifact) is the same phenomenon that page covers from the backtesting-methodology side (deflated Sharpe, walk-forward validation).
Verified against
52 claims checked against these sources · 3 refuted and removed
- onlinelibrary.wiley.com/doi/abs/10.1111/jofi.12365
- papers.ssrn.com/sol3/papers.cfm
- sciencedirect.com/science/article/abs/pii/S0304405X19301618
- microalphas.com/factor-decay
- papers.ssrn.com/sol3/papers.cfm
- global-q.org/uploads/1/2/2/6/122679606/houxuezhang2020rfs.pdf
- link.springer.com/article/10.1007/s11408-022-00417-8
- blankcapitalresearch.com/learn/jegadeesh-titman-momentum
- alphaarchitect.com/the-factors-that-plague-factor-investing
- psychiatrist.com/jcp/harking-cherry-picking-p-hacking-fishing-e…
What links here
Source: Sinapsi — verified compositional memory, queryable by LLMs. Query this wiki live from your assistant over MCP, or build your own verified wiki (public, or private for your team). CC BY 4.0 — reuse with attribution to Sinapsi.