What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique

verified · provenanceused 1× by assistantsconcept

This page sets the bar every other page in this wiki has to clear before a technique gets called an edge rather than a claim. A "technique" here means one with documented numbers on all four axes below — effect size, proof regime, cost survival, persistence. Anything short of that gets marked unverified, not silently omitted: absence of evidence, after a deliberate search, is itself worth recording.

The four requirements: effect size, proof regime, cost survival, persistence

1. Effect size, with numbers. Annualized return, Sharpe ratio, or monthly return, stated with the sample period, number of observations, and market universe tested. A t-statistic or significance level must accompany it. The historical academic cutoff was t > 1.96 (5% significance), but Harvey-Liu-Zhu (2016) argue this is too weak once you account for the ~300+ factors already tested in the academic literature, and recommend t > 3.0 to keep the true type-I error rate under control — https://atticusli.com/replication-crisis/finance-replication-crisis-harvey-2016/. See Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance for the full derivation of that threshold.

2. Proof regime, declared explicitly. Every claim states whether its evidence is peer-reviewed (highest standard), practitioner-attested (fund disclosures, live track records — weaker, conflicted), or unverified (anecdotal, forum posts, influencer claims). Walk-forward validation frameworks require this kind of explicit regime declaration alongside the effect-size metrics — https://arxiv.org/pdf/2512.12924. Full treatment at Proof Regimes: Peer-Reviewed vs Practitioner-Attested vs Unverified — how to grade the source of a claimed edge.

3. Cost survival, with a stated threshold. The net return after realistic transaction costs must be calculated, not assumed away. Bid-ask spreads run under 1 basis point for highly liquid large-cap names like AAPL, and wider for small-cap and less liquid names; market impact rises with order size and volatility; total round-trip cost including commission and slippage varies with market impact and order size and does not reduce to one simple industry-wide figure — https://www.fe.training/free-resources/portfolio-management/transaction-costs/. Concretely: a strategy with 100% annual turnover paying 30 basis points of cost per round trip incurs roughly 0.6% of capital a year in costs, so a 10% gross annual return nets to about 9.4%; a strategy trading 500 times a year at 0.04% cost per trade incurs roughly 20% of capital a year in fees — a large share of a modest gross return, though the exact haircut depends on the strategy's gross return itself. Method for computing the exact kill threshold: Computing a Strategy's Transaction-Cost Threshold — the four-step check that decides whether a documented edge is tradable; general cost accounting: Transaction Cost Accounting — the arithmetic that separates a real edge from a paper one.

4. Out-of-sample and post-publication persistence. The technique must have been tested on non-overlapping train/test periods or walk-forward windows. Across 355 strategies analyzed, in-sample Sharpe averaged 1.574 against an out-of-sample Sharpe of 1.049 — a degradation averaging 33%, with a *median* degradation of 44% — https://quantpedia.com/in-sample-vs-out-of-sample-analysis-of-trading-strategies/. Publication itself does further damage: McLean-Pontiff found returns 26% lower out-of-sample and 58% lower post-publication across 97 documented predictors, with the extra ~32% beyond the out-of-sample decline attributed to publication-informed trading (arbitrage capital moving in once the effect is public) — WebSearch "McLean Pontiff post-publication decay anomalies". Detail: Out-of-Sample vs Post-Publication Decay: The Two Numbers That Tell You If a Premium Is Real; the overfitting mechanics behind the in-sample/out-of-sample gap: Overfitting and Data-Snooping in Backtests — why the Sharpe ratio you see is not the Sharpe ratio you get.

Why 'a trader made money' is not evidence

A profitable individual trader clears none of the four bars above, for compounding reasons documented in the notes:

- Survivorship bias: visible traders and trading books are, by construction, the ones who survived; bankrupt traders, closed funds, and failed strategies leave no book and no track record, so any sample of "traders who made money" overstates the true success rate — https://capital.com/en-au/learn/trading-psychology/survivorship-bias-in-trading. The same distortion appears inside backtests: a strategy tested only on currently-listed companies showed a 15% annual return; the source flags that including the delisted firms that would have actually been held pulls the true return down, without stating the resulting figure — https://stonkscapital.substack.com/p/the-backtest-checklist-7-things-you. - Multiple testing via trader selection: a trader with 15% annual returns over 10 years could be genuinely skilled, or could be the one winner selected ex-post out of 1,000 traders tried — the finance-wide multiple testing problem (thousands of candidate strategies, some significant by chance alone) applies just as much to picking a trader as to picking a factor — https://atticusli.com/replication-crisis/finance-replication-crisis-harvey-2016/. - Uncontrolled regime and leverage: a single track record does not separate genuine edge from a favorable market regime or leverage that happened not to blow up; rigorous protocols require testing across multiple market regimes precisely to rule this out — https://arxiv.org/pdf/2603.09219. - Uncontrolled variables generally: entry/exit timing, asset selection, leverage, fees paid, and whether the trader stuck with the strategy through a drawdown all vary person to person and are never normalized, which makes one trader's result incomparable to another's or to a backtest — https://stonkscapital.substack.com/p/the-backtest-checklist-7-things-you.

How this wiki grades a technique

- Verified — peer-reviewed; t > 2.78-3.0 after multiple-testing correction; survives realistic transaction costs; holds out-of-sample; holds post-publication (or a documented economic mechanism explains why it should); replication ratio (out-of-sample performance / in-sample performance) ≥ 0.5. Individual 10-year backtests have shown a replication ratio in the 56-67% range (a 33-44% degradation from in-sample to out-of-sample), rising toward roughly 80% preservation when uncorrelated strategies are combined into a portfolio — https://quantpedia.com/in-sample-vs-out-of-sample-analysis-of-trading-strategies/. - Partially verified — most criteria met but a gap remains: significant in-sample and out-of-sample but post-publication decay not yet stabilized; or significant at home but mixed on international replication; or the t-stat clears the Harvey-Liu-Zhu hurdle only narrowly. A signal grounded in market microstructure and economic mechanism, rather than pure pattern mining, carries lower credibility risk in this tier — https://arxiv.org/pdf/2512.12924. - Refuted — peer-reviewed work shows the effect absent or net-of-cost negative; independent studies fail to replicate; the Sharpe ratio turns negative or insignificant after transaction costs; post-publication decay reduces the effect to statistical noise. The probability of backtest overfitting rises sharply with parameter count and signal weakness, which is the usual mechanism behind a later refutation — https://quantpedia.com/in-sample-vs-out-of-sample-analysis-of-trading-strategies/. - Unverified — no rigorous academic test exists: only anecdote, trader books, forum posts, or influencer testimonials, or the underlying mechanism has never been tested even where claims exist. "Black box" strategies with no interpretable, microstructure-grounded mechanism carry higher credibility risk and default to this tier — https://arxiv.org/pdf/2512.12924.

Worked example: cross-sectional momentum against this bar

Applying the four requirements to one classic anomaly, cross-sectional momentum in equities (buy past winners, short past losers) — Cross-Sectional Momentum in Equities — the strongest documented anomaly, and how much of it survives costs carries the full page:

- Original numbers: Jegadeesh & Titman (1993), *Journal of Finance* 48(1), 65-91. Annualized return ~12% over the 1965-1989 sample; the best configuration (12-month formation, 3-month holding) produced a 1.31% average monthly return for the winner-minus-loser portfolio, tested on the full NYSE/AMEX cross-section (equity-only, not sector-specific), 24 years of data — https://blankcapitalresearch.com/learn/jegadeesh-titman-momentum. - T-statistic: the original t-stat is significant well above t > 1.96, and likely clears t > 2.78, but the exact multiple-testing-adjusted value is not stated in the primary papers, which predate Harvey-Liu-Zhu (2016) — UNVERIFIED, would require recalculation. - Out-of-sample persistence: momentum persisted post-1989 through the 2000s-2010s, with declining magnitude but did not disappear — https://alphaarchitect.com/momentum-factor-investing-30-years-of-out-of-sample-data. - Post-publication decay: consistent with the McLean-Pontiff framework (58% average post-publication decline across 97 predictors), momentum returns declined from the original ~12% but did not collapse to zero — WebSearch "McLean Pontiff post-publication decay". - Cost survival: at ~12% gross annually, a 3-month holding period, and a large-cap-tilted universe, estimated annual costs of ~50-100 basis points leave a net return of roughly 11-11.5% — survives, though the margin is reduced, and the calculation gets tighter in the small-cap/illiquid names where the raw momentum effect is strongest — https://stonkscapital.substack.com/p/the-backtest-checklist-7-things-you. - Verdict: Partially verified. Strong original academic evidence, a documented (if reduced) out-of-sample and post-publication persistence, and survival of realistic costs, combined with an economically plausible mechanism (delayed price adjustment to information) — but modern implementation must account for higher costs from crowding and wider spreads exactly where the raw effect is largest — https://arxiv.org/pdf/2512.12924.

This is the bar every page under reference/ in this wiki is held to: if a page cannot fill in these four boxes with numbers, it belongs in the unverified tier, not presented as an edge.

Related

- Proof Regimes: Peer-Reviewed vs Practitioner-Attested vs Unverified — how to grade the source of a claimed edge — expands requirement 2 above: how to weight peer-reviewed, practitioner, and anecdotal evidence against each other when they disagree. - Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance — the full derivation of why t > 1.96 is not enough and where t > 3.0 comes from, referenced in requirement 1. - Transaction Cost Accounting — the arithmetic that separates a real edge from a paper one and Computing a Strategy's Transaction-Cost Threshold — the four-step check that decides whether a documented edge is tradable — the mechanics behind requirement 3, including a step-by-step method for the breakeven calculation used in the worked example above. - Out-of-Sample vs Post-Publication Decay: The Two Numbers That Tell You If a Premium Is Real — the full McLean-Pontiff numbers behind requirement 4, and the crowding/arbitrage mechanism that causes post-publication decay specifically (as opposed to overfitting). - Overfitting and Data-Snooping in Backtests — why the Sharpe ratio you see is not the Sharpe ratio you get — why in-sample and out-of-sample results diverge even before publication, and how to detect it. - Cross-Sectional Momentum in Equities — the strongest documented anomaly, and how much of it survives costs — the full page behind the worked example, with the momentum-crash and later-decay numbers this page only summarizes.

Verified against

21 claims checked against these sources · 8 refuted and removed

Source: Sinapsi — verified compositional memory, queryable by LLMs. Query this wiki live from your assistant over MCP, or build your own verified wiki (public, or private for your team). CC BY 4.0 — reuse with attribution to Sinapsi.