Proof Regimes: Peer-Reviewed vs Practitioner-Attested vs Unverified — how to grade the source of a claimed edge
Before this wiki accepts a number for a technique, it has to know which regime produced it: a peer-reviewed paper, a practitioner track record, or an unverified claim. The three regimes carry different, non-overlapping failure modes — knowing which one you're reading tells you which bias to check for before trusting the number.
Peer-reviewed academic studies: what they guarantee and what they don't
Peer review means 2-3 field experts screen a manuscript against scientific standards before publication, with authority to block work that doesn't meet them — https://libguides.tulane.edu/c.php?g=1513936&p=11324740.
What it does NOT guarantee — multiple testing
Harvey, Liu, and Zhu (2016) reviewed 316 published factors claiming to predict cross-sectional stock returns and showed the conventional t > 2.0 threshold is far too loose once you account for testing hundreds of factors across half a century of research; their headline recommendation, t > 3.0, is now the de facto bar for new factors in academic finance — https://www.nber.org/papers/w20592, https://atticusli.com/replication-crisis/finance-replication-crisis-harvey-2016/. See Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance for the mechanics of why the threshold has to move.
What it does NOT guarantee — replication
Hou, Xue, and Zhang (2020) retested 202 characteristic signals (452 factor portfolios) and found only 35% cleared acceptable empirical standards — 65% failed to replicate — https://replicationnetwork.com/2017/06/14/hou-xue-zhang-replication-controversies-in-finance-accounting/. Jensen, Kelly, and Pedersen (2023) extended the test to 93 countries and found an 82.4% replication rate, a large gap from Hou-Xue-Zhang that shows replication success is itself sensitive to methodological choices — https://onlinelibrary.wiley.com/doi/10.1111/jofi.13249. Neither number should be read as "the" replication rate; both are evidence the rate depends heavily on how strict the test is. See The Factor Zoo and Finance's Replication Crisis — why most published factors are false discoveries for the full scale of the problem.
What it does NOT guarantee — practitioner relevance
a study of 6,000+ papers found practitioners rated academic research relevance at only 1.11 on a 5-point scale, and 90% of marketing articles received zero news citations; the academic-practitioner correlation varies by field — strong in International Business (0.86) and Entrepreneurship (0.76), weak in Information Systems (0.07) and, notably for this wiki, Finance (0.43) — https://pmc.ncbi.nlm.nih.gov/articles/PMC10699644/. Publication itself is also biased toward positive results: statistically significant findings are more likely to be published, published fast, and published more than once — https://methods.cochrane.org/bias/reporting-biases — though bias detection has been improving in more recent meta-analyses — https://onlinelibrary.wiley.com/doi/10.1002/sim.6525.
Reading rule for this wiki
a peer-reviewed number is a starting point, not a verdict. Check its t-stat against 3.0 (not the older 2.0 convention), check whether it has been independently replicated, and treat "published" as "survived one filter," not "true."
Practitioner track records and fund disclosures: strengths and conflicts of interest
Practitioner evidence has one strength academic papers can't match: live capital, real execution, real costs. It has three biases academic evidence mostly avoids.
Selection bias
hedge fund managers choose whether to report to commercial databases, and funds that do report significantly outperform non-reporters — reporting itself is correlated with having good numbers to show, which inflates the apparent skill of "the hedge fund universe" — https://academic.oup.com/rfs/article-abstract/26/1/208/1592622, https://breakingdownfinance.com/finance-topics/alternative-investments/hedge-fund-database-biases/.
Survivorship and delisting bias
managers have discretion over when a failing fund is removed from a database, so failures are often not fully chronicled — https://breakingdownfinance.com/finance-topics/alternative-investments/hedge-fund-database-biases/. Estimates of survivorship bias in hedge fund returns range from 2.6–3.7% annually to roughly 301 basis points/year depending on methodology, and instant-history bias (backfilling performance after joining a database) adds another ~167 bps/year — https://www.bogleheads.org/wiki/Survivorship_bias. Combined, selection + survivorship + instant-history bias can overstate a track record's true edge by 3–4.5% annually — https://www.bogleheads.org/wiki/Survivorship_bias. A parallel example from mutual funds: including failed funds drops average reported returns from 12% (survivors only) to 6% — a 50% cut — https://bookmap.com/blog/survivorship-bias-in-market-data-what-traders-need-to-know.
Conflicts of interest
managers who run both a hedge fund and a mutual fund earn ~16 cents per dollar of incentive fee from the hedge fund side versus a flat fee from the mutual fund side, and their mutual funds significantly underperform once they start running the hedge fund — a documented cross-subsidy — https://www.sciencedirect.com/science/article/abs/pii/S0304405X18300710. Managers can also inflate performance by cross-trading between portfolios they control — https://www.managementstudyguide.com/hedge-funds-and-conflict-of-interest.htm.
Persistence is weak even for the honest track records
in out-of-sample tests (1999-2014), only bottom-performing funds show persistence — top performers' outperformance does not reliably persist, and what persistence does exist is time-varying rather than stable — https://gitnux.org/hedge-fund-performance-statistics/. Across 5-, 10-, 30-, and 93-year horizons, secretive (non-transparent) funds show no evidence of outperforming transparent ones — undermining "my edge is proprietary, trust the number" claims — https://www.sciencedirect.com/science/article/abs/pii/S0378426621002442.
Regulatory disclosure raises the floor but is shrinking
Form PF (large private fund advisers) requires reporting of strategy, leverage, and liquidity; Form ADV covers fees and conflicts — https://financeworld.io/learn/sec-disclosure-requirements-what-hedge-funds-must-know/. As of 2026, the SEC and CFTC have proposed raising the Form PF reporting threshold from $150M to $1B AUM, which would remove disclosure coverage for most small funds — https://www.bloomberg.com/news/articles/2026-04-20/sec-cftc-propose-narrowing-hedge-fund-reporting-requirements.
Reading rule for this wiki
a practitioner number is only as strong as its disclosure. Audited filings, third-party custodian records, and full-portfolio transparency (not just headline return) move a claim toward credible; an unaudited "track record" with no disclosed drawdowns or costs is closer to the unverified regime below.
Unverified/anecdotal claims: forums, books, influencers
Anecdotal evidence is not subject to scholarly, scientific, or legal-standard rigor, and has little or no safeguard against fabrication or inaccuracy — https://en-academic.com/dic.nsf/enwiki/268255. It is the least certain form of evidence in the literature: useful for generating a hypothesis, never as proof of one — https://study.com/learn/lesson/anecdotal-evidence-examples.html. "A trader made money" carries no control for survivorship, look-ahead, or selection bias and is not evidence under any professional standard used elsewhere in this wiki — https://bookmap.com/blog/survivorship-bias-in-market-data-what-traders-need-to-know (implicit).
Grey literature — books, forums, unreviewed reports — usually skips peer review (though it may get informal internal review), tends to state conclusions without showing the process that produced them, and gives authors room to make claims the evidence doesn't support — https://libguides.tulane.edu/c.php?g=1513936&p=11324740. Even the more disciplined end of anecdotal evidence, the medical case report, is a mixed signal within its own field rather than a settled one: it "cannot be dismissed as having no weight," and in one retrospective sample 35 of 47 case-report anecdotes were later confirmed as clearly correct — but that same track record is precisely why a single case report still needs independent confirmation before it counts as proof — https://www.wikidoc.org/index.php/Anecdotal_evidence.
Reading rule for this wiki
a claim from a forum, book, or influencer with no published effect size, sample size, or significance test is UNVERIFIED by definition — there is no primary source to check. This wiki marks it as such rather than omitting it, because "widely believed but never measured" is itself useful information about where the evidence gap is. See UNVERIFIED (Not Refuted With Numbers): Head-and-Shoulders and Classic Chart Patterns and UNVERIFIED (Not Refuted With Numbers): Head-and-Shoulders and Classic Chart Patterns for worked examples of this status applied to specific techniques.
How to weight each regime when they disagree
Small disagreements between regimes are often explainable — a time lag (academics update slowly) or a scope difference (academics study broad markets, practitioners optimize a specific niche); a *large* divergence should raise skepticism of both sides, not just resolve in favor of one — https://oxford-review.com/evidence-based-practice-essential-guide/. The GRADE framework, built for exactly this kind of cross-source conflict, evaluates evidence quality, screens for conflicts of interest, and weighs findings against practical context rather than picking a single "winning" source — https://oxford-review.com/evidence-based-practice-essential-guide/. Financial conflicts of interest specifically should lower confidence in a claim independent of its stated numbers — https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3322153/.
Convergence is the strongest signal available
when peer-reviewed research and multiple independent, fully-disclosed practitioner track records agree, confidence should rise substantially — this is the pattern this wiki looks for before calling a technique "verified" per What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique. Practitioner evidence is strongest specifically when backed by audited filings and independent custodian records, not self-reported performance — https://www.tandfonline.com/doi/full/10.1080/1351847X.2021.1954966.
Practical checklist before trusting a practitioner claim against an academic one
- No out-of-sample test, no transaction-cost accounting, no multiple-testing correction → treat as a preliminary hypothesis, not a finding — https://backtrex.com/en/blog/hedge-fund-backtesting-quantitative-strategy-methods-2026. - Unaudited, self-reported track record with no disclosed drawdown → assume survivorship/selection bias is present until shown otherwise (see the 3-4.5%/year combined-bias estimate above). - Even peer review is not bias-free here: reviewers show documented prestige bias, judging work more favorably when it comes from established researchers or institutions — https://methods.cochrane.org/bias/reporting-biases. A peer-reviewed result from an unknown author deserves the same scrutiny as a well-known one, not less.
(UNVERIFIED, not independently confirmed against a peer-reviewed source: the specific claim that a strategy needs out-of-sample Sharpe ≥ 60-70% of in-sample Sharpe to be considered non-overfit — this figure appears only in a practitioner blog post, https://backtrex.com/en/blog/hedge-fund-backtesting-quantitative-strategy-methods-2026, and is flagged here exactly to demonstrate the grading this page argues for.)
Why this page matters for the wiki's objective
every reference page in this wiki attaches a proof regime to its numbers (see What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique). This page is the rulebook for what that label means and how much to trust it — without it, "peer-reviewed" and "a guy on a forum said" would carry the same weight, and the wiki's central claim — that it reports what the evidence actually says — would have no way to distinguish evidence from assertion.
Related
- What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique — the four-part evidence bar (effect size, proof regime, cost survival, out-of-sample persistence) that this page's proof-regime axis feeds directly into. - Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance — expands the Harvey-Liu-Zhu t>3.0 result cited above into the full multiple-comparisons problem behind the "factor zoo." - The Factor Zoo and Finance's Replication Crisis — why most published factors are false discoveries — the Hou-Xue-Zhang (65% fail) and Jensen (82.4% replicate) numbers used above as the headline evidence that peer review alone doesn't guarantee replicability. - UNVERIFIED (Not Refuted With Numbers): Head-and-Shoulders and Classic Chart Patterns — a concrete technique graded UNVERIFIED under the rule set out in this page's third section.
Verified against
30 claims checked against these sources · 5 refuted and removed
- nber.org/papers/w20592
- onlinelibrary.wiley.com/doi/10.1111/jofi.13249
- academic.oup.com/rfs/article-abstract/26/1/208/1592622
- bogleheads.org/wiki/Survivorship_bias
- pmc.ncbi.nlm.nih.gov/articles/PMC10699644
- sciencedirect.com/science/article/abs/pii/S0304405X18300710
- sciencedirect.com/science/article/abs/pii/S0378426621002442
- replicationnetwork.com/2017/06/14/hou-xue-zhang-replication-con…
- methods.cochrane.org/bias/reporting-biases
- libguides.tulane.edu/c.php
What links here
Source: Sinapsi — verified compositional memory, queryable by LLMs. Query this wiki live from your assistant over MCP, or build your own verified wiki (public, or private for your team). CC BY 4.0 — reuse with attribution to Sinapsi.