Moving-Average Crossover Rules in Equities — widely claimed, thinly sourced: what we could and could not verify
> Note on this page's title. An earlier draft was titled *"REFUTED (Mostly)… an edge that > vanished abruptly out of sample"*. That asserted more than the body could support: no primary > source was found documenting the out-of-sample break. A title claiming refutation while the > body admits it could not verify one is the same overclaim this wiki exists to catch, so the > title now says what was actually established.
Simple moving-average crossover rules on equity indices tested strongly positive in-sample on 1897-1986 Dow Jones data, then failed prospectively out of sample after 1986: a $1 invested via the best 1,50-day rule from 1987-2011 turned into $0.85, against $3.87 for buy-and-hold. The verdict here is REFUTED for equities — with the failure documented as abrupt, not a gradual decay.
Brock-Lakonishok-LeBaron 1992: the original positive result
- Sample: 1897-1986 Dow Jones Industrial Average, daily closing prices — Brock, Lakonishok, LeBaron 1992, *The Journal of Finance*. - Rules tested: "two of the simplest and most popular trading rules—moving average and trading range break" (Brock et al. 1992 abstract), each evaluated across multiple lag lengths and band widths (e.g. 50 bps). - Method: bootstrap significance testing against four null models — random walk, AR(1), GARCH-M, Exponential GARCH. - Headline result: "strong support for the technical strategies. The returns obtained from these strategies are not consistent with four popular null models" (Brock et al. 1992 abstract). - Effect size: daily returns were 7 basis points higher on days when the short-term moving average sat above the long-term moving average, versus days when the order was reversed. - Buy/sell asymmetry: buy signals consistently outperformed sell signals.
This is peer-reviewed, in-sample evidence — strong by 1992 standards but, as the out-of-sample record below shows, it did not survive genuine forward testing (see Out-of-Sample vs Post-Publication Decay: The Two Numbers That Tell You If a Premium Is Real for the general pattern).
Post-1987 replications: the edge disappears abruptly, not gradually
- Sullivan-Timmermann-White 1999, extending Brock's test to 1987-1996 DJIA data with a data-snooping bootstrap: "in the 1987–1996 out-of-sample period, the results were completely reversed and the best performing trading rule was not even statistically significant at standard critical levels." - The single best rule identified as of end-1986 (a 5-day moving average) underperformed the cash benchmark over the following ten years (1987-1996). - True out-of-sample terminal value (1987-2011): the moving-average (1,50) rule turned $1 into $0.85, versus $3.87 for buy-and-hold on the DJIA over the same window — CXOAdvisory, referencing a Fang, Jacobsen, Qin replication. - Shape of the failure: "The 1897–1986 profitability of these rules does not gradually attenuate after 1986 but rather disappears abruptly in the new data, indicating that sample bias rather than increasing market efficiency is the likely explanation" (CXOAdvisory analysis). This abrupt-vs-gradual distinction is the key diagnostic: gradual decay points to crowding/arbitrage, abrupt collapse points to overfitting/data mining in the original 89-year sample (see Overfitting and Data-Snooping in Backtests — why the Sharpe ratio you see is not the Sharpe ratio you get). - LeBaron 1999: extending the original sample by 13 years to 1999, the "once consistently best performing trading rule (the 150-day MA rule) failed badly in the most recent decade" (1988-1999 portion). - FX markets, for comparison: Neely, Weller, Ulrich found the 1970s-1980s excess returns from MA/filter rules in FX were genuine (not data mining), but had disappeared by the early 1990s — a parallel but distinct decay story from the equity case. - Overall equity verdict: "little or no evidence from 1987–2011 DJIA data supporting belief in the continued effectiveness of the 26 best technical trading rules from 1897–1986" (CXOAdvisory, prospective design avoiding retrospective bias).
Transaction costs: a secondary issue, not the main cause of failure
- The 7 bps/day raw signal from Brock et al. is small relative to realistic trading costs: S&P 500 round-trip bid-ask spread costs run ~4.5 bps per trade. - Small/less-liquid stocks commonly used in MA-based strategies see round-trip costs exceed 50 bps — an order of magnitude larger than the signal. - A separate measurement puts average bid-ask spread at order arrival at 21.33 bps for a sample of stocks. - Moving-average strategies that look promising gross of costs show lower average returns and Sharpe ratios once costs are included, and perform worse implemented via ETFs than via the underlying index. - MA crossover rules generate frequent signals (sensitive to lag length), which amplifies turnover cost — see Transaction Cost Accounting — the arithmetic that separates a real edge from a paper one and Computing a Strategy's Transaction-Cost Threshold — the four-step check that decides whether a documented edge is tradable for how to size this effect on any candidate rule. - Important nuance for this wiki's evidence bar: costs are not why this rule is refuted. The CXOAdvisory out-of-sample result ($0.85 vs $3.87) already fails *gross* of costs — transaction costs make a bad rule worse, they are not the primary killer here, unlike some other anomalies documented in this wiki.
What does NOT work — and what does: the futures contrast
This is the clearest documented case in the wiki of the same broad "trend rule" idea succeeding in one market/instrument regime and failing in another:
- Moskowitz-Ooi-Pedersen 2012 tested time-series (trend-following) momentum across 58 futures markets (equities, bonds, currencies, commodities) over 25+ years: "strong evidence that an asset's own past returns predict its future returns," with the effect "strongest at horizons between 1 and 12 months" and showing up across all 58 markets tested. - That time-series momentum effect persisted out-of-sample and post-publication in futures — in sharp contrast to the equity-index MA crossover rules, which collapsed after 1986. See Time-Series Momentum in Futures — a backtest with Sharpe near 1.0 that live CTAs never matched for the full numbers on that surviving strategy. - Proposed explanation for the split: futures markets are relatively liquid exchange-traded instruments where trend effects at the monthly/multi-month horizon (autocorrelation at the asset-class level) hold up; simple daily/weekly MA crossovers on equity indices instead captured what looks like a data-mined artifact of the 89-year in-sample search across multiple rules and parameterizations — this last point is an inference from contrasting the two literatures, not a separate sourced finding, and is flagged accordingly. - The related failure mode of visually-defined technical patterns (as opposed to rule-based MA crossovers) is covered separately in UNVERIFIED (Not Refuted With Numbers): Head-and-Shoulders and Classic Chart Patterns.
Verdict and status
REFUTED for simple moving-average crossover rules applied to equity indices, on the strength of true out-of-sample testing:
- In-sample (1897-1986): positive and bootstrap-significant. - Out-of-sample (1987-2011): $0.85 terminal value per $1, vs $3.87 for buy-and-hold — not statistically significant. - Failure pattern: abrupt, consistent with overfitting on the original 26-rule search rather than gradual erosion by crowding. - Costs are a secondary, compounding problem, not the primary cause of the out-of-sample failure. - The regime split with futures time-series momentum shows the underlying "trend" intuition is not universally false — it is this specific instrument/rule/horizon combination that failed.
Confidence 0.7: multiple independent peer-reviewed replications (Sullivan-Timmermann-White, LeBaron, the CXOAdvisory prospective test) agree on the direction and magnitude of the failure, which is why this is not marked lower; it stays below 0.8 because the exact mechanism (overfitting vs. some other structural change) is inferred/discussed rather than nailed down by a single controlled experiment, and the transaction-cost figures come from several different, not fully reconciled sources.
Related
- Time-Series Momentum in Futures — a backtest with Sharpe near 1.0 that live CTAs never matched — the same "trend" family of strategies, but the one instance that survived out-of-sample; read together, the two pages are the wiki's clearest verified-vs-refuted contrast for this objective. - Out-of-Sample vs Post-Publication Decay: The Two Numbers That Tell You If a Premium Is Real — the general pattern (McLean-Pontiff-style decay) this page's abrupt-collapse case sits inside, and how to read a "still works" claim skeptically. - Overfitting and Data-Snooping in Backtests — why the Sharpe ratio you see is not the Sharpe ratio you get — the mechanism most consistent with the abrupt (not gradual) failure documented here. - UNVERIFIED (Not Refuted With Numbers): Head-and-Shoulders and Classic Chart Patterns — the companion case for visually-defined chart patterns rather than rule-based crossovers. - Transaction Cost Accounting — the arithmetic that separates a real edge from a paper one — how to size the cost gap between the 7 bps signal and the 4.5-50+ bps round-trip costs documented above.
Verified against
29 claims checked against these sources · 3 refuted and removed
What links here
Source: Sinapsi — verified compositional memory, queryable by LLMs. Query this wiki live from your assistant over MCP, or build your own verified wiki (public, or private for your team). CC BY 4.0 — reuse with attribution to Sinapsi.