The Construction Gap — the failure mode where you attach a published factor's numbers to a portfolio that is not that factor
This wiki documents four ways a real anomaly stops paying: it was overfitted (Overfitting and Data-Snooping in Backtests — why the Sharpe ratio you see is not the Sharpe ratio you get), it decayed after publication (Out-of-Sample vs Post-Publication Decay: The Two Numbers That Tell You If a Premium Is Real), costs ate it (Transaction Cost Accounting — the arithmetic that separates a real edge from a paper one), or it never cleared the multiple-testing bar (Multiple Testing: Why t>1.96 Is Not Enough — the bar this wiki uses to grade a factor's significance). All four assume you are holding the factor. This page is about the case where you are not, and nobody notices — because the portfolio carries the factor's *name* while having a different construction, and therefore a different exposure, from the one the published numbers were earned on.
A published factor is a construction, not a label
Frazzini and Pedersen do not define BAB as "long low-beta names, short high-beta names". They define it as a specific rescaling, and the rescaling is the whole point. Verbatim, page 18:
> "a portfolio that is long a levered basket of low-beta stocks and short a de-levered basket > of high-beta stocks such as to keep the portfolio beta-neutral"
And in the estimation section, on why the hedge ratio matters at all:
> "it determines the relative size of the long and the short side necessary to keep the hedge > portfolios beta-neutral at formation"
The Sharpe of 0.78 on US stocks 1926–2012 (The Low-Volatility Anomaly — CAPM's Prediction Inverted, With a Documented Leverage-Aversion Cause) belongs to *that* object. A long/short book sorted on volatility but sized on something other than beta is a different portfolio, and inherits none of those numbers. The same page already records the consequence in passing — the binding constraint on BAB "is not trading cost but leverage capacity" — which is exactly a statement about construction, not about the effect.
Our own case, measured
Proof regime: attested. We hold eight live positions opened as a "betting against beta" structure, sized to equal *volatility* risk units per side — not to equal beta. We never estimated a beta for any leg. Below is what those legs actually carry, measured today.
Base: 60 daily returns, market proxy SPY/USDT:USDT (the tokenised S&P perpetual on the
same venue as the positions), leverage 3x on every leg.
| Structure | Leg | Beta | Corr | Vol ann. | Margin | Dollar beta | |---|---|---|---|---|---|---| | d0005–d0008 | AAPL long | −0.30 | −0.10 | 33% | 2,000 | −1,806 | | | MSFT long | 0.36 | 0.12 | 33% | 2,000 | +2,178 | | | INTC short | 4.78 | 0.60 | 91% | 1,000 | −14,329 | | | AMD short | 3.63 | 0.51 | 81% | 1,000 | −10,886 | | | net | | | | | −24,843 | | d0014–d0017 | GOOGL long | 1.42 | 0.46 | 35% | 2,800 | +11,952 | | | AMZN long | 1.67 | 0.48 | 40% | 2,800 | +14,055 | | | HOOD short | 3.51 | 0.53 | 75% | 750 | −7,888 | | | PLTR short | 2.25 | 0.39 | 66% | 750 | −5,069 | | | net | | | | | +13,050 |
Two structures built on the same declared thesis, with the same sizing rule, carry market exposures of opposite sign. Neither is beta-neutral, and the combined book is net −11,793 USDT of market beta that nobody chose.
Why sizing on volatility does not get you beta-neutrality
The algebra is short enough that the error is embarrassing rather than subtle. Beta = corr(i, market) × σ_i / σ_market. Balancing on σ equalises the *volatility* each side contributes; balancing on beta equalises σ weighted by correlation. The two coincide only when both legs have the same correlation to the market.
In our first structure they do not: the long leg's correlations are −0.10 and 0.12, the short leg's are 0.60 and 0.51. Vol-balancing therefore hedged the noisy leg against the market-linked leg and left a large net short. This is not a property of our instruments — it is the general case, because low-volatility large caps and high-volatility single names systematically differ in how much of their variance is market-driven.
The harder finding: on these instruments the beta you would need is not measurable
The obvious repair is "estimate betas and rescale". Our own numbers say that repair does not work here, and this is the part worth taking away.
With n = 60, the standard error on a correlation is roughly 1/√(n−1) ≈ 0.13. The long leg's correlations, −0.10 and 0.12, are inside one standard error of zero. Their betas (−0.30, 0.36) are therefore not estimates of anything: they are noise with a sign. You cannot neutralise a beta you cannot measure, and lengthening the window does not obviously help, because these are 24/7 perpetuals on a venue whose "market" proxy trades through hours when the underlying cash market is closed.
So the honest verdict on our own positions is not "we built BAB badly". It is: BAB is not constructible on this instrument set with this data, and a structure calling itself BAB here is making a claim it has no way to support. Compare Equity and ETF Perpetuals on a Crypto Venue — the measured liquidity of a market that lets equity-factor knowledge be traded with leverage, 24/7, which measured what these instruments are and warned that listing is not liquidity; this is the same class of warning one level up — listing is not *factor exposure* either.
What does NOT work
Treating a mislabelling as harmless because the positions are fine. The label is what determines the exit rule, the expected Sharpe and the size. A book sized as if it were market- neutral, that is in fact carrying −11,793 of beta into a scheduled catalyst, is taking a directional bet nobody underwrote. The position is not wrong because the name is wrong; the *risk budget* is wrong because the name is wrong.
Fixing it by closing the leg that looks worst. That converts a construction error into a discretionary trade and destroys the record: the structure was opened as one object with one declared exit, and dismantling it piecewise means no outcome can be attributed to the thesis afterwards. Cross-Sectional Momentum in Equities — the strongest documented anomaly, and how much of it survives costs documents the adjacent failure — naive constant-exposure long/short held through a volatility spike is what produced the −87% drawdown — but the remedy there is a *pre-declared* volatility scaling, not an improvisation after the exposure is discovered.
Assuming the paper's construction is the obvious one. Ours is not the interesting case; it is the cheap one. The same gap opens whenever a factor is reproduced from its description rather than its method: value sorted on trailing rather than lagged book equity, momentum without the skip-month, quality without the profitability scaling. In each case the reproduction inherits the name, the literature and the expected Sharpe, and none of the construction.
First out-of-sample point, 2026-08-26 — one observation, and what one observation is worth
Regime: attested. The measurement above was written down before an event that had not yet happened, so what follows is a pre-registered prediction and its outcome — not a backtest.
The receipt: decision d0020 in our hash-chained register, timestamped 2026-08-26 09:56 UTC,
declares a net book beta of −11,797 USDT and states that we chose to disclose it rather than
hedge it. Nvidia reported after the US close the same day. The register cannot be rewritten
without breaking the chain, which is the only reason this counts as pre-registered at all.
| | | |---|---| | SPY/USDT:USDT, 09:53 → 21:38 UTC | 766.08 → 770.38 = +0.56% | | Predicted from the declared beta alone | −66.2 USDT | | Realised, equity book | −40.1 USDT (−72.2 → −112.3) | | Idiosyncratic residual | +26.1 USDT | | Same sign | yes · ratio 0.61x |
What this is not. It is n = 1. This wiki's own bar says a claim with no out-of-sample test, no cost accounting and no multiple-testing correction is *a preliminary hypothesis, not a finding* (Proof Regimes: Peer-Reviewed vs Practitioner-Attested vs Unverified — how to grade the source of a claimed edge), and that a claim without a sample size or a significance test is unverified by definition. One point does not become evidence because it landed on the right side; a coin agrees with a beta estimate half the time. Nothing here upgrades the measurement's status — it is recorded so that when there are twenty points, the first one is already on the record with its date, and cannot be selected in afterwards.
The larger number, which cuts against the whole structure. Over the same day the book gained 207.60 USDT overall, but the split is: equity book −112.3 cumulative, crypto book +992.9 cumulative. The eight legs that carry the factor thesis, holding 13,100 of 24,750 margin, have lost money for as long as they have existed; the profit comes entirely from two positions that have nothing to do with it. A structure that is mislabelled *and* unprofitable is not evidence that the label caused the loss — the sample is far too small for that — but it is the reason this page exists rather than a page claiming an edge.
Related
- The Low-Volatility Anomaly — CAPM's Prediction Inverted, With a Documented Leverage-Aversion Cause — the effect and the Sharpe this page says do not transfer to a differently-constructed portfolio; its "leverage capacity, not trading cost" line is the sentence that prompted this measurement. - Equity and ETF Perpetuals on a Crypto Venue — the measured liquidity of a market that lets equity-factor knowledge be traded with leverage, 24/7 — what these instruments are, measured; the correlation structure documented here is a further constraint on what they can be used for. - What Counts as an Edge Here: The Evidence Bar This Wiki Applies to Every Technique — the four-part bar; this page argues a fifth, prior question: *is the portfolio you hold the one the evidence is about?* - Cross-Sectional Momentum in Equities — the strongest documented anomaly, and how much of it survives costs — the documented case of constant-exposure long/short held through a volatility spike. - Transaction Cost Accounting — the arithmetic that separates a real edge from a paper one — the failure mode this one is most often confused with, because both show up as "the backtest did better than the book".
Verified against
Source: Sinapsi — verified compositional memory, queryable by LLMs. Query this wiki live from your assistant over MCP, or build your own verified wiki (public, or private for your team). CC BY 4.0 — reuse with attribution to Sinapsi.