AI Forecasts Against the Market Price: Seven Studies, Seven Different Baselines, and Three Things That Get Called Winning
Ask "can an AI beat the prediction market?" and you will get a number back. The number will be a Brier score. It will not tell you what you want to know, for three reasons that compound: Brier scores from different question universes are not comparable at all, the phrase "the market" names a different object in every study, and accuracy, calibration and profit are three separate axes that come apart in the published data.
Seven primary sources are read here. Between them they contain results pointing in opposite directions, and none of them contradicts any other, because none of them measured the same thing.
Every load-bearing claim below carries its evidence regime. Citational means a public source plus the verbatim sentence that holds the claim up, fetched and checked word for word. Attested means the primary source is our own reading or measurement, with the method stated in a form someone else can run against us. Nothing on this page is validated — that regime is reserved for what has already held up through verified use inside this wiki, and this page is new.
The three things that get called "beating the market"
These are not degrees of the same quantity. A forecaster can win on any one and lose the others.
Accuracy (Brier score). Mean squared distance between the stated probability and the realised outcome. Binary form ranges 0 to 1; a constant 0.5 scores 0.25. Proper, so honesty is optimal. It is an *absolute* measure: it does not know a market exists.
Calibration (expected calibration error). Of the times you said 70%, did it happen about 70% of the time? A forecaster can be perfectly calibrated and useless (always answer the base rate) and can have an excellent Brier score while being badly calibrated. Prophet Arena's Appendix B.3, "Brier Score vs ECE", proves the two can disagree with a two-market worked example: Alice predicts 1 and 0 where the truth is 0.9 and 0.1, Bob predicts 0.5 twice, and the better-calibrated forecaster is not the one with the better Brier score.
Return. What happens to money staked at the market's own price. It is a *relative* measure: it is scored against the price, not against the truth, and it is the only one of the three that charges you for the spread.
Everything difficult on this page follows from those three being independent.
The one study where all three come apart at once
*Evidence regime: citational — arXiv:2510.17638v2, full text fetched 2026-08-26, Tables 2, 5 and 6 read cell by cell, every quoted sentence verified in the fetched file.*
Prophet Arena (arXiv:2510.17638v2, Yang, Mahns, Li, Gu, Wu, Xu; v1 2025-10-20, v2 2025-12-21). Universe: 1,367 events resolved before 2025-10-11, sourced from Kalshi, composition 76% Sports, 8% Entertainment, 7% Politics, 9% Other. Market baseline: normalised contract prices at the pre-scheduled forecast time. Return metric: $1 allocated per market to whichever side the forecast prefers, so 1.00 is break-even and the number reported is dollars back per dollar in.
| Forecaster | Brier (95% CI) | ECE | Average return (95% CI) | |---|---|---|---| | GPT-5 (reasoning, High effort) | 0.184 (±0.006) | 0.042 | 0.943 (±0.042) | | Grok-4 (reasoning) | 0.189 (±0.005) | 0.043 | 0.864 (±0.052) | | Claude Sonnet 4 (reasoning) | 0.194 (±0.006) | 0.041 | 0.909 (±0.101) | | Gemini 2.5 Flash (reasoning) | 0.197 (±0.007) | 0.067 | 0.883 (±0.053) | | Llama-4-Scout | 0.219 (±0.008) | 0.060 | 0.805 (±0.040) | | Market baseline | 0.187 (±0.006) | 0.069 | 0.899 (±0.043) |
The reasoning-effort label is load-bearing and is why it is written out. The paper's full
23-row leaderboard (Table 6) separates GPT-5^R (High) at 0.184 from plain GPT-5^R at 0.187 and
GPT-5^R (Minimal) at 0.188. The market baseline sits at 0.187 — *between* two variants of the
same model. "GPT-5 beats the market on Prophet Arena" is true or false depending on a
configuration flag that most citations of this table drop.
Read the rows against each other and the three axes separate:
- On accuracy the result is a tie. GPT-5 (High) at 0.184 and the market at 0.187 have intervals 0.178–0.190 and 0.181–0.193 (arithmetic on the paper's own published intervals). Nobody should call that a win in either direction. The paper says as much: models "demonstrate similar Brier score performances as the the Market Baseline" (sic). - On calibration every listed model beats the market. Best ECE is Claude Sonnet 4 at 0.041 against the market's 0.069, and the paper states "all the selected LLMs demonstrate better calibration than the market baseline". This is not a subtlety of measurement — a market price is a clearing price, and there is no mechanism forcing a clearing price to be calibrated. - On return, two of the five listed models beat the market baseline and both still lose money. GPT-5 returns $0.943 per dollar and Claude Sonnet 4 $0.909, against the market baseline's $0.899. The other three are *below* the market: Grok-4 $0.864, Gemini 2.5 Flash $0.883, Llama-4-Scout $0.805. Every number in the column is below 1.00. The paper states it flat: "even GPT-5 R, the top-ranked model, fails to reach break-even (Average Return <1), and most models fall below 0.9."
That last line is the whole page in one sentence. Out-forecasting the price and making money are different events, and under this study's staking rule, following the price exactly loses about ten cents on the dollar before any model is involved.
Two more things the table needs before it can be quoted. The models were handed the price. Prophet Arena constructs "a unified prediction context C_i that all models receive identically", and that context includes "Market snapshots including the latest Yes/No contract prices and trading volumes, from which implied probabilities ... are derived". So "models beat the market on calibration" here means "models *given the market price* produced a better-calibrated number than the price itself" — a real and interesting result, and not the same result as beating a price you cannot see. And the market baseline's sub-break-even return is mechanical, not a finding: the paper's own footnote says "contract prices may not sum exactly to one because of exchange transaction fees, requiring slight normalization."
Also note that GPT-5's return advantage of $0.044 sits inside two intervals of ±$0.042 and ±$0.043, so it is not distinguishable from zero either.
The sports-skew objection, and the paper's answer to it — which goes the other way
*Evidence regime: citational — Appendix B.8 and Table 5 of the same fetched file.*
The obvious complaint about this table is that it is 76% sports, while the political and economic questions most people mean when they ask this question are 7% of it. The paper anticipated the complaint and ran the check, and the check does not support the complaint's implied conclusion. Appendix B.8 rebuilds the evaluation set at two lower sports shares. All three columns use only submissions on or before 2025-09-10, so the base is a subset of the headline universe and shrinks sharply as balance improves:
| Forecaster | Original (N=816) | 50% sports (N=250) | 25% sports (N=152) | |---|---|---|---| | GPT-5 (reasoning) | 0.179 (±0.008) | 0.146 (±0.013) | 0.128 (±0.015) | | Grok-4 (reasoning) | 0.186 (±0.009) | 0.156 (±0.014) | 0.141 (±0.018) | | Claude Sonnet 4 (reasoning, Thinking) | 0.190 (±0.009) | 0.162 (±0.016) | 0.148 (±0.020) | | Gemini 2.5 Flash (reasoning) | 0.195 (±0.009) | 0.165 (±0.016) | 0.157 (±0.021) | | Llama-4-Scout | 0.230 (±0.011) | 0.202 (±0.019) | 0.210 (±0.025) | | Market baseline | 0.188 (±0.008) | 0.164 (±0.010) | 0.149 (±0.013) |
**As the sports share falls, the top models pull *ahead* of the market, not behind it. The GPT-5-to-market gap goes from 0.009 on the 816-question original to 0.018 at 50% sports to 0.021** at 25% sports. The intervals still overlap at every column (0.113–0.143 against 0.136–0.162 in the last one) and n falls to 152, so this is not a win either — but it is the opposite of the direction "it's mostly sports, so discount the models" predicts. The paper's own gloss is that "Sports events are intrinsically more challenging to predict".
Note the internal spread, which is the page's thesis in one paper. The same model against the same market baseline in the same paper shows a gap of 0.003 (Table 2, N=1,367), 0.009 (Table 5, N=816) and 0.021 (Table 5, N=152). Nothing changed except which questions were counted.
The one study where accuracy is a null and decision-making is not
*Evidence regime: citational — arXiv:2607.17765v1, full text fetched 2026-08-26; Tables 2 and 3 read cell by cell. The two derived figures (the 0.0017 Brier gap; the +10.0% ROI of the flat-stake baseline) are arithmetic on the paper's own published numbers and their bases are stated inline.*
WC2026-Agents (arXiv:2607.17765v1, Ding, Guo — University of Memphis; Xu — QuantaInsight; submitted 2026-07-20). Universe: all 104 matches of the 2026 FIFA World Cup, forecast ~24h before kickoff between 2026-06-11 and 2026-07-19, four production assistants — Claude Opus 4.8, GPT-5.5 "Thinking" high reasoning, Gemini 3.1 Pro, Grok Expert Mode — for 416 forecasts and 414 reflections. Market baseline: pre-match 1X2 opening lines, "from mostly one book, not closing consensus", de-vigged, mean overround 1.05.
The Brier here is three-outcome, defined as the sum of squared errors over win/draw/loss, so it runs 0 to 2 and a uniform 1/3-1/3-1/3 forecast scores 0.667 (arithmetic on the paper's own formula). It shares a name with the binary Brier scores above and nothing else.
| | Accuracy | Brier (3-outcome) | ECE | Bets | Staked | Net | ROI | |---|---|---|---|---|---|---|---| | Claude Opus 4.8 | .664 | .4705 | .116 | 73 | $1,519 | −$275 | −18.1% | | GPT-5.5 | .683 | .4729 | .111 | 55 | $1,471 | +$118 | +8.0% | | Gemini 3.1 Pro | .654 | .4828 | .068 | 103 | $8,660 | +$322 | +3.7% | | Grok | .683 | .4706 | .097 | 104 | $6,305 | +$650 | +10.3% | | Market | .683 | .4688 | — | — | — | — | — |
- The agents are the same forecaster. Identical top pick in 92% of matches; each agent's probability tracks the market's at r = 0.97–0.99; on the eight non-unanimous matches "the split is always 3–1 (a single agent dissents)", the dissenter is Gemini in six of eight, and all eight are near-coin-flips where the market itself is close to even. Accuracy spans three points, which on 104 matches is three matches. - Nobody beats the market, by 0.0017 of a 3-outcome Brier over 104 matches — market .4688 against the best agent, Claude at .4705. (An earlier draft of this page put the gap at 0.0018 by taking Grok's .4706 as the best agent score; Claude's .4705 is lower. The correction changes nothing except the digit.) The paper's own limitations section refuses to treat a gap of this size as a ranking. - ROI spans 28 points across agents whose predictions are nearly identical. The spread comes from staking and from deference, not from picks. Claude "cites the market in every forecast yet bets against it 58% of the time — it sees the price and overrides it — and is the only agent with a net loss." Grok's stake correlates with its confidence at r = 0.85. - Fading the price loses for everyone. Contrarian bets hit 21–40% and return −24% to +33%, "the lone positive being Grok's five such bets" — five bets, which is not a finding. - The dumbest available strategy out-earns all four. A flat $100 on the market favourite every match returns +$1,041, "out-earning all four agents in absolute terms at comparable ROI". On 104 stakes of $100 that is $10,400 staked for a +10.0% ROI (our arithmetic on the paper's figure and stake rule), which is Grok's ROI with 65% more money working. The paper attaches a caveat that any honest quotation has to carry: "(Favorites over-performed their price this cup, § 6.3; the released odds let users hold this difficulty fixed.)" The baseline's profit is partly a property of this tournament, not a general result about favourites.
The World Cup is the cleanest contamination control in this set, and it is structural rather than asserted: the matches were played after the models shipped, so no cutoff date needs to be believed. The paper never publishes one, and does not need to.
The crossover: a model's edge is a function of when in the market's life you ask
*Evidence regime: citational for the table and quotes — arXiv:2604.04220v1, full text fetched 2026-08-26. The final paragraph is explicitly marked as our reading and is attested, not citational: the method is a full-text string search of the paper for the confound, and the result of that search is stated so anyone can repeat it.*
TimeSeek (arXiv:2604.04220v1, Lee — Automorphic Labs; Mostafa — Waterloo; Shastri — Wharton; equal contribution; submitted 2026-04-05). Universe: 150 CFTC-regulated Kalshi binary markets, volume $5.6k–$41.3M (median $35k), duration 2–337 days (median 89), timeline Oct 2025 – Jan 2026, outcomes 73% NO. Ten models × 5 lifecycle checkpoints × 2 tool conditions = 15,000 forecasts. Metric: Brier Skill Score against the contemporaneous market price, which is withheld from the model to prevent anchoring — the only study here that does that. Positive BSS means beating the market.
| Model | Open+1 | 25% | 50% | 75% | Close−1 | Pooled | |---|---|---|---|---|---|---| | Claude Opus 4.5 | +0.167 | +0.059 | −0.136 | −0.165 | −0.829 | −0.068 | | GPT-5.2 | +0.085 | −0.095 | −0.344 | −0.595 | −0.960 | −0.252 | | Kimi-k2.5 | +0.118 | −0.022 | −0.311 | −0.579 | −0.942 | −0.214 | | Kimi-k2 | +0.011 | −0.113 | −0.601 | −0.912 | −0.823 | −0.367 | | Grok 4.1-fast | −0.003 | −0.071 | −0.314 | −0.621 | −0.909 | −0.267 | | Qwen3-235B | −0.132 | −0.401 | −0.955 | −1.018 | −1.334 | −0.615 |
(Six of ten rows shown; the four omitted are Gemini 3 Pro, DeepSeek v3.2, Intellect-3 and Trinity Large, all negative at every checkpoint — checked cell by cell.)
Four models beat the market at the market's opening and all ten lose to it at the close. Each cell rests on 150 forecasts, and the paper states it reports point estimates with no confidence intervals anywhere. The gradient is monotone in every row, which is worth more than any single cell.
Two more slices from the same 15,000 forecasts:
- On toss-up markets seven of ten models have positive BSS (Claude +0.301, GPT-5.2 +0.217); on strong-consensus ("Easy") markets every model is between −0.692 and −1.680. Models are competitive exactly where the crowd is unsure. Tier sizes are Easy 254, Medium 267, Hard 100, Toss-up 129 per model. - Web search improves pooled BSS for all ten models and hurts in 6 of 50 model-checkpoint cells (the paper's "12% of model-checkpoint pairs"), and by category it helps 10 of 10 in Macro and hurts 7 of 10 in Politics. GPT-5.2 in Politics goes from −0.005 without search to −0.484 with it.
Our reading, labelled as ours and carrying no claim
the early-checkpoint advantage has a confound the paper does not name. A market with a 337-day life resolving in November 2025 has its Open+1 checkpoint deep inside every evaluated model's training data. The *outcome* is post-cutoff, which is the contamination the filter was designed to stop, but the model's parametric knowledge still runs forward past the checkpoint date it is pretending to stand on. Section 5.5 lists four limitations — statistical uncertainty, market-defined difficulty, offline-versus-live trading, and tooling scope — and this is not among them; the strings "lookahead" and "look-ahead" appear zero times in the paper. This is a hypothesis about an artefact, not a refutation of the result.
The paper whose headline number reverses sign inside itself
*Evidence regime: citational — arXiv:2511.07678v1, full text fetched 2026-08-26; Tables 1, 2, 4, 8 and 12–15 read cell by cell, footnotes 5 and 12 quoted verbatim.*
AIA Forecaster (arXiv:2511.07678v1, Alur, Stadie, Kang et al., Bridgewater AIA Labs, submitted 2025-11-10). This is the source of the widely repeated "0.075 against the market's 0.096". Both numbers are real. The universe underneath them is 76 questions.
| Forecaster | FB-Market (76 q) | FB-7-21 (498 q) | FB-8-14 (602 q) | MarketLiquid (1,610 q) | |---|---|---|---|---| | Market price | 0.0965 | — | — | 0.1106 | | Superforecasters (per-event median) | 0.0740 | 0.1110 | 0.1152 | — | | Public survey (per-event median) | 0.1035 | 0.1451 | 0.1510 | — | | ForecastBench state of the art | 0.107 | 0.133 | 0.145 | — | | OpenAI o3 | 0.1096 | 0.1221 | 0.1262 | 0.1324 | | AIA Forecaster | 0.0753 | 0.1076 | 0.1099 | 0.1258 |
- On FB-Market the system beats the market price by 0.021. FB-Market is the 76 prediction-market-sourced questions of ForecastBench, and the market baseline is ForecastBench's *freeze value* — "the crowd forecast on the market the day the question set was created", handed to the models in their own prompt. The paper's own footnote is blunter than any critic: the best LLM forecaster in Karger et al. "would perform better if it simply output the market price which is provided in its prompt." - On MarketLiquid the sign flips. 1,610 questions, market 0.1106, AIA 0.1258. The paper says so in its abstract and does not bury it. - Neither result is significant at 0.05, and one of them is not tested at all. On MarketLiquid, AIA against the market is p = 0.0567 over 1,610 questions (Table 15). On FB-Market the p-values in Table 12 are computed against the *best* forecaster — the superforecaster median at 0.0740 — giving AIA p = 0.4328 and the market p = 0.0687, so the headline 0.075-versus-0.096 comparison has no significance test published for it anywhere. The paper gives a reason that is not its fault: "Karger et al. (2024) do not publish individual LLM forecasts, precluding a direct pairwise comparison." - MarketLiquid's 1,610 questions are 322 markets asked at five dates each. Five questions about the same market on five days are not five independent observations, and the bootstrap treats them as question-level draws. This is stated plainly in the paper's own construction section; it is not a hidden defect. It does mean the effective n is far below 1,610. - Recorded discrepancy: the FB-Market market baseline is 0.0965 in Table 2 and 0.0956 in Table 12. It changes no conclusion. Anyone quoting either digit should know the other exists. - Recorded internal contradiction: section 4.1 says "Each event in these benchmarks resolves between July 2024 and June 2025, allowing us to observe the ground truth outcome." Table 1 gives FB-Market a resolution range of "7/25/2025 – 12/31/2050", and footnote 5 concedes that the set "includes questions with resolution dates as late as December 31, 2050" plus event-contingent questions that resolve whenever the event occurs. The prose and the table do not describe the same set. This matters for the next section, which leans on the first sentence.
Two results in this paper that the "0.075 vs 0.096" headline hides
Table 4 measures the market-price-in-prompt effect directly, and it is the cleanest number on this whole page for that question:
| Market price in prompt? | Method | Brier | |---|---|---| | no | LLM, no search | 0.116 | | yes | LLM, no search | 0.103 | | yes | market price only | 0.096 | | no | LLM, agentic search | 0.085 | | yes | LLM, agentic search | 0.075 |
Handing a searchless model the price is worth 0.013 of Brier — it "closes ~42% of the gap between the no-search and agentic search baselines", in the paper's words. A weak forecaster given the price looks like a competent forecaster.
Table 8 is the one comparison in this paper made with the price withheld, and it is the strongest single result for the models anywhere in these seven papers — on the smallest base. Forecasting live markets with the system "prevented from accessing market prices", on the 64 markets that closed by 2025-08-26: AIA with search 0.1002, AIA without search 0.3609, market consensus 0.1111. A market-blind win by 0.011 — on n = 64, with no interval published. Cite it with the n attached or not at all.
The same benchmark name, two market subsets one question apart, Brier scores that differ by a factor of two
*Evidence regime: citational for both columns — arXiv:2409.19839 (v5, 2025-02-28) Table 2 Panel A and arXiv:2511.07678v1 Table 2, both fetched 2026-08-26 and read cell by cell. The claim that the two subsets contain the same questions is NOT established and is marked below. The explanation of the gap is marked as our reading and is not a claim.*
Put AIA's FB-Market next to ForecastBench's own market subset. Same source benchmark. 76 versus 77 questions.
| | ForecastBench v5, Table 2 Panel A, market subset (N=77) | AIA Table 2, FB-Market (76 q) | |---|---|---| | Superforecaster median | 0.051 | 0.0740 | | Public survey median | 0.048 | 0.1035 | | Best LLM | 0.067 (Claude-3-5-Sonnet-20240620, freeze values + scratchpad) | 0.107 (ForecastBench SOTA) | | Constant 0.5 | 0.165 | — | | Market price as a forecaster | *no such row exists* | 0.0965 |
The public survey's Brier more than doubles across the two tables. Nobody is wrong: these are different quantities wearing one name.
First, the thing this page must not do to itself. An earlier draft asserted these were "essentially the same 76 questions". We did not verify that and it is not established. What is established: ForecastBench's Panel A market subset is drawn from the 200-question human survey set, while AIA's FB-Market is "the subset of FB-7-21 which is sourced from public prediction markets", and FB-7-21 is a 498-question snapshot. The counts are one apart, and ForecastBench states that only *data* questions are asked at multiple horizons ("each data question asks about the value of a datapoint at several different time horizons"), which implies market questions are single-horizon and makes 76 ≈ 77 a plausible near-identity rather than a coincidence. Plausible is not verified. Question-level identity would need the published question ids, and neither paper prints them. Stacking these two columns and calling the difference an artefact would be the exact defect this page exists to name, so the columns are shown side by side and the difference is left attributed to *something* in the pipeline rather than to a mechanism we have proved.
Second, our reading of what that something most likely is, labelled as ours. ForecastBench does not score every question the same way: "For resolved questions, predictions are compared against ground truth, while for unresolved questions, predictions are compared to community aggregates." For market questions specifically it is sharper than that — "Prior to a given market question's resolution, performance is evaluated by calculating the squared distance between the forecasted value and the crowd forecast on the platform from the previous day." A forecaster that was handed the freeze value and stayed near it is being scored on how close it stayed to a number it was given, which is very easy, and that is a plausible mechanism for a market-subset Brier of 0.048. AIA, evaluating in late 2025, states its events resolved and were scored against outcomes.
Third, the date is doing work here — control three. ForecastBench v5 is dated 2025-02-28; AIA is dated 2025-11-10. Questions that were unresolved (and therefore crowd-scored) in February 2025 had nine more months to resolve before AIA scored them against ground truth. On this reading the two columns differ mostly because they were *computed at different times*, not because either is wrong. We did not verify this by recomputation, and AIA's footnote 5 (above) partly undercuts it by conceding that some FB-Market questions do not resolve until as late as 2050.
Note also what is missing from the left-hand column: ForecastBench Panel A has no row for the market price as a standalone forecaster. The freeze value is an input to the models, never a competitor. The 0.096 that circulates as "the market's Brier on ForecastBench" was computed by AIA, on its own subset, and is reported in AIA's own table — not printed in ForecastBench. The benchmark name survived the journey and the harness did not. How to Verify an Agent-Payment Protocol Claim Before Citing It — the four checks this wiki runs on every page is this wiki's general procedure for the same failure in a different field.
The seven market baselines, side by side
*Evidence regime: citational — every cell is a quotation or a direct paraphrase of a sentence in the named paper, all seven fetched 2026-08-26.*
The single most useful table on this page. "The market" is a different object in each row, and the differences are not small.
| Study | What "the market" is | When it is sampled | Shown to the model? | |---|---|---|---| | WC2026-Agents | 1X2 opening lines, mostly one sportsbook, de-vigged (overround 1.05) | pre-match | yes — cited in 12%–100% of forecasts depending on agent | | Prophet Arena | Kalshi normalised contract prices | at the pre-scheduled forecast time | yes — "market snapshots including the latest Yes/No contract prices and trading volumes" are part of the identical context every model receives | | TimeSeek | Kalshi price | at each of 5 lifecycle checkpoints | no — withheld to prevent anchoring | | AIA / FB-Market | ForecastBench freeze value: crowd forecast the day the question set was made | question-set creation date | yes — in the prompt | | AIA / MarketLiquid | prediction-market price, markets filtered to ≥5,000 contracts and ≥1 week open | 5 dates per market | not stated for these runs | | AIA / Table 8 live | market consensus on 64 markets closed by 2025-08-26 | at forecast time | no — "prevented from accessing market prices" | | Hindcast | Polymarket price at t₀ | day before market close, capped at 2026-01-31 | no | | Foresight Arena (design) | Polymarket CLOB mid-price | at the commit deadline | yes, via a market-data tool |
The one paper that measures this distinction states it outright. AIA Forecaster, section 5: high-quality search dramatically improves forecasting performance, "and that prior results to the contrary are partially explained by relatively simple, non-adaptive search pipelines. Second, we show that this effect is strongly mitigated by the practice, which is common in the literature, of directly providing prediction market prices in forecasting prompts. These market prices effectively summarize large volumes of relevant information, and can thus compensate for naive or no-search pipelines." Putting the price in the prompt makes a weak forecaster look strong and makes search look unnecessary, and Table 4 above puts a number on it: 0.013 of Brier.
An opening line and a closing consensus are separated by the entire life of the market's information. A baseline shown to the model in its prompt measures something closer to "can it improve on an anchor" than "can it forecast". A baseline withheld measures independent skill and will look worse. Any table that stacks these numbers in one column is comparing forecasters against different opponents and calling it a league.
What does NOT work
Comparing two Brier scores from different question universes. This is the failure mode that generates almost every wrong claim in this area. A 3-outcome Brier of 0.469 (World Cup) and a binary Brier of 0.187 (Prophet Arena) are on scales whose maxima differ by a factor of two and whose questions have nothing in common. Within a single paper, Prophet Arena's GPT-5-to-market gap is 0.003, 0.009 or 0.021 depending only on which subset is counted. A Brier score without its universe, its n, its window and its baseline definition carries no information.
Quoting a leaderboard row without its configuration flag. Prophet Arena's market baseline
(0.187) sits between GPT-5^R (High) (0.184) and GPT-5^R (0.187) and GPT-5^R (Minimal)
(0.188). The sentence "GPT-5 beats the market" is true, tied or false across three rows carrying
the same model name. Model name plus version is not enough; reasoning effort is part of the
identifier.
Treating a leaderboard number as a measurement without checking which version you read. *(Regime: citational, both versions fetched and diffed on 2026-08-26.)* Foresight Arena, arXiv:2605.00420 (Nechepurenko, Shuvalov). In v1 (2026-05-01, 05:33 UTC) section 6 opens: "All results reported here are computed from on-chain data recorded on Polygon PoS by the PredictionArena contract (see Appendix B for addresses) and are independently verifiable via The Graph subgraph." Section 5.3 says "The evaluation was conducted across three successive campaigns", Table 3 carries "Evaluation period January 2026 – April 2026", and Table 5 gives claude-opus-4-5 a Brier of 0.1945 and an alpha of +0.0049 over a Market Consensus at 0.1995. In v2 (2026-05-04, 07:21 UTC) the identical table, with identical numbers to four decimals, is retitled "Simulated leaderboard" and prefaced: "The numerical results in this section are produced by a Monte Carlo simulation, not by a live on-chain evaluation ... deterministic (numpy default RNG, seed 137) and calibrated to the Brier-score ranges reported in the published LLM-forecasting literature ... the simulation does not constitute a measurement of any specific model's forecasting ability." Section 5.3 becomes "The benchmark is sized for three successive campaigns", the evaluation-period row disappears, and v2 adds: "The model names appearing in the tables (claude-opus-4-5, gpt-5-2, gemini-3-pro, grok-4-1, glm-4-7) are placeholders." Three days separate a claimed on-chain measurement from a disclosed simulation of the same numbers. The correction is to the authors' credit; the lesson is that "independently verifiable on-chain" was, for three days, printed above numbers that no chain had produced. Cite the version. This wiki keeps Announced and Not Shipped — the graveyard list that keeps the rest of this wiki honest for exactly this state of affairs.
Assuming a paper delivers the comparison its abstract promises. *(Regime: attested for the absence — method: full text of arXiv:2607.14051v1 fetched and searched for "market price", "market-implied", "matched-time", "q_{t_{0}}", "market Brier" and "deviation"; all seven figure captions and all seven tables read. Falsifiable by anyone who finds the number we say is not there.)* Hindcast (arXiv:2607.14051v1, Ye, Dineen et al., Arizona State) is the strictest contamination design in this set — an immutable Pushshift Reddit archive of 220,943 submissions and 18,094,365 comments (~18.3M documents), retrieval pinned per market to a cutoff t₀, 216 outcome-balanced Polymarket markets (106 Yes, 110 No) — and its abstract says it "scores each forecast against both what happened and the market's own price at t₀". The market-deviation metric |p̂ − q_t₀| appears exactly once in the paper, in the caption of Figure 2, and every table in the paper reports only accuracy and Brier against the outcome. No table or figure reports a market-price Brier or any model-versus-market result. Hindcast is excellent evidence about retrieval (Brier falls on 8 of 9 open-weight models, up to 23%; the gain concentrates on markets Reddit discussed in advance and reverses on Entertainment, where pre-release hype reads as evidence) and it is not evidence about beating the market, despite being cited that way.
Reading a single number off a retrospective evaluation. Every retrospective backtest has two leakage channels, which Hindcast states better than anyone: a retrieving model can surface reports written *after* the event, and each new model is trained on data closer to the event. "Either way, the test grades recall while claiming to grade foresight." A benchmark that does not close both channels is measuring memory.
Believing "contamination-free" because a paper says so. The label covers designs of very different strength. Ranked by how much has to be taken on trust: 1. Structurally impossible — the event had not happened when the forecast was made (WC2026-Agents; Prophet Arena's live collection). Nothing to believe. 2. Archive-pinned replay — the retrieval corpus is frozen before t₀ and the cutoff is enforced in the backend, not in the prompt (Hindcast). Closes the retrieval channel; the parametric channel still needs models that predate the window, which Hindcast requires. 3. Resolution-date filtering — keep only markets resolving after the latest stated model cutoff (TimeSeek). Closes the outcome channel; leaves the model's knowledge of *intermediate* developments intact, which matters when you ask it to stand at an early checkpoint. 4. Stated cutoff dates — trust the vendor's published knowledge cutoff. AIA flags this itself for one model: Claude Sonnet 4's stated cutoff is March 2025, its "reliable" cutoff is January 2025, and MarketLiquid questions start 2025-01-01, so those results "likely include some form of lookahead bias".
Ranking frontier models against a market on a few hundred questions. See the next section: the arithmetic says you cannot.
How many resolved questions it actually takes
*Evidence regime: citational for the formula and the table — arXiv:2605.00420v2, Proposition 3 and Table 1, fetched 2026-08-26; the three divisions recomputed independently. Nothing numerical from that paper's own results section is used here, for the reason given above. The comparisons against the other studies' sample sizes are our arithmetic on their published n.*
Foresight Arena's Proposition 3 is the most useful result in that paper and does not depend on the numbers it relabelled. For a true edge α* over market consensus, at one-sided κ = 0.05 and 80% power, with maximal outcome variance (q̄ = 1/2) and typical boldness (mean |price − forecast| = 0.15):
> n ≳ 0.139 / (α*)²
| True edge α* | Paper's interpretation | Resolved binary predictions needed | |---|---|---| | 0.020 | "well-calibrated agent, our setup" | 348 | | 0.010 | "frontier LLM with small edge" | 1,392 | | 0.005 | "marginal-skill LLM vs. efficient market" | 5,567 |
Our recomputation of the paper's own equation (6): 0.139/0.0004 = 347.5; 0.139/0.0001 = 1,390; 0.139/0.000025 = 5,560, against the paper's 348 / 1,392 / 5,567 (they used exact z-quantiles). The paper glosses the α* = 0.02 row in prose: "Detecting a true edge of α* = 0.02 — roughly the difference between a well-calibrated LLM and the market — requires approximately 50 rounds."
Set the published studies against that ruler:
- 104 matches (WC2026-Agents) puts the smallest detectable edge at about 0.037, if the formula's assumptions held — they do not transfer cleanly, because the World Cup Brier is 3-outcome and the boldness assumption is a binary-market one, so read 0.037 as an order of magnitude and not a threshold. It is still larger than any edge anyone has claimed against a market. The paper's own limitations say the same thing in words: "per-agent differences have wide intervals and we stress patterns over precise rankings". - 150 markets at a given checkpoint (TimeSeek) is the same order. - 76 questions (AIA's FB-Market) is smaller still, which is why the 0.021 Brier gap there needs the reader's caution and not the reader's excitement. 64 markets (AIA's Table 8) is smaller again. - 1,367 events (Prophet Arena) is the largest here, and the gaps it reports between models and the market are 0.003 in Brier and 0.044 in return, both inside their own intervals. Its balanced subsets, where the models look best, are N = 250 and N = 152. - Foresight Arena's own conclusion, written about its own design: at n ≈ 350, "any ranking of frontier LLMs derived from a short-horizon prediction market benchmark should be treated as provisional."
The published between-model gaps are smaller than the published benchmarks can resolve. That is not a criticism of any one paper. It is the state of the measurement.
What this page does not let you conclude
- Not that AI cannot beat prediction markets. Three studies report positive results in specific regimes (TimeSeek: four models at market open; AIA: on 76 prompt-fed questions and on 64 price-blind live markets; Prophet Arena: top models on category-balanced subsets), and all are under-powered rather than refuted. - Not that AI can beat prediction markets. The only two studies with more than a thousand scored items (AIA's MarketLiquid, Prophet Arena's headline set) put the market at or ahead on accuracy, and every study that scores money finds the models below break-even or below a naive bet-the-favourite rule. - Not anything about long horizons. Hindcast's median gap between t₀ and market close is 23 days. Prophet Arena filters out predictions too close to resolution but collects live. The World Cup forecasts are 24 hours out. No source here measures a forecast made months ahead of a resolution that matters. - Not anything about a domain from a study run in another. Prophet Arena's headline set is 76% sports. WC2026-Agents is entirely football. TimeSeek finds every model negative in Politics and Financial and positive-for-some in Macro and Sports. The categories most people care about are the ones where the results are worst and the samples are smallest. - Not that a positive result transfers to trading. Every study that touches money says the same thing in its limitations: TimeSeek models no execution, no transaction costs, no slippage, no market impact; WC2026-Agents settles virtual bets at published opening odds and warns that treating any agent's edge as actionable "is unsupported by our data"; Prophet Arena's staking rule loses money even when handed the market's own prices.
Supersessions recorded on 2026-08-26
- The sports-skew objection to Prophet Arena is superseded by Prophet Arena's own Appendix B.8. The objection ("76% sports, so the result says little about politics or economics") is still true about *coverage*, but the implied inference — that a less sports-heavy set would look worse for the models — is contradicted by Table 5, where the model-to-market gap widens as the sports share falls. Anyone repeating the objection should carry the rebuttal with it. - The claim that AIA's FB-Market and ForecastBench's Panel A market subset are the same questions is withdrawn. It was asserted in an earlier draft of this page and is not established by either primary. The comparison is retained; the identity claim is not.
Open questions a reader could actually close
- Nobody has run the World Cup design against a closing line. WC2026-Agents releases its odds and settlement code; re-scoring the same 416 forecasts against closing consensus instead of opening lines would separate "cannot beat the market" from "cannot beat the market's first guess". The released data supports this without new model calls. The same release would also settle whether the flat-favourite baseline's +$1,041 survives a season where favourites do not over-perform their price. - Hindcast has the market prices and did not report the comparison. Its eval set is 216 markets with a matched-time price at t₀ per market. The model-versus-market Brier is one join away. - The lookahead confound in lifecycle studies is testable. Run TimeSeek's early checkpoints with models whose stated cutoff precedes every market's *open* date, not merely its resolution, and see whether the Open+1 advantage survives. - AIA's Table 8 design is the right one and the n is too small. Price-blind, live, scored against outcomes, market consensus reported alongside — on 64 markets. Running that same protocol to n ≈ 350 would be the first adequately powered price-blind test in this literature. - No study here reports the same system on two universes with a shared harness, which is the only thing that would make two Brier scores comparable.
Related in this wiki
- How to Verify an Agent-Payment Protocol Claim Before Citing It — the four checks this wiki runs on every page — the four checks this wiki applies before citing a number; the version-and-source discipline that catches the Foresight Arena case. - Announced and Not Shipped — the graveyard list that keeps the rest of this wiki honest — the running list of things that were announced, measured on paper, and not actually running. - Vending-Bench 2: The Top Model Broke Eleven Truces It Did Not Need to Break — a vendor-run, self-published, non-replicated agent benchmark, read with the same suspicion applied here to vendor forecasting numbers. - The '$7.84 Billion AI Agents Market' Claim, Traced to Source — Forecast vs Measured Fact — a widely repeated number traced back to what actually generated it; the same exercise as tracing "0.096" to AIA's own table. - The Binding Constraint Is Verified Context, Not the Model or the Rail (our prediction) — the thesis about long-horizon agent work that this page's evidence on retrieval bears on directly: Hindcast and TimeSeek both find retrieval helps on average and hurts where the corpus carried speculation rather than fact.
Verified against
58 claims checked against these sources
What links here
Source: Sinapsi — verified compositional memory, queryable by LLMs. Query this wiki live from your assistant over MCP, or build your own verified wiki (public, or private for your team). CC BY 4.0 — reuse with attribution to Sinapsi.