Running Campaign Experiments (Drafts and Experiments) — Proving a Change Works Before You Bet the Live Campaign's Budget on It

verified · provenanceused 0× by assistantshowto

An experiment splits traffic between an unchanged original campaign and a draft with your proposed change, running both simultaneously so you get statistical evidence of impact before committing budget to the change account-wide. Skipping this step for a direct campaign edit means no comparison arm and no evidence — you only ever see the after, never a simultaneous before, so you can't separate the change's effect from everything else that moved (seasonality, competitors, your own other edits) in the same period.

When to use Drafts & Experiments vs a direct edit

Use an experiment whenever you want to isolate the impact of a change without risking the whole campaign's performance (support.google.com/google-ads/answer/10682377). An experiment allocates a portion of the original campaign's traffic and budget and runs the draft simultaneously with the base campaign for a set duration (support.google.com/google-ads/answer/6318742) — that simultaneity is the whole point: it controls for anything happening in the market at the same time. A direct edit to the live campaign skips the comparison step entirely and so cannot provide statistical evidence of impact, whereas an experiment can measure results and show impact before you apply the change permanently (support.google.com/google-ads/answer/6318742). "Drafts" are for preparatory work only (staging the change); "experiments" are the mechanism that actually measures it (support.google.com/google-ads/answer/10682377).

Traffic split

Google recommends a 50% split between original and experiment campaign as giving "the best comparison" (support.google.com/google-ads/answer/6261395). Custom experiments support only cookie-based or search-based traffic assignment (support.google.com/google-ads/answer/6261395); traffic splits cannot be changed after setup — a different split percentage requires a brand-new experiment (support.google.com/google-ads/answer/13826584). For audience-list experiments using cookie-based splits, keep the list at 10,000+ users for accurate results (support.google.com/google-ads/answer/6261395).

Minimum runtime

start with 2–3 weeks to gather data, and extend if results are inconclusive (support.google.com/google-ads/answer/6318747). Google's own best-practice guidance is stronger: run for at least 4–6 weeks to reach statistical significance and cover full conversion cycles (support.google.com/google-ads/answer/13826584). Two exceptions with their own clocks:

- Performance Max experiments auto-discard the first 7 days of data to absorb ramp-up, so both arms are compared fairly only after that window (support.google.com/google-ads/answer/13826584). - Smart Bidding strategy tests (Smart Bidding Exploration) need 2–3 conversion cycles or 7 days — whichever comes first — for the baseline to stabilize, then 6+ weeks of running (support.google.com/google-ads/answer/16294686). See Bidding Strategies: Manual, Smart Bidding, and When Each Applies — matching the strategy to the conversion data you actually have for what "learning period" means for the bid strategy itself, separate from the experiment clock. - App Uplift experiments should run 30 days if possible, to maximize the chance of a conclusive result (support.google.com/google-ads/answer/14074599).

What Google's significance report tells you — and its limit

Google computes significance with Jackknife resampling over bucketed data — 20 buckets each for control and treatment — to estimate sample variance, then runs two-tailed significance testing at a 95% confidence interval (support.google.com/google-ads/answer/9232676); Jackknife is used because it gives "a high level of coverage" (support.google.com/google-ads/answer/9232676). The Experiments page then shows an estimated difference (e.g. "+10% more clicks") with a confidence interval range (e.g. "95% chance the true difference is +10% to +20%") — support.google.com/google-ads/answer/6318747. The default confidence interval shown is 80%, and it's adjustable (support.google.com/google-ads/answer/6318747). Results resolve into three buckets: "Control campaign winner," "Treatment campaign winner," or "No clear winner," depending on whether there's enough data (support.google.com/google-ads/answer/6318747). "Inconclusive" is a data-volume problem, not a verdict: the fix is to extend runtime and, where possible, pick a higher-volume campaign to run the test on (support.google.com/google-ads/answer/13826584).

Hard limit

Google's methodology does not let you aggregate several experiments after the fact to recompute combined statistics — advertisers don't have access to the underlying user-level data needed to rebuild the buckets and re-run the algorithm (support.google.com/google-ads/answer/9232676). Each experiment's result stands alone; you cannot pool three inconclusive experiments into one conclusive one.

What does NOT work

- Running several experiments at once. They can interfere with each other and skew the result — run them one after another instead, so each gets clean data (support.google.com/google-ads/answer/13826584). - Testing more than one variable per experiment. Keep every arm identical except a single variable (audience, bid strategy, keyword match type, etc.) — support.google.com/google-ads/answer/13826584. - Ending an experiment manually before its scheduled end date. A manually terminated experiment will not auto-apply, even if it looked like a clear winner (support.google.com/google-ads/answer/13826584). - Trying to change the traffic split mid-experiment. It's fixed at setup; changing the split requires starting a new experiment (support.google.com/google-ads/answer/13826584). - Editing the original campaign or the experiment while it's running. Both are technically permitted for budget adjustments, but any other change makes results harder to interpret (support.google.com/google-ads/answer/13826584). - Running the test over an unrepresentative period. Seasonal demand swings (or Smart Bidding Exploration windows that overlap them) bias the read — isolating the test from seasonality is the reason to run an experiment in the first place (support.google.com/google-ads/answer/16294686). - Testing on a low-traffic campaign. Higher-volume campaigns produce more reliable data; low-traffic campaigns are prone to inconclusive results regardless of runtime (support.google.com/google-ads/answer/13826584). This is the same signal-thinning problem covered for automated bidding generally in What Does NOT Work: Broad Match and Smart Bidding Without Guardrails — where automation burns budget instead of saving it.

Applying or discarding an experiment safely

Once an experiment ends, apply it two ways: "Update your original campaign" pushes the experiment's changes into the original campaign, or "Convert to a new campaign" spins the experiment into its own live campaign and pauses the original (support.google.com/google-ads/answer/7457118). Either way, performance data for both the original and the experiment is preserved after applying, so the historical record for both arms stays intact (support.google.com/google-ads/answer/7457118). Only experiments that reached their scheduled end date (or that already have sufficient data) are eligible to apply — a manually terminated experiment does not auto-apply, per the mistake above (support.google.com/google-ads/answer/13826584). To discard an experiment instead, simply change its end date so it stops running at the close of the day (support.google.com/google-ads/answer/13826584).

For this wiki's objective

experiments are the account-safe way to test the changes recommended elsewhere in this wiki — a new bid strategy, a Performance Max asset-group change, a match-type shift — without risking the live campaign's spend on an unverified guess. Before applying a PMax change permanently, see How to Launch Performance Max Without Cannibalizing Brand or Wasting Spend for the controls to check first; after applying any experiment, the change should show up in the account's change history, which is the first place to look during a later How to Audit an Existing Google Ads Account (Priority Order) — the sequence that stops you from trusting broken data.

Related

- Bidding Strategies: Manual, Smart Bidding, and When Each Applies — matching the strategy to the conversion data you actually have — an experiment's runtime and a Smart Bidding strategy's own learning period are two different clocks that both have to finish before you trust the numbers; conflating them is how advertisers misread "inconclusive" as "it doesn't work." - What Does NOT Work: Broad Match and Smart Bidding Without Guardrails — where automation burns budget instead of saving it — the same low-volume/thin-signal problem that makes a Smart Bidding strategy unreliable is what makes a low-traffic experiment inconclusive; both trace back to not having enough data for the algorithm (or the significance test) to work with. - How to Launch Performance Max Without Cannibalizing Brand or Wasting Spend — Performance Max experiments have their own 7-day ramp-up discard rule; run the experiment before rolling PMax changes out account-wide. - How to Audit an Existing Google Ads Account (Priority Order) — the sequence that stops you from trusting broken data — an applied experiment leaves a trace in change history, which is where a later audit should look for performance shifts that correlate with it.

Verified against

53 claims checked against these sources

Source: Sinapsi — verified compositional memory, queryable by LLMs. Query this wiki live from your assistant over MCP, or build your own verified wiki (public, or private for your team). CC BY 4.0 — reuse with attribution to Sinapsi.