Incrementality Testing: Geo Tests & Holdouts, No Data Team
Incrementality testing is the quickest way to answer the question you already have in the back of your mind: are your ads bringing in new sales, or are they mostly getting credit for sales you would have earned anyway? With tracking getting fuzzier thanks to privacy updates and platform changes, you need a way to measure cause and effect without relying on perfect attribution. The good news is you can run clean geo tests and holdouts with reporting you probably already have, plus a calendar and a spreadsheet.
In this guide, we will walk you through how to run a credible lift test without a data team, what to watch so you do not misread results, and how to reduce risk when you pause spend in a control group. If you run Google Ads, Meta Ads, or Amazon Ads for an ecommerce brand, this is the approach you can actually put into play.
Incrementality testing: what it really measures (and why ROAS can mislead you)
Most reporting dashboards are built on attribution. The platform looks at clicks and views and tries to assign credit for a conversion. That can be useful for day to day optimization, but it does not prove that your ads caused the sale.
Here is the issue you have likely seen in the real world. Someone searches your brand because a friend recommended you, they buy, and your branded search campaign takes the win. Or someone was already planning to reorder, sees a retargeting ad on Meta, and the platform reports a great ROAS. It might be accurate that the ad was involved, but it does not mean the ad created net new revenue.
Incrementality testing is about isolating what changed because ads were running. You compare outcomes with ads versus outcomes without ads, then measure the difference as incremental lift.
This matters more now because user level tracking is less dependable. iOS changes and cookie deprecation have made platform attribution less trustworthy, which is exactly why more brands are moving toward lift testing.
How incrementality testing works: test vs control and simple lift math
Every lift test has the same core structure:
Test group: keeps ads running as usual
Control group: has those ads paused or meaningfully reduced
If the test group beats the control group, the gap is your lift.
A straightforward calculation looks like this:
Incremental lift (%) = (Test conversions − Control conversions) ÷ Control conversions
Quick example: your test markets drive 1,120 purchases while control drives 1,000 purchases. Lift is 12%. In everyday terms, it suggests about 120 purchases were created by paid ads, not just “counted” by them.
Once you have lift, you can translate it into business outcomes like incremental revenue, incremental contribution margin, or incremental ROAS. That is where this becomes more than a measurement exercise. It becomes a cleaner way to make budget calls without politics.
Geo incrementality testing: the most practical setup when you have no data team
If you are not running user level experiments or you do not have a measurement engineer on call, geo testing is your best starting point. You split locations into two sets, keep ads active in one set, and hold out spend in the other set for a defined window. Then you compare performance.
Triple Whale has a helpful overview of why geo based holdouts are one of the most accessible incrementality testing methods for lean teams at Triple Whale: Incrementality testing methods.
Why geo tests work in the real world:
You can use data you already trust like Shopify or your ecommerce platform, plus geo breakouts in Google Ads and Meta Ads.
You are not depending on fragile user matching or perfect tags.
You can run them fast enough to influence next month’s budget, not next quarter’s.
The tradeoff is discipline. Geo tests are easy to do, and also easy to mess up if your markets are mismatched or your holdout leaks.
Where incrementality testing usually surprises you (especially branded search and retargeting)
One of the best reasons to run incrementality testing is that it often changes how you view channel performance. Digital Applied reviewed 225 DTC incrementality tests and found platform reported ROAS can differ a lot from true incremental ROAS, with branded search and retargeting often over credited for conversions that were already happening at Digital Applied: Incrementality testing and causal lift.
This does not mean branded search or retargeting are pointless. It means they are often better at capturing demand than creating it. If you only optimize on in platform ROAS, you can accidentally starve the campaigns that generate new customers and overspend on the ones that mop up existing intent.
We see this pattern a lot:
Meta prospecting looks “expensive” in platform, but shows strong lift when tested properly.
Google branded search looks like a hero in last click reporting, but the incremental lift is smaller than expected.
Retargeting looks clean and efficient, but the control group keeps converting anyway.
Incrementality is what gives you the confidence to rebalance budgets toward what actually grows the business.
How to run a geo lift test step by step (4 to 6 weeks and a spreadsheet)
This is the approach we recommend when you want a credible result without advanced tooling. It is simple, but it is not casual. You will still want to treat it like an experiment.
Pick one lever to test. Choose one channel or one campaign type. Common first tests are Meta prospecting, Meta retargeting, Google non brand, or Google branded search.
Choose 4 to 8 geos with stable volume. States or DMAs work well. If you pick tiny markets, you will get noise instead of signal.
Match test and control based on history. Look back 6 to 12 weeks and pair markets that had similar revenue, orders, and efficiency before the test. You want the groups to look boringly similar.
Lock the window. Plan on 4 to 6 weeks for most ecommerce brands. If you are lower volume, go 6 to 8 weeks.
Define a clean holdout rule. In control markets, pause the campaigns you are testing or cap budgets close to zero. Keep everything else consistent.
Track one primary outcome. Usually purchases and revenue. If you are lead gen, track qualified leads or closed won revenue if you can.
Calculate lift and sanity check it. Compare test versus control for purchases, revenue, and blended efficiency metrics like MER.
If you want a straightforward reporting setup that keeps these tests tied to business outcomes, our post on building a paid media dashboard is a good companion read: Paid media dashboard: 12 numbers leaders need.
Geo holdout mistakes that quietly ruin incrementality testing
Most failed tests do not fail because the math is hard. They fail because the setup is messy. Fusepoint’s guide makes a good point that sample size and stable conditions matter more than most teams expect at Fusepoint Insights: Mastering incrementality testing.
Here are the problems we see most often when a geo holdout “doesn’t work”:
Too little conversion volume. If each group only generates a few purchases per week, you are basically guessing. Expand geos or extend the test.
Bad timing. Big promos, pricing changes, product launches, or inventory swings can drown out the impact of ads.
Leaky controls. Control markets still get spend because targeting is not tight, Performance Max spills across locations, or settings are not actually enforcing exclusions.
Too many moving parts. If you change creative, landing pages, and offer strategy mid test, you cannot untangle what caused lift.
Missing the blended view. Lift is only half the story. You still need to know if the incremental sales were worth the spend given your margins.
That last point is worth slowing down for. Incrementality testing alone will not fix budget allocation unless you interpret lift alongside blended metrics like MER. We agree. Lift tells you what was caused by ads. MER helps you decide if that outcome was efficient enough to scale.
How to lower the risk of a holdout test (without making it useless)
Let’s be honest. The hardest part of a geo holdout is the moment you pause spend and think, “We are really doing this.” You are. And you should. A short controlled pause is often cheaper than months of over funding campaigns that look great inside the platform but do not move the business.
Ways to reduce downside while keeping the test credible:
Start with a smaller holdout. Use 10 to 20% of your revenue footprint, not half the map.
Test one channel at a time. Keep the rest of your machine running so you are not creating chaos.
Pick representative markets. Do not hold out your single best state just to “make the test faster.” That is a good way to lose internal support.
Write the rules down before you start. Decide the length, primary KPI, and decision thresholds up front.
In practice, many brands find the short term hit is smaller than expected. That is often the first clue that attribution has been flattering a campaign for a while.
Turning lift into budget decisions you can defend
Once the test ends, do not stop at “lift was 8%.” Your CFO and your founder want to know what you are changing next.
If lift is strong and blended efficiency holds: scale gradually and keep monitoring MER and contribution margin.
If lift is weak but platform ROAS is high: you likely have a demand capture campaign. Consider reducing spend, narrowing the audience, or reframing it as a lower priority tactic.
If lift is positive but MER worsens: the channel might be incremental but too expensive. That points to creative, offer, feed quality, or audience strategy.
If results are unclear: extend the window, choose higher volume geos, or test a different lever. Inconclusive is usually a design signal, not a dead end.
If you are also weighing other measurement approaches, we broke down what to use and when in our guide here: Marketing mix modeling vs multi touch attribution for ecommerce measurement. Most teams do best with a mix. Use incrementality tests to validate channels and tactics, then rely on consistent cross channel reporting to run the business week to week.
Incrementality testing checklist for lean ecommerce teams
If you want a quick pre flight checklist before you run your first geo holdout, use this:
Scope is tight: one channel or campaign type per test.
Geos are solid: 4 to 8 markets with stable volume and clean targeting boundaries.
Timing is clean: avoid promos, launches, and pricing shifts.
Holdout is enforced: clear exclusions or budget caps so control stays unexposed.
Primary KPI is chosen: purchases and revenue, or qualified leads.
Decision rules are set: what lift triggers scale, hold, or cut.
If you want help setting this up across Google and Meta in a way your finance team will trust, take a look at how we run profit first paid media at PPC Boost services. If Google is your main growth lever, our Google Ads management approach is here: Google Ads management. For social, our Meta Ads management page is here: Meta Ads management.
FAQ: incrementality testing, geo tests, and holdouts
What is incrementality testing in paid media?
Incrementality testing measures how many conversions or how much revenue your ads caused that would not have happened otherwise. It is designed to answer the “would we still have gotten this sale without ads?” question.
What is a lift test?
A lift test is a common type of incrementality test where you compare a test group that receives ads to a control group that does not. The difference between the two is the lift.
How long should a geo incrementality test run?
For most ecommerce brands, 4 to 6 weeks is a solid window. If your conversion volume is low, run it 6 to 8 weeks so you are not making a call based on randomness.
How big should the holdout be?
If you are new to this, start with 10 to 20% of your revenue footprint. You can run larger holdouts later once leadership is comfortable with the process.
Which channel should you test first?
Start where you have the most spend or the most uncertainty. Meta prospecting, Google non brand, branded search, and retargeting are common starting points. Branded and retargeting often look excellent in platform, so they are worth validating early.
Can you run incrementality testing on Amazon Ads?
Yes, but geo controls can be trickier because Amazon demand and inventory can shift across regions. You can still test by isolating regions where feasible or by holding out specific tactics while keeping everything else steady.
What if the test shows low incrementality?
Do not treat it like a verdict on paid media. Treat it like a map. Low lift can mean you are over invested in demand capture, your creative is not doing enough work, or your offer is not strong. Reallocate, improve inputs, and retest.
Conclusion: incrementality testing is a practical growth skill now
Incrementality testing is no longer reserved for enterprise teams with data scientists and custom tooling. With a thoughtful geo setup, a clean holdout, and basic lift math, you can get to the truth fast and cut the waste that attribution tends to hide.
If you want a second set of eyes on what to test first, how to pick markets, or how to turn lift into profit first budget moves, PPC Boost can help you design and run the test as part of your paid media strategy. You will make fewer guesses, keep more control, and scale decisions you can actually defend. If you want to see what it is like to work with us, you can also check out our reviews on Clutch.

