Skip to content
Measurement 9 min read

How to Run Incrementality Tests on Meta and ChatGPT Ads

Platform dashboards tell you what they claim. Incrementality testing tells you what actually happened. Here is how to run simple tests on Meta and ChatGPT ads.

Every ad platform has attribution built in. Meta says your ads drove 500 purchases this month. ChatGPT says 120. Google says 400. Add them up and you get 1,020 attributed purchases. But your store only shipped 600 orders.

This is not fraud. It is overlapping attribution. The same customer may have seen a Meta ad, searched on Google, asked ChatGPT a question and then typed your URL directly. Each platform takes credit. The total is impossible.

Incrementality testing cuts through this. Instead of asking "how many purchases did the platform claim," it asks "how many purchases would not have happened without this spend." That is the question that actually drives profit.

This guide covers what incrementality means, why it matters, simple test designs you can run without a data science team, and how to interpret results for Meta and ChatGPT ads.

What incrementality means

Incrementality is the difference between what happened with ads and what would have happened without them.

If you spend $10,000 on Meta ads and your store does $50,000 in revenue, that does not mean Meta drove $50,000. Some of those customers would have bought anyway through organic search, direct traffic, email or word of mouth. Incrementality measures the lift: the additional revenue caused by the ads.

Suppose you ran a perfect experiment and found that without Meta ads, your store would have done $38,000. The incremental revenue from Meta is $12,000, not $50,000. Your true return on that $10,000 is 1.2x, not 5x.

That difference matters. A brand making decisions based on 5x reported return will spend very differently from one that knows the true return is 1.2x.

Why platform attribution overcounts

Platform attribution has two structural problems:

  1. Credit for existing intent. If someone was already planning to buy and happens to see your ad first, the platform takes credit for a sale it did not cause.
  1. Overlap with other channels. A customer who sees ads on three platforms gives credit to all three. Last-click models reduce this but do not eliminate it.

ChatGPT ads have an additional wrinkle: they appear in research contexts. Someone asking "best running shoes for flat feet" may be deep in a purchase journey that started elsewhere. If they click your ad and buy, ChatGPT reports the conversion. But that person might have found you through another channel anyway.

None of this means platform data is useless. It helps you compare ads within a platform. But it cannot tell you whether the platform itself is earning its budget. For that, you need incrementality.

The simplest incrementality test: geo holdouts

You do not need a statistics degree to run a basic incrementality test. The simplest version uses geography.

How it works

  1. Split your market into regions. For example, if you sell across the US, divide states into two groups that are as similar as possible in historical performance.
  1. Keep ads running in one group (test). Turn them off in the other (holdout).
  1. Run for 2 to 4 weeks. Long enough to see real purchasing behavior, short enough to limit risk.
  1. Compare results. Did the test group outperform the holdout by more than historical variance?

Example setup for Meta

Say you run Meta ads nationally. You pick 10 states for the holdout, chosen to match your test states in historical revenue per capita. You turn off Meta ads in those 10 states using exclusion targeting.

After three weeks, you see:

GroupAd spendRevenueRevenue per capita index
Test (40 states)$15,000$72,0001.00
Holdout (10 states)$0$14,4000.80

The holdout group did 80% as well per capita without ads. That suggests 20% of sales in the test group were incremental, meaning Meta drove about $14,400 of the $72,000 (20% of $72,000). Your incremental return on $15,000 spend is roughly 0.96x, not the 4.8x that $72,000 / $15,000 would imply.

This is a simplified illustration. Real tests need larger sample sizes and longer durations to reach statistical confidence. But even a rough estimate is better than taking platform attribution at face value.

Applying geo tests to ChatGPT ads

ChatGPT ads support geographic targeting by country and region. You can run the same type of test: hold out a set of regions, keep ads running elsewhere, and compare lift.

Because ChatGPT ad budgets are often smaller and conversion volumes lower than Meta, you may need longer test periods to accumulate enough data. A four-week test with at least a few hundred conversions in each group is a reasonable minimum. If your volume is too low for geo tests, consider the time-based alternative below.

Time-based incrementality tests

If your volume is too low for geo splits, or if geo targeting is impractical, you can test over time.

How it works

  1. Establish a baseline. Run ads at a steady budget for 4 weeks and track total revenue.
  1. Turn off ads for 2 to 4 weeks. Keep everything else constant: email, organic, other channels.
  1. Compare revenue. Did revenue drop during the off period? By how much?
  1. Turn ads back on. Did revenue recover?

Example for ChatGPT ads

You run ChatGPT ads at $75 per day for four weeks. Total revenue during that period is $28,000. You then pause ChatGPT ads for three weeks. Revenue during the off period is $22,400. You turn ads back on and revenue returns to roughly the same level.

If we assume no other changes, the revenue drop of $5,600 over three weeks suggests ChatGPT ads contributed about $5,600 of incremental revenue. You spent $75 x 21 = $1,575 during that equivalent time window while ads were running. That implies an incremental return of about 3.6x on the paused spend.

This method is messier than geo holdouts. Seasonality, promotions, competitor activity and other factors can affect revenue week to week. But it is doable at small scale and gives directional insight.

Conversion lift studies (platform-run)

Meta offers a built-in tool called Conversion Lift that runs randomized controlled tests. Some users are shown your ads; a control group is held out. Meta compares conversion rates between the two.

Pros

  • Randomized at the user level, which is statistically cleaner than geo splits.
  • Meta handles the measurement and reporting.

Cons

  • Requires significant spend and volume (Meta recommends substantial budgets to reach statistical significance).
  • You are trusting Meta to measure its own effectiveness.
  • Not available to all advertisers.

If you qualify and have the budget, Conversion Lift is worth running at least once to calibrate your other methods. But it is not a replacement for independent measurement.

ChatGPT ads do not currently offer an equivalent built-in lift study. You will need to use geo or time-based methods.

Interpreting incrementality results

Once you have results, what do you do with them?

Calculate incremental return

Incremental return = Incremental revenue / Ad spend

If your geo test shows $12,000 of incremental revenue from $10,000 of spend, your incremental return is 1.2x. Compare that to your break-even threshold. A brand with 50% contribution margin breaks even at 2x. At 1.2x, you are losing money on a first-order basis.

Use the break-even ROAS calculator and the MER calculator to set your own thresholds.

Adjust for lifetime value

First-order break-even is not the whole story. If customers return for repeat purchases, you can tolerate a lower first-order return. A brand with a 2.5x LTV to CAC ratio can afford higher acquisition costs than the first order alone suggests.

Compare channels

Run incrementality tests on each major channel over time. You may find that Meta shows 3x in its dashboard but 1.5x incremental, while ChatGPT shows 2x in its dashboard and 1.8x incremental. The channel with lower reported return may actually be more efficient.

Use incrementality to set budgets

If you know the incremental return curve for a channel, you can find the point of diminishing returns. Spend until the marginal dollar returns less than your threshold. Our guide to how to calculate ad budget covers this logic in more detail.

Common mistakes in incrementality testing

Mistake 1: Running tests too short

Purchases do not happen instantly. Someone who sees an ad today may buy next week. A one-week test misses delayed conversions. Run at least two weeks, preferably four.

Mistake 2: Changing other variables

If you pause Meta ads and also launch a new email campaign, you cannot isolate what caused any change. Keep other channels constant during the test.

Mistake 3: Ignoring confidence intervals

Small sample sizes produce noisy results. A test that shows 1.5x incremental return might actually be anywhere from 0.8x to 2.2x if the confidence interval is wide. If possible, consult someone with statistics background to assess significance.

Mistake 4: Testing at the wrong spend level

Incrementality often varies with spend. The first $5,000 you spend may be highly incremental (reaching the most motivated buyers), while the next $5,000 is less so. A test at $10,000 per week will give different results than one at $2,000 per week. Test at or near your actual operating spend.

Mistake 5: Forgetting halo effects

Pausing ads may reduce brand awareness, which affects organic search and direct traffic. A simple revenue comparison may undercount ad value if there is a halo. Conversely, ads may cannibalize organic by capturing searches that would have happened anyway. These effects are hard to measure perfectly but worth acknowledging.

A simple incrementality testing calendar

You cannot test every channel every month. Here is a practical cadence for a brand spending across Meta and ChatGPT:

QuarterTest
Q1Meta geo holdout (3 weeks)
Q2ChatGPT time-based test (3 weeks on/off)
Q3Meta time-based test during historically stable period
Q4No tests during holiday peak; rely on prior learnings

Adjust based on your seasonality and volume. The goal is to calibrate incrementality at least once per channel per year, more often if you make major spend changes.

How SecondWin fits into measurement

SecondWin builds and manages ChatGPT ad campaigns for consumer brands. It starts from the proven messages in your niche (identified by looking at the longest-running ads in Meta's Ad Library) and writes original ChatGPT ads around the questions your buyers ask.

Because ChatGPT is a newer channel with less mature attribution, we recommend tracking performance through blended metrics like MER (marketing efficiency ratio) and periodic incrementality tests. SecondWin does not claim credit for sales; ad spend is billed by OpenAI directly to you, and you see results in your own store data.

Want to see what that would look like for your brand? Run a free analysis of your website to see the buyer questions and ad concepts we would start with. Plans are on the pricing page. SecondWin is independent and not affiliated with OpenAI.

FAQ

What is the difference between attribution and incrementality?

Attribution assigns credit to touchpoints in a conversion path. Incrementality measures whether those touchpoints actually caused conversions that would not have happened otherwise. Attribution is useful for comparing performance within a platform. Incrementality is necessary for knowing whether the platform itself is worth the spend.

How much budget do I need to run an incrementality test?

It depends on your conversion volume. For geo tests, you need enough conversions in both test and holdout regions to see a statistically meaningful difference. Roughly, aim for at least a few hundred conversions per group over the test period. For time-based tests, you need stable baseline data and enough volume to detect a lift or drop. Brands with very low volume may need longer test periods or alternative approaches like matched-market analysis.

Can I trust Meta's Conversion Lift study?

Meta's Conversion Lift uses randomized user-level holdouts, which is methodologically sound. The limitation is that Meta is measuring its own value, and some advertisers prefer independent verification. Running your own geo or time-based tests alongside Conversion Lift gives you a second data point.

How do I run incrementality tests on ChatGPT ads?

ChatGPT ads support geographic targeting, so you can run geo holdout tests by excluding certain regions from campaigns. You can also run time-based tests by pausing campaigns and comparing revenue during the off period. Because ChatGPT ad volumes are often smaller than Meta, you may need longer test durations to accumulate enough data for reliable conclusions.

Should I stop running ads during a holdout test?

In the holdout region or period, yes. That is the point: you need a true control to measure lift. The risk is lost sales during the test. Choose holdout regions or periods that represent a small enough share of your business that the short-term cost is acceptable for the long-term learning.

Free tools

All 36 tools