Incrementality testing · Geo-lift

Every channel claims the sale. Find out which one caused it.

We design, run, and read geo-lift tests. Media runs in some markets, the rest are held out, and you get the revenue your spend actually caused. Every test is sized before launch, so the answer comes back readable.

See if your brand can run a test See a sample readout
  • Sized before a dollar goes out
  • Measured on your own sales data
  • A readout in 6 to 8 weeks
01 — The gap

Reported ROAS and incremental ROAS are different numbers.

Attribution gives an ad credit for every sale it was near. Plenty of those customers were already coming: they searched the brand, clicked a retargeting ad on the way, and bought what they'd have bought anyway. The only way to separate caused from correlated is to hold markets out and compare.

In the sample test below, the platform credited the campaign with 3.4x ROAS. The holdout measured 1.20x.

02 — Why run one

One test changes how the next year of budget gets spent.

CausedWould have happened anyway
01

Cut spend that isn't causing sales.

Most programs have a line item riding on sales that would have happened anyway. It looks great in-platform and does nothing for revenue. A test finds it.

Old capStill paying back
02

Scale what works without guessing.

Know the incremental return before you raise a budget cap, not after. A measured iROAS turns "let's try a bit more" into a plan.

Meta
Google
Email
Test
03

Give finance a number it can use.

One measured result with a confidence interval beats three dashboards each claiming the same sale. It's the number to put in the plan.

CTV, by clicks
CTV, by geo test
04

Test channels platforms can't measure.

CTV, podcasts, out-of-home, and upper-funnel video rarely get fair credit in click-based attribution. A geo test measures them the same way it measures search.

Diagrams are illustrative.

What that looks like Composite scenarios

Based on patterns we see often, not any single client's numbers.

  1. A

    A brand's Meta and Google buyers were each hitting their own channel targets, but nobody was checking whether the combined spend was incremental. A geo-lift test showed a chunk of a six-figure monthly program was paying for sales that would have happened anyway, and the budget moved to a real test elsewhere.

  2. B

    A brand had capped its best campaign at the same budget for two years out of caution. An incrementality read showed it still paying back well past that ceiling, so scaling it funded itself within weeks.

03 — What you can test

Any channel you can turn up or down by market.

If spend can be changed in some markets and left alone in others, it can be tested, including the channels click-based attribution handles worst.

CTV and streaming

The channel attribution misses most. Measured on sales in the markets that saw it, not on view-through claims.

Paid social at a higher budget

Before you double Meta or TikTok, find out whether the next dollar still pays back.

Branded search

How much of it would have come in anyway? Pull it back in a few markets and watch what happens to total sales.

YouTube and upper-funnel video

Awareness spend that never gets credit for a click, measured on the revenue it moves.

Offline media

Direct mail, radio, podcasts, and out-of-home. If it runs by market, it can be read by market.

New-customer acquisition

Point the same test at first orders only, when the question is growth rather than total sales.

Two channels at once Multi-cell

One test window, three answers.

Run two channels in separate groups of markets against one shared holdout. You learn what each channel adds on its own and which one earns more per dollar, without two tests months apart under different conditions.

  • Cell A CTV
  • Cell B second channel
  • Holdout neither

Splitting markets splits statistical power. We check each cell can be read on its own before recommending it.

Budget levels 3-cell

Find out what the next dollar is worth.

Run the same channel at two budget levels in separate groups of markets, against one shared holdout. You get the return at today's budget and the return on the extra spend, which is the number that decides whether to raise it.

A channel can pay back well on average and still waste its last dollars. Average ROAS hides that. A budget test shows it.

Incremental revenue by spend levelIllustrative
  • Cell A current budget
  • Cell B 2x budget
  • Holdout none

Here the extra spend returns about half what the first dollars did. Worth knowing before the cap goes up.

Platforms grade their own homework. A holdout market has no stake in the answer.
Why the test only needs your sales data
04 — How it works

Three steps. The first one decides whether the other two are worth doing.

Design
Run · spend on in test markets
Post-period
Read
Wk 0Wk 2 · 4–8 weeks+2 weeksReadout
  1. 01

    Design

    Using your daily sales by market, we find test markets whose history can be matched, day for day, by a weighted blend of other markets. Then we size the budget so the lift you care about is detectable. If it isn't, we say so before you spend anything.

  2. 02

    Run

    Spend goes into the test markets for 4 to 8 weeks while everything else stays exactly as it was, plus two weeks afterward so delayed purchases still count. No pixels, no platform cooperation, no tracking prerequisites: the test only needs your own sales data.

  3. 03

    Read

    We compare what the test markets did against what the control says they would have done. Orders and revenue are read side by side. You get the lift, a confidence interval, incremental ROAS, and a plain statement of how sure we are.

05 — Sample test design

What you see before a dollar goes out.

From a real run of our testing platform on Aurelia, an illustrative direct-to-consumer brand with realistic daily sales by market. The brand is fictional. The method and every number below are what the platform produced. The design is sized to what the data can support, not worked backward from a budget.

Sample — geo-lift test designIllustrative brand
Test: CTV added in 10 markets · 45 daysControl modeled from holdout markets
Test markets1017% of revenue
Duration45 daysApr 6 – May 20, 2026
Smallest readable lift5%80%+ chance of detecting it
Recommended spend$26,000≈ $578/day across test markets
06 — Sample readout

What you get when the test ends.

Same brand, same markets, after the campaign ran. The control is the forecast of what the test markets would have done without the spend.

Sample — geo-lift readoutIllustrative brand
Test: CTV added in 10 markets · 45 daysControl modeled from holdout markets
Incremental ROAS 1.20x

3.4xPlatform-reported

CTV caused about $1.20 of revenue for every $1 spent in the test markets. The platform reported $3.40.

Test revenue$313.9K
−
Control (modeled)$282.8K
=
Lift in test markets$31.1K +11%Significant90% CI $14.6K–$47.6K
Change in revenue$31.1K
÷
Spend added$26,000
=
Incremental ROAS1.20xSignificant90% CI 0.56x–1.83x
07 — What's included

From fit check to next test, in one engagement.

You keep control of the media buy. We design the test, tell you exactly what to change and where, watch it while it runs, and read the result.

  1. 01
    Fit check

    From a sales export: whether a test will produce a readable answer for your brand, and on what dates.

  2. 02
    Test design

    Test markets, duration, spend by market, the smallest lift it can detect, and a placebo check on past data.

  3. 03
    Launch plan and midpoint check

    What to turn on, what to hold steady, and a directional read halfway through.

  4. 04
    Readout

    Lift, incremental ROAS with a confidence interval, reported vs. measured, and a plain recommendation.

  5. 05
    What to test next

    The next question worth answering, based on what this one showed.

08 — How to check our work

No logo wall. A method you can audit instead.

Design record · lockedSample
Locked
—
Question
Does CTV cause incremental revenue?
Primary metric
—
Markets
—
Decision rule
Scale if iROAS ≥ 1.0x at 90% confidence
  1. Written down before launch.

    Markets, metric, and the decision rule are locked and dated before spend starts. The answer can't be worked backward.

  2. Tested on nothing first.

    Every design is rerun on past windows with no campaign. You see how often it finds a lift that isn't there.

  3. The daily series is yours.

    The readout ships with the test and control numbers by day, so your analyst can recompute the lift.

  4. Run by Wildlight Media.

    The ecommerce media team behind it. wildlightmedia.com

09 — Is it a fit?

Some brands shouldn't run a geo test yet. We'll tell you which one you are.

A test that can't produce a readable answer isn't cheap, it's wasted. Before any budget moves, we check your data against these conditions and tell you plainly whether a test will work, what it would take, or what to do instead.

  1. Enough daily volume in the markets you would test. If your biggest markets see only a few orders a day each, a real lift can disappear into ordinary day-to-day noise.

  2. Revenue that isn't driven by a handful of huge orders. If one purchase can move a market's month, we measure orders instead of revenue.

  3. A season with signal. Testing in your slowest months can work, but the lift has to clear a higher bar. We'll show you the bar for your dates.

  4. Control markets you can leave alone. The holdout only works if spend there stays steady for the length of the test.

10 — Questions

What people ask before a first test.

What data do you need?

Daily sales by location for as far back as you have it, ideally two years. A Shopify or order-system export by ZIP code is enough; we roll it up to TV markets. No pixels, no platform access.

How long does a test take?

Usually 4 to 8 weeks of spend plus two weeks for late conversions to settle. The design tells you the exact length for your brand and season.

How much do I need to spend?

Enough to move sales in the test markets by more than their normal noise. The design sizes it for you. For many brands that lands in the tens of thousands over the test, spent only in the test markets.

Do I have to pause my other advertising?

No. Everything else keeps running as usual in every market. The test only changes one thing, in the test markets, so its effect can be isolated.

What if the result isn't significant?

Then the channel didn't move sales by as much as the test could detect, which is a finding in itself. Because we size the test before launch, "inconclusive" means the effect was small, not that the test was too weak to see a meaningful one.

Why measure orders instead of revenue?

Orders are steadier. Revenue can swing on a handful of large purchases in one market on one day, which widens the confidence interval. We read both; when big orders are common, orders lead and revenue is recovered with your average order value.

Can we test two channels at once?

Yes, with a multi-cell design: each channel gets its own group of markets against a shared holdout. It answers which channel earns more per dollar in one window. The catch is that each cell has less volume behind it, so we confirm every cell can be read on its own first.

Can a test tell us how much to spend?

Yes. A three-cell budget test runs the same channel at two spend levels in different markets, plus a shared holdout. It measures the return at your current budget and the return on the added budget, so you can see where more spend stops paying back.

Why keep measuring after the campaign ends?

Some channels, CTV especially, drive purchases days after someone sees the ad. A two-week window after spend stops lets those orders land, so the channel isn't undercounted.

How is this different from a platform lift study?

A platform's own lift study is graded by the platform, inside its own walls, on its own users. A geo test is measured on your sales data, the same way for every channel, so results can be compared side by side.

How is this different from media mix modeling?

MMM estimates every channel at once from history and is only as good as its assumptions. A geo test is an experiment: it changes one thing on purpose and measures the result. The two work well together; a test result is the best way to calibrate an MMM.

11 — Check your fit

See if your brand can run a test.

A few numbers are enough for a first answer. You'll hear back within two business days with a yes, a no, or what it would take.

Monthly online revenue
Orders per day (typical)