We design, run, and read geo-lift tests. Media runs in some markets, the rest are held out, and you get the revenue your spend actually caused. Every test is sized before launch, so the answer comes back readable.
Attribution gives an ad credit for every sale it was near. Plenty of those customers were already coming: they searched the brand, clicked a retargeting ad on the way, and bought what they'd have bought anyway. The only way to separate caused from correlated is to hold markets out and compare.
In the sample test below, the platform credited the campaign with 3.4x ROAS. The holdout measured 1.20x.
Most programs have a line item riding on sales that would have happened anyway. It looks great in-platform and does nothing for revenue. A test finds it.
Know the incremental return before you raise a budget cap, not after. A measured iROAS turns "let's try a bit more" into a plan.
One measured result with a confidence interval beats three dashboards each claiming the same sale. It's the number to put in the plan.
CTV, podcasts, out-of-home, and upper-funnel video rarely get fair credit in click-based attribution. A geo test measures them the same way it measures search.
Diagrams are illustrative.
Based on patterns we see often, not any single client's numbers.
A brand's Meta and Google buyers were each hitting their own channel targets, but nobody was checking whether the combined spend was incremental. A geo-lift test showed a chunk of a six-figure monthly program was paying for sales that would have happened anyway, and the budget moved to a real test elsewhere.
A brand had capped its best campaign at the same budget for two years out of caution. An incrementality read showed it still paying back well past that ceiling, so scaling it funded itself within weeks.
If spend can be changed in some markets and left alone in others, it can be tested, including the channels click-based attribution handles worst.
The channel attribution misses most. Measured on sales in the markets that saw it, not on view-through claims.
Before you double Meta or TikTok, find out whether the next dollar still pays back.
How much of it would have come in anyway? Pull it back in a few markets and watch what happens to total sales.
Awareness spend that never gets credit for a click, measured on the revenue it moves.
Direct mail, radio, podcasts, and out-of-home. If it runs by market, it can be read by market.
Point the same test at first orders only, when the question is growth rather than total sales.
Run two channels in separate groups of markets against one shared holdout. You learn what each channel adds on its own and which one earns more per dollar, without two tests months apart under different conditions.
Splitting markets splits statistical power. We check each cell can be read on its own before recommending it.
Run the same channel at two budget levels in separate groups of markets, against one shared holdout. You get the return at today's budget and the return on the extra spend, which is the number that decides whether to raise it.
A channel can pay back well on average and still waste its last dollars. Average ROAS hides that. A budget test shows it.
Here the extra spend returns about half what the first dollars did. Worth knowing before the cap goes up.
Platforms grade their own homework. A holdout market has no stake in the answer.
Using your daily sales by market, we find test markets whose history can be matched, day for day, by a weighted blend of other markets. Then we size the budget so the lift you care about is detectable. If it isn't, we say so before you spend anything.
Spend goes into the test markets for 4 to 8 weeks while everything else stays exactly as it was, plus two weeks afterward so delayed purchases still count. No pixels, no platform cooperation, no tracking prerequisites: the test only needs your own sales data.
We compare what the test markets did against what the control says they would have done. Orders and revenue are read side by side. You get the lift, a confidence interval, incremental ROAS, and a plain statement of how sure we are.
From a real run of our testing platform on Aurelia, an illustrative direct-to-consumer brand with realistic daily sales by market. The brand is fictional. The method and every number below are what the platform produced. The design is sized to what the data can support, not worked backward from a budget.
Same brand, same markets, after the campaign ran. The control is the forecast of what the test markets would have done without the spend.
3.4xPlatform-reported
CTV caused about $1.20 of revenue for every $1 spent in the test markets. The platform reported $3.40.
You keep control of the media buy. We design the test, tell you exactly what to change and where, watch it while it runs, and read the result.
From a sales export: whether a test will produce a readable answer for your brand, and on what dates.
Test markets, duration, spend by market, the smallest lift it can detect, and a placebo check on past data.
What to turn on, what to hold steady, and a directional read halfway through.
Lift, incremental ROAS with a confidence interval, reported vs. measured, and a plain recommendation.
The next question worth answering, based on what this one showed.
Markets, metric, and the decision rule are locked and dated before spend starts. The answer can't be worked backward.
Every design is rerun on past windows with no campaign. You see how often it finds a lift that isn't there.
The readout ships with the test and control numbers by day, so your analyst can recompute the lift.
The ecommerce media team behind it. wildlightmedia.com
A test that can't produce a readable answer isn't cheap, it's wasted. Before any budget moves, we check your data against these conditions and tell you plainly whether a test will work, what it would take, or what to do instead.
Enough daily volume in the markets you would test. If your biggest markets see only a few orders a day each, a real lift can disappear into ordinary day-to-day noise.
Revenue that isn't driven by a handful of huge orders. If one purchase can move a market's month, we measure orders instead of revenue.
A season with signal. Testing in your slowest months can work, but the lift has to clear a higher bar. We'll show you the bar for your dates.
Control markets you can leave alone. The holdout only works if spend there stays steady for the length of the test.
Daily sales by location for as far back as you have it, ideally two years. A Shopify or order-system export by ZIP code is enough; we roll it up to TV markets. No pixels, no platform access.
Usually 4 to 8 weeks of spend plus two weeks for late conversions to settle. The design tells you the exact length for your brand and season.
Enough to move sales in the test markets by more than their normal noise. The design sizes it for you. For many brands that lands in the tens of thousands over the test, spent only in the test markets.
No. Everything else keeps running as usual in every market. The test only changes one thing, in the test markets, so its effect can be isolated.
Then the channel didn't move sales by as much as the test could detect, which is a finding in itself. Because we size the test before launch, "inconclusive" means the effect was small, not that the test was too weak to see a meaningful one.
Orders are steadier. Revenue can swing on a handful of large purchases in one market on one day, which widens the confidence interval. We read both; when big orders are common, orders lead and revenue is recovered with your average order value.
Yes, with a multi-cell design: each channel gets its own group of markets against a shared holdout. It answers which channel earns more per dollar in one window. The catch is that each cell has less volume behind it, so we confirm every cell can be read on its own first.
Yes. A three-cell budget test runs the same channel at two spend levels in different markets, plus a shared holdout. It measures the return at your current budget and the return on the added budget, so you can see where more spend stops paying back.
Some channels, CTV especially, drive purchases days after someone sees the ad. A two-week window after spend stops lets those orders land, so the channel isn't undercounted.
A platform's own lift study is graded by the platform, inside its own walls, on its own users. A geo test is measured on your sales data, the same way for every channel, so results can be compared side by side.
MMM estimates every channel at once from history and is only as good as its assumptions. A geo test is an experiment: it changes one thing on purpose and measures the result. The two work well together; a test result is the best way to calibrate an MMM.
A few numbers are enough for a first answer. You'll hear back within two business days with a yes, a no, or what it would take.