BusinessDigital MarketingNewswireStartupsTechnology

How to Run a CTV Incrementality Test (and Why Most Fail)

▼ Summary

– Most CTV incrementality tests fail because they compare ad viewers to non-viewers, capturing targeting effects rather than true ad impact, requiring a pre-decided control group withheld from all advertising.
– For B2B, the account is the correct unit of randomization, splitting the target account list into test and control groups to align with how B2B reporting is structured.
– The strongest testing method is an account-matched holdout with hard suppression, while media mix modeling is a fallback for planning, not causal proof.
– Tests often fail due to insufficient sample size for rare B2B conversions; if the test cannot reach statistical significance, use an upper-funnel metric or postpone testing.
– Report results to CFOs in terms of incremental pipeline, cost per incremental opportunity, and payback period, not lift percentages.

Most connected TV (CTV) incrementality tests are doomed before a single ad airs. The failure isn’t due to poor media performance; it stems from generating a seemingly convincing number that proves nothing at all.

This issue grows more critical with each passing quarter as budget holders demand stronger evidence that CTV delivers. According to the IAB’s 2026 Outlook Study, cross-platform measurement became a top priority for 72% of advertisers this year, up from 64% in 2025. Marketers must demonstrate that streaming investments produce results that search and social cannot replicate.

The most transparent way to tackle this challenge is through an incrementality test. This method involves withholding a control group from advertising, exposing the remaining audience, and measuring the outcome difference. While the concept sounds simple, execution proves difficult. Even the most well-funded advertising teams have struggled to answer the incrementality question. If the biggest spenders find it tough, the rest of us need a reliable approach.

Here is how to design a CTV incrementality test for B2B that withstands scrutiny, avoid the common errors that quietly undermine most tests, and convert the result into a metric your CFO will trust.

Incrementality isn’t attribution

Attribution focuses on identifying which ad exposure deserves credit for a conversion. Incrementality, by contrast, asks a harder question: Would that conversion have happened anyway without the ad? This distinction exposes the weaknesses of CTV measurement.

The most common testing method compares people who saw the ad with those who didn’t, then reports the difference as uplift. This seems rigorous but is fundamentally flawed. Viewers of your CTV ad are inherently different from non-viewers. They stream more content, may already be interested in your product, and are deliberately targeted. This measurement captures the effects of targeting, not the ad’s impact. It’s like crediting someone’s fitness to gym attendance simply because they chose to go.

True incrementality requires deciding, before any impression runs, which accounts you will intentionally avoid advertising to. That withheld group is the control. Their behavior represents what would have happened without your ad exposure. Everything the exposed group does above that baseline is incremental. If you skip this step, no later analysis can compensate.

Pick the right unit to randomize

Unlike most digital marketing, CTV lacks a persistent cookie. Ads serve to households and IP addresses, with multiple people potentially watching the same screen simultaneously. This makes a clean individual-level split impossible. Move up a level instead.

B2B still targets individuals across channels (such as named buying committee members on LinkedIn, via email, or in your ABM program), but the committee makes purchase decisions as a group. ABM already reports at the account level. Complex B2B deals involve multiple decision-makers, not just one. Randomize where the decision lands. The account is the correct randomization unit.

For account-based programs, split your target account list into two groups. Half the accounts are eligible for CTV advertising, while the other half are excluded from all advertising, including streaming. This aligns with how B2B reporting is typically structured, focusing on accounts and their progression through the sales pipeline.

Choose the testing method and understand what it proves

Not all testing methods are equal. From strongest to weakest proof:

An account-matched holdout with hard suppression is the strongest test most B2B teams can run. Exclude the control accounts’ IP addresses and device IDs from every related line item, not just the test campaign.

A geo-matched market test uses a difference-in-differences design, highlighting the change in test markets minus the change in control markets over the same period.

Media mix modeling is a fallback, not a test. It measures correlation between spend and outcomes. It’s useful for planning, but don’t present it as causal proof for a single channel.

Select your metric before testing begins. Avoid early-stage metrics like impressions, as they indicate delivery rather than impact.

For B2B, the primary metric should be qualified pipeline or opportunities generated from target accounts. You can also use faster signals, such as target account site visits, increases in branded search, improved paid social metrics, and demo requests, to gauge whether the test is working before the pipeline matures.

Ensure your test size is statistically significant

The main reason most incrementality tests fail isn’t the media. B2B conversions are both rare and slow. If your baseline opportunity rate is low and you hold back only a small number of accounts, you won’t have enough conversions to detect any meaningful improvement. The test will then indicate “no effect,” regardless of how well CTV performed.

Before committing, check whether the test can reach statistical significance. Three factors determine this: how often these accounts convert today, how large a lift you’d need to call the result real, and how many accounts you can place on each side of the split.

Unfortunately, when conversions are rare and the account list is short, no sample will surface an effect, even if CTV is working. If that’s your situation, you have two options: make an upper-funnel metric the primary one and treat pipeline as directional, or don’t run the test yet. A test too small to detect an effect is worse than no test at all, because it produces a confident wrong answer.

Protect the control group

Keep two things in mind when conducting incrementality tests. First, measure both the test and control groups before the campaign starts to ensure they’re on parallel trends. If the control group already performed better, any perceived lift from the campaign may be misleading.

Second, protect against data leakage. Factors such as shared corporate IP addresses, co-viewing, and users exposed to multiple active campaigns within the same account can contaminate your control group. Enforce suppression across all line items and monitor it closely during the test.

Also, allow enough time for the campaign to take effect. Upper-funnel signals typically take weeks to emerge, while the sales pipeline requires the campaign duration plus an appropriate lookback period aligned with your sales cycle. If you measure a two-week test against pipeline, it may seem like a failure simply because the pipeline hasn’t had enough time to develop.

Why disciplined teams keep their budget

A clean incrementality test only matters if the result lands with the person who controls the budget. Your CFO doesn’t think in lift percentages. They think in pipeline, cost per opportunity, and how long the spend takes to pay for itself. Report incremental pipeline, cost per incremental opportunity, and payback period. Those are the numbers you can defend in a budget meeting.

The teams that keep their CTV budget through 2026 won’t be the ones with the best-looking dashboards. They’ll be the ones who decided, before the campaign ever launched, which accounts they were willing to leave alone. None of that is glamorous. It’s the difference between knowing CTV worked and hoping it did. Build the control group first. The rest follows.

(Source: MarTech)

Topics

ctv incrementality 98% test design flaws 95% control group setup 92% randomization unit 88% testing methods 87% statistical significance 85% metric selection 82% control group protection 80% budget justification 78% b2b advertising context 76%
Show More