The reliable way to know whether advertising works is to withhold it from a random portion of the audience and compare. That method is not free, and the price is what stops most teams from using it.
The cost is forgone revenue, not fees
A holdout group does not receive the advertising, so it buys less. That gap is the measurement, and it is also a direct loss for the duration of the test.
Unlike a research budget, the cost does not appear as a line item. It appears as lower sales in a period someone is accountable for.
Which means the person who must approve the test is often the person whose numbers it will damage, an incentive problem no methodology can solve.
Group size trades cost against certainty
A small holdout costs little and produces a result too noisy to act on. A large one gives a clear answer and forgoes more revenue.
The required size depends on how large an effect the test needs to detect, and detecting small effects requires disproportionately large groups.
Deciding in advance what size of effect would change a decision keeps the test proportionate, and skipping that step is how teams run tests that could never have concluded anything.
Contamination is easy and invisible
Held-out customers still see the brand through channels the test does not control, and members of the same household frequently land in different groups.
Each of these narrows the measured gap, so a contaminated test understates the true effect rather than overstating it.
A result showing little effect from a leaky test is therefore not evidence that the advertising did nothing, though it is routinely read that way.
Duration has to cover the purchase cycle
For an infrequently purchased category, an effect takes months to appear in sales, and a test concluded after four weeks measures almost nothing.
Long tests are harder to protect, since campaigns change and someone eventually asks why a group is being excluded.
Matching test length to the buying cycle rather than to the reporting calendar is the discipline, and it is the one most often abandoned under pressure.
Why it remains worth the money
Platform-reported results credit conversions that would have happened regardless, and the size of that overstatement is unknowable without a comparison group.
A single well-run holdout frequently reveals that a channel is worth a fraction of its reported value, or a multiple of it.
Either finding redirects spending by more than any amount of dashboard optimization, which is why the cost is best understood as the price of a decision rather than of a report.