A pilot campaign produced results strong enough to justify a national rollout, and the rollout produced almost nothing. The difference between the two was not the advertising, and it was visible in advance if anyone had looked for it.

The result that did not replicate

The pilot ran in a single city for eight weeks. Sales in that market rose clearly against the preceding period and against the rest of the country.

The effect was large enough to survive the usual scepticism. It was not seasonal, it did not coincide with a promotion, and it decayed at a plausible rate after the flight ended.

The rollout used the same creative, the same channel mix, and proportionally similar weight in twelve other markets.

Two of those markets responded. The others produced movement indistinguishable from noise.

When a result fails to replicate that comprehensively, the cause is usually a condition in the original market rather than a flaw in the execution.

Distribution was the hidden variable

The pilot city had unusually dense retail availability for the product. It was stocked in most relevant stores, at eye level, in the format the ads showed.

Across the other markets, availability varied enormously. In several, the product was carried by a minority of stores and often in a different pack size.

Advertising can create an intention to buy. It cannot create a shelf.

A viewer persuaded in a market with thin distribution walks into a shop, does not find the product, and substitutes. The persuasion is spent on a competitor.

Every element of the campaign performed identically in both cases. Only the last step differed, and the last step is the one that produces revenue.

Media weight against availability

The rollout allocated budget by population, which is the standard method and is wrong whenever availability varies.

Spending in proportion to population assumes that a persuaded person is equally likely to complete a purchase everywhere. That assumption fails hard in physical retail.

A more defensible allocation weights by availability, so markets with thin distribution receive less until distribution improves.

That feels backwards to teams who want to build markets where the brand is weak, and it is still correct in the short term.

Advertising into a market that cannot fulfil the demand converts budget into awareness for the category, which competitors collect.

Density and word of mouth

The pilot city had a second advantage that is easy to overlook. The product had a visible existing user base there, concentrated enough that buyers encountered it socially.

Advertising performs better where it confirms something the viewer has already half seen. It performs worse where it is the only source of information.

This is why pilot markets are so often chosen badly. Teams pick a market where the brand already does well, because the pilot is more likely to succeed there.

A pilot chosen for its likelihood of success measures the market, not the campaign.

The useful pilot runs in a market that resembles the average of where the rollout will go, including its weaknesses.

Why the creative got the credit

The creative was the most visible thing that changed, so it received the explanation. This happens by default in almost every post-campaign review.

Nothing else about the pilot was new. The channels were familiar, the budget was ordinary, and the only novel element was the work itself.

Attributing the result to the novel element is a reasonable instinct and it skips the question of what conditions the result depended on.

The review also had no data on availability, because retail distribution sat with a different team and was not part of the marketing report.

Variables that live in another department are invisible to analysis in a way that has nothing to do with their importance.

The rollout that failed anyway

Once the rollout was underway, the weekly reporting showed flat sales and healthy media metrics. Reach was delivered, frequency was on plan, and creative engagement matched the pilot.

That combination is the signature of a fulfilment problem rather than a communication problem. The message was received and could not be acted on.

It took several weeks to reach that reading, because the first response to flat sales is always to adjust the media.

Adjusting the media consumed most of the remaining budget without changing the outcome, which is the usual course.

The campaign was eventually paused in the markets with the thinnest distribution and the remaining budget concentrated where the product could actually be bought.

What we check before scaling now

Before any pilot is scaled, we now write down what was true in the pilot market that might not be true elsewhere. The list is usually short and usually contains the answer.

Availability is the first item for anything sold physically. Delivery coverage, service area, and stock depth are the equivalents in other categories.

Competitive intensity is the second, because a market where a competitor is spending heavily behaves differently from one where nobody is.

Existing brand penetration is the third, since advertising into an established base is a different task from advertising into an unfamiliar one.

If any of those differ materially, the pilot result is a measurement of that market and cannot be extrapolated without a second test elsewhere.

Where geography is not the explanation

Not every failure to replicate is about place. The same pattern appears across time, where a result depends on a condition that has since changed.

A campaign that worked during a period of high category demand may fail identically when demand normalises, with no geographic variation involved.

It also appears across audience segments, where a result driven by an existing customer base disappears when the same creative is shown to strangers.

The general form is the same in each case. A result is a product of the work and the conditions, and only one of those travels.

Separating them requires a second test under different conditions, which is slower than scaling immediately and considerably cheaper than scaling wrongly.