In 2012 a large online marketplace stopped bidding on searches containing its own name, and lost 0.5 percent of its clicks. Not 0.5 percent of paid clicks. Half a percent of total clicks, because organic results caught almost everything the ads had been catching.

The published finding is worded precisely: “almost all (99.5 percent) of the forgone click traffic from turning off brand keyword paid search was immediately captured by natural search traffic”. And the authors note that 99.5 percent is a lower bound on retention, since some of the remainder probably shifted to people typing the address directly.

That is the single most useful experiment ever run on this question, and its second half is harder still. Everything below comes from randomized or quasi-randomized field experiments, not from platform reporting, and the gap between the two is the subject.

The brand keyword result, and how it was measured

The design matters, because it is what makes the number trustworthy.

The company stopped all bidding on queries containing its brand name on two search engines while continuing to bid on a third, which served as a control. That allows a difference-in-differences estimate rather than a before-and-after comparison. Raw traffic from the treated engine fell 5.6 percent. After controlling for seasonality using the untreated engine, the actual loss was 0.529 percent of clicks, or roughly 1.5 percent of paid clicks alone.

A replication on the third engine four months later showed a 3.2 percent fall, and the authors flag the weakness honestly: there was no valid control during that test, and they expect a proper control would reduce the estimate, as it did in the first case. A separate German test with a control confirmed the pattern.

The mechanism is not mysterious. Someone searching for your brand name has already decided to come to you. The advertisement is buying a click you were going to receive anyway, and the only thing it changes is which line of the results page they click.

Effect of switching off brand keyword advertising at a large online marketplaceEffect of switching off brand keyword advertising at a large online marketplace, from a field experiment published in Econometrica in two thousand fifteen. In March two thousand twelve the company stopped all bidding on search queries containing its own brand name on two search engines while continuing to bid on a third engine, which therefore served as a control group and allowed a difference in differences estimate rather than a simple before and after comparison. Raw click traffic from the treated engine fell five point six percent immediately after the suspension. After controlling for seasonality using the untreated engine, the actual loss was zero point five two nine percent of total clicks, which corresponds to about one point five percent of paid clicks alone. The published finding states that almost all, ninety nine point five percent, of the forgone click traffic from turning off brand keyword paid search was immediately captured by natural search traffic, and the authors note that ninety nine point five percent is a lower bound on retention because some of the remaining half percent probably shifted to users typing the web address directly rather than searching. A replication on the third engine four months later showed a three point two percent fall in total referred traffic, and the authors flag honestly that there was no valid control group during that test and that they expect a proper control would reduce the estimate as it did in the first case, while a separate test in Germany with a control group confirmed the pattern. The mechanism is that a person searching for a brand name has already decided to visit that brand, so the advertisement buys a click the company was going to receive anyway and changes only which line of the results page the person clicks.Brand keyword ads switched offRaw fall in traffic from the treated engine-5.6%Controlled against the untreated engine-0.5%What the paper says”almost all (99.5 percent) of the forgone click traffic from turning off brand keyword paid search wasimmediately captured by natural search traffic”And 99.5 percent is a lower bound: some of the rest probably went to typing the address directly.The 0.5 percent of total clicks lost is about 1.5 percent of paid clicks alone.
Raw traffic fell 5.6 percent. Controlled against an untreated search engine, the real loss was 0.5 percent of clicks. Source : Blake, Nosko and Tadelis, Consumer Heterogeneity and Paid Search Effectiveness, Econometrica 83(1), 2015 (2015)

The non-brand experiment, and the number that should be famous

The main test was larger. The company bid on more than 100 million keywords, and the researchers switched advertising off in 68 of the 210 United States market areas for 60 days, against 142 controls. Attributed sales in the treated areas fell by more than 72 percent, confirming the manipulation worked.

Then they estimated the return on investment five different ways on the same data. The results are worth setting out in full, because the spread is the lesson:

MethodEstimated return
Naive regression+4,173%
Regression with market and date fixed effects+1,632%
Instrumental variables-22%
Instrumental variables, second specification-63%
Difference in differences-63%

The experimental estimate is negative 63 percent, with a confidence interval from negative 124 to negative 3, so it is statistically distinguishable from zero. The naive regression on the identical data reports a return of over four thousand percent.

That gap is not a curiosity. It is the difference between what an advertising dashboard shows and what the spending actually did.

The explanation the authors give is arithmetic. Even users buying more than fifty times a year arrive through a paid click on about 4 percent of their purchases. So a large share of the budget was being spent to intercept people who were already on their way.

Five estimates of the same advertising return on investment computed from the same experimental dataFive estimates of the same advertising return on investment, all computed from the same underlying data collected during a sixty day field experiment. The experiment switched paid search advertising off across sixty eight of the two hundred and ten United States designated market areas, against one hundred and forty two control areas, with the company bidding on more than one hundred million keywords. Attributed sales in treated areas fell by more than seventy two percent, confirming the manipulation took effect. The spend figure used was fifty one million dollars of annual United States search engine marketing, against United States gross revenue of two thousand eight hundred and eighty point six four million dollars, both restated from public filings. A naive ordinary least squares regression on the data produced an estimated return on investment of positive four thousand one hundred and seventy three percent, with a confidence interval from four thousand one hundred and thirty nine to four thousand two hundred and five percent. Adding market area and date fixed effects produced positive one thousand six hundred and thirty two percent, with a confidence interval from six hundred and ninety seven to two thousand two hundred and sixty five percent. An instrumental variables specification produced negative twenty two percent with a very wide confidence interval spanning negative two thousand one hundred and sixty eight to positive one thousand one hundred and ninety one percent. A second instrumental variables specification produced negative sixty three percent with a confidence interval from negative one hundred and twenty four to negative three percent. The difference in differences estimate produced the same negative sixty three percent with the same confidence interval. The experimental estimate is therefore negative sixty three percent and is statistically distinguishable from zero, while the naive regression on identical data reports a return exceeding four thousand percent. The explanation given by the authors is arithmetic: even users purchasing more than fifty times a year arrive through a paid click on only about four percent of their purchases, so a large share of the budget intercepted people who were already on their way.Same data, five answerslossgainNaive regression+4,173%With fixed effects+1,632%Instrumental variables-22%Instrumental variables, second-63%Difference in differences-63%The experimental estimateNegative 63 percent, confidence interval -124 to -3, so distinguishable from zero. 68 treated market areas, 60 days.
The naive method reports over four thousand percent. The experimental method reports negative sixty three. Source : Blake, Nosko and Tadelis, Econometrica 83(1), 2015, Table 1 (2015)

Who the advertising actually reached

The heterogeneity analysis is the part that makes the result usable rather than merely deflating.

The researchers split users into eleven segments by how many purchases they had made in the previous twelve months. The pattern:

  • The effect was strongest among users who had never purchased, and this was the only segment statistically distinguishable from the rest.
  • It fell rapidly with purchase frequency, approaching zero for regular buyers.
  • By recency, the effects were “tightly estimated zeros” for people who had purchased within 0, 30 or 60 days. The effect only returned, and became significant, for users absent more than a year.
  • Advertising did produce a statistically significant increase in new registrations.

The authors summarize the mechanism in one sentence: “search advertising only works if the consumer has no idea that the company has the desired product.”

That is a usable rule. Paid search pays when it informs someone who did not know, and it wastes money when it intercepts someone who did.

Concentration of the paid search advertising effect by customer purchase historyConcentration of the measured paid search advertising effect by customer purchase history, from the heterogeneity analysis of a large field experiment. The researchers interacted the treatment with eleven segments defined by the number of purchases a user had made in the twelve months preceding the experiment. The effect was strongest among users who had never purchased, and that segment was the only one statistically distinguishable from the others. The effect decreased rapidly with purchase frequency, with an estimated slope of negative zero point zero zero three eight and a standard error of zero point zero zero one four, approaching zero for regular buyers. Analyzed by recency rather than frequency, the effects were described as tightly estimated zeros for users who had purchased within zero, thirty or sixty days, and the effect only returned and became statistically significant for users who had been absent for more than a year, with a spline model breaking at ninety days estimating an increase of zero point zero two percentage points per month of absence. Advertising did produce a statistically significant increase in new registrations. The authors summarize the mechanism by stating that search advertising only works if the consumer has no idea that the company has the desired product. The practical rule that follows is that paid search pays when it informs someone who did not know, and wastes money when it intercepts someone who did. The arithmetic explanation given for the negative overall return is that even users purchasing more than fifty times a year arrive through a paid click on about four percent of their purchases, so a large share of the budget is spent on people already on their way.Where the effect livedNever purchasedstrongestThe only segment statisticallydistinguishable from the rest.Bought within 60 dayszero”tightly estimated zeros” at0, 30 and 60 days.Absent over a yearreturnsSignificant again, rising witheach month of absence.The authors’ one-sentence mechanism”search advertising only works if the consumer has no idea that the company has the desired product.”Even users buying more than fifty times a year arrive through a paid click on about 4 percent of purchases.
Strongest among people who had never bought. A tightly estimated zero for anyone who had bought in the last sixty days. Source : Blake, Nosko and Tadelis, Econometrica 83(1), 2015 (2015)

Why your dashboard disagrees, and why more data does not fix it

The same divergence has been measured on social advertising, deliberately and at scale.

Researchers ran 15 randomized advertising experiments on a major social platform, covering 500 million user-experiment observations and 1.6 billion impressions, with randomization at the user level rather than the cookie level. They then discarded the control group and re-estimated the effect using the observational methods advertisers actually use: exact matching, propensity scores, stratification, regression, inverse probability weighting and difference in differences, with four increasing sets of covariates.

Their conclusion: “the point estimates in 7 of the 14 studies with a checkout-conversion outcome are consistently off by more than a factor of three.”

In one study the true lift was 72.8 percent and the naive exposed-versus-unexposed comparison reported 316 percent. In another the true lift was 2.4 percent and the best observational estimate was 1,306 percent.

The follow-up, four years later, is larger and no more encouraging: 663 experiments, 7.9 billion observations, more than 5,000 user-level variables, and the median experiment covering 7.3 million users. Median true lifts by funnel stage were 29, 18 and 5 percent. The best observational methods reported 83, 58 and 24 percent, and a propensity method reported 173, 176 and 64 percent. The authors concluded that “despite having access to large-scale experiments and rich user-level data, we are unable to reliably estimate an ad campaign’s causal effect”, and called it a data problem rather than a model problem.

And the cleanest demonstration of why is older and simpler. In a one-day experiment, 35 million people saw a banner and 1.8 million saw an unrelated public service announcement instead. The true effect on brand searching was plus 5.4 percent. The exposed-versus-unexposed comparison reported plus 1,198 percent, and adding a week of behavioral controls only brought it down to 872 percent.

The reason is that the control group, which saw nothing relevant at all, showed the same spike. People who are browsing actively do more of everything, including searching and buying. The authors named it activity bias and summarized it memorably: “one could erroneously ‘show’ that nearly any browsing behavior is caused by nearly any other browsing behavior.”

Comparison between experimental and observational estimates of the same advertising effectsComparison between experimental and observational estimates of the same advertising effects, from three studies. The first ran fifteen randomized advertising experiments on a major social platform covering five hundred million user experiment observations and one point six billion impressions, with randomization at the user level rather than the cookie level so that cross device tracking was guaranteed and control contamination prevented. The researchers then discarded the control group and re estimated the effect using observational methods advertisers actually use, namely exact matching, propensity scores, stratification, regression, inverse probability weighting and difference in differences, across four increasing sets of covariates. They concluded that the point estimates in seven of the fourteen studies with a checkout conversion outcome are consistently off by more than a factor of three. In one study the true lift was seventy two point eight percent while the naive exposed versus unexposed comparison reported three hundred and sixteen percent, and exact matching on age and sex still reported two hundred and twenty two percent. In another the true lift was two point four percent while the best observational estimate was one thousand three hundred and six percent. The follow up study covered six hundred and sixty three experiments, seven point nine billion observations and more than five thousand user level variables, with a median experiment covering seven point three million users over thirty days, and found median true lifts by funnel stage of twenty nine, eighteen and five percent against observational estimates of eighty three, fifty eight and twenty four percent using double machine learning and one hundred and seventy three, one hundred and seventy six and sixty four percent using stratified propensity score matching, concluding that despite access to large scale experiments and rich user level data the authors were unable to reliably estimate a campaign’s causal effect. The third study ran a one day campaign in which thirty five million people saw a banner while one point eight million saw an unrelated public service announcement instead, finding a true effect on brand searching of plus five point four percent while the exposed versus unexposed comparison reported plus one thousand one hundred and ninety eight percent, and adding a week of behavioral controls only brought that down to eight hundred and seventy two percent. The control group that saw the irrelevant announcement showed the same behavioral spike, a phenomenon the authors named activity bias.Truth against dashboardRandomized estimate against observational estimate, same campaignsRetailer campaign, true lift+73%the same, measured observationally+316%Purchase lift across 663 experiments, median+5%the same, best observational method+24%Banner effect on brand search, true+5.4%the same, exposed vs unexposed+1,198%Why the last one is decisiveThe control group saw an unrelated public service announcement, and showed the same spike.
Three studies, one pattern. The control group that saw an irrelevant ad showed the same behavioral spike. Source : Gordon, Zettelmeyer, Bhargava and Chapsky, Marketing Science 38(2), 2019; Gordon, Moakler and Zettelmeyer, same journal 42(4), 2023; Lewis, Rao and Reiley, WWW 2011 (2023)

Why you probably cannot run this test yourself

The obvious reply is to run your own experiment. The arithmetic makes that harder than it sounds, and someone has published exactly how hard.

A study of 25 large field experiments, each with more than 500,000 unique users and collectively 2.8 million dollars of advertising spend, computed what precision those experiments actually achieved and what would be needed for a decision.

  • To reliably distinguish a highly profitable campaign, at plus 50 percent return, from one that merely breaks even, the median campaign would need to be 9 times larger.
  • To resolve a 10 percentage point difference in return, which is an ordinary threshold for an investment decision, the median campaign would need to be 62 times larger, which the authors call “nearly impossible for a campaign of any realistic size”.
  • Their summary: “informative advertising experiments can easily require more than 10 million person-weeks”.

The reason is variance. Individual purchase behavior has a coefficient of variation around 10. A campaign delivering a 25 percent return had to raise average spending per person by 35 cents, on a variable with a mean of 7 dollars and a standard deviation of 75. The resulting explanatory power is on the order of five millionths.

The median standard error on return on investment across those retail experiments was 26 percent, giving a confidence interval roughly 100 percentage points wide.

So for most companies the honest position is not that the test failed. It is that the test cannot be run at the size available, and that anyone reporting a precise return on advertising for a mid-size business is reporting attribution, not measurement.

What none of this proves about your store

Two honest limits, because they cut in your favor.

The eBay result is about a brand everyone already knows. The authors say so themselves, and the sentence deserves quoting because almost nobody who cites this study includes it: “This may not be true for small and new entities that have no brand recognition.” Their own mechanism predicts the opposite result for an unknown store. Advertising works when it informs, and nobody is searching for your brand name if they have never heard it.

And there is one large randomized experiment on small businesses. A study gave the standard advertising package of a local review platform, free, to 7,209 randomly selected restaurants out of 18,294, for three months. The results were positive and specific: plus 19 percent page views, plus 14 percent direction requests, plus 7 percent calls, plus 7 percent clicks to the restaurant’s own site, plus 5 percent reviews. Effects disappeared immediately when the advertising stopped. And, mirroring the eBay finding exactly, national chains gained less than independents with comparable attributes.

But read that study’s own limits before quoting its return figure. The effects on takeout orders and reservations were not statistically significant, and the authors say they were underpowered on those measures. The 88 percent return figure that circulates comes from an appendix exercise correlating revenue with page views on pre-experiment tax data matched to only 13 percent of the sample, and the authors describe it themselves as “a back-of-the-envelope calculation rather than as an accurate or precise estimate of the returns to advertising.”

So even the best small-business evidence measures attention and intent, not revenue.

Measured effects of a randomized advertising experiment on small restaurantsMeasured effects of a randomized advertising experiment conducted on small restaurants, and the outcomes it did not establish. The experiment covered eighteen thousand two hundred and ninety four restaurants, of which seven thousand two hundred and nine were randomly assigned to receive the standard advertising package of a local review platform free of charge for three months in the autumn of two thousand fifteen, with randomization at restaurant level within four subsamples and a mean treated fraction per postal code of two point three percent to avoid over treating any local market. The measured effects were an increase of nineteen percent in platform page views, fourteen percent in direction requests, seven percent in telephone calls, seven percent in clicks through to the restaurant’s own website, and five percent in customer reviews. Effects disappeared immediately when the advertising stopped. Mirroring a finding from a much larger marketplace experiment, national chains gained less than independent restaurants and local chains with comparable attributes, while gains were larger for better rated establishments with longer review histories and more baseline organic traffic. Paid views converted worse than organic views, with a forty six percent lower conversion rate. Critically, the study did not causally measure revenue: the effects on takeout orders and reservations were not statistically significant and the authors state they were underpowered on those measures, with a minimum detectable effect of about zero point zero four six standard deviations for five thousand eight hundred and sixty seven restaurants and zero point two standard deviations for the three hundred and six restaurants with reservation data. The eighty eight point two percent return on investment figure that circulates comes from an appendix exercise correlating revenue with page views using state tax data from before the experiment, matched to only eight hundred and thirty five restaurants or thirteen percent of the subsample, and the authors describe it themselves as a back of the envelope calculation rather than an accurate or precise estimate of the returns to advertising.7,209 restaurants, randomly given free advertisingPlatform page views+19%Direction requests+14%Telephone calls+7%Customer reviews+5%Takeout orders and reservationsnot significantTwo findings that mirror the marketplace experimentNational chains gained less than independents. Effects stopped the moment the advertising stopped.The return figure people quoteAn appendix estimate on pre-experiment tax data matched to 13 percent of the sample, which the authorscall “a back-of-the-envelope calculation”.
Real effects on attention and intent. No significant effect on orders, and the return figure is an appendix estimate. Source : Dai, Kim and Luca, Frontiers: Which Firms Gain from Digital Advertising?, Marketing Science 42(3), 2023; NBER Working Paper 30925 (2023)

What to do with a store-sized budget

Separate brand queries from everything else, and treat them as a different question. For an established name, brand ads mostly buy clicks that organic search would deliver. For a name nobody knows, there is nothing to cannibalize.

Judge non-brand search on new customers, not on total conversions. The experiment says the effect concentrates on people who have never bought and on people absent more than a year. Your reports probably mix those in with repeat buyers who arrived through a paid click out of habit. Separating the two in the reporting before a non-brand budget goes live is the groundwork of our B2B paid acquisition work, since the distinction cannot be recovered afterwards.

Read your attribution windows before arguing about the numbers. The default click window on one platform is 30 days. The view-through window on the other is one day and is triggered by an impression, not by anyone watching anything. Neither number is wrong; they are simply not the same measurement.

Do the one test that is affordable. Full incrementality measurement is out of reach at store scale, but switching one campaign off in some geographies and not others, for long enough, is not. Google publishes the method, and its own worked example is instructive: in one case each dollar bought one third of an incremental click, an incremental cost per click of three dollars against a reported cost of two dollars forty, understating the true cost by 20 percent.

And expect the honest answer to be uncomfortable. No public institution publishes a cost per click or a cost per acquisition for any industry, and neither platform publishes benchmarks by sector. Every average you have seen came from an agency describing the accounts it happens to manage. Your own before-and-after, with a geography held back, is worth more than all of them.