Paid or Organic for an Online Store: The Experiments
A retailer switched off its brand-keyword ads and organic search caught 99.5 percent of the clicks. Its measured return on the rest was negative 63 percent.
In 2012 a large online marketplace stopped bidding on searches containing its own name, and lost 0.5 percent of its clicks. Not 0.5 percent of paid clicks. Half a percent of total clicks, because organic results caught almost everything the ads had been catching.
The published finding is worded precisely: “almost all (99.5 percent) of the forgone click traffic from turning off brand keyword paid search was immediately captured by natural search traffic”. And the authors note that 99.5 percent is a lower bound on retention, since some of the remainder probably shifted to people typing the address directly.
That is the single most useful experiment ever run on this question, and its second half is harder still. Everything below comes from randomized or quasi-randomized field experiments, not from platform reporting, and the gap between the two is the subject.
The brand keyword result, and how it was measured
The design matters, because it is what makes the number trustworthy.
The company stopped all bidding on queries containing its brand name on two search engines while continuing to bid on a third, which served as a control. That allows a difference-in-differences estimate rather than a before-and-after comparison. Raw traffic from the treated engine fell 5.6 percent. After controlling for seasonality using the untreated engine, the actual loss was 0.529 percent of clicks, or roughly 1.5 percent of paid clicks alone.
A replication on the third engine four months later showed a 3.2 percent fall, and the authors flag the weakness honestly: there was no valid control during that test, and they expect a proper control would reduce the estimate, as it did in the first case. A separate German test with a control confirmed the pattern.
The mechanism is not mysterious. Someone searching for your brand name has already decided to come to you. The advertisement is buying a click you were going to receive anyway, and the only thing it changes is which line of the results page they click.
The non-brand experiment, and the number that should be famous
The main test was larger. The company bid on more than 100 million keywords, and the researchers switched advertising off in 68 of the 210 United States market areas for 60 days, against 142 controls. Attributed sales in the treated areas fell by more than 72 percent, confirming the manipulation worked.
Then they estimated the return on investment five different ways on the same data. The results are worth setting out in full, because the spread is the lesson:
Method
Estimated return
Naive regression
+4,173%
Regression with market and date fixed effects
+1,632%
Instrumental variables
-22%
Instrumental variables, second specification
-63%
Difference in differences
-63%
The experimental estimate is negative 63 percent, with a confidence interval from negative 124 to negative 3, so it is statistically distinguishable from zero. The naive regression on the identical data reports a return of over four thousand percent.
That gap is not a curiosity. It is the difference between what an advertising dashboard shows and what the spending actually did.
The explanation the authors give is arithmetic. Even users buying more than fifty times a year arrive through a paid click on about 4 percent of their purchases. So a large share of the budget was being spent to intercept people who were already on their way.
The heterogeneity analysis is the part that makes the result usable rather than merely deflating.
The researchers split users into eleven segments by how many purchases they had made in the previous twelve months. The pattern:
The effect was strongest among users who had never purchased, and this was the only segment statistically distinguishable from the rest.
It fell rapidly with purchase frequency, approaching zero for regular buyers.
By recency, the effects were “tightly estimated zeros” for people who had purchased within 0, 30 or 60 days. The effect only returned, and became significant, for users absent more than a year.
Advertising did produce a statistically significant increase in new registrations.
The authors summarize the mechanism in one sentence: “search advertising only works if the consumer has no idea that the company has the desired product.”
That is a usable rule. Paid search pays when it informs someone who did not know, and it wastes money when it intercepts someone who did.
Why your dashboard disagrees, and why more data does not fix it
The same divergence has been measured on social advertising, deliberately and at scale.
Researchers ran 15 randomized advertising experiments on a major social platform, covering 500 million user-experiment observations and 1.6 billion impressions, with randomization at the user level rather than the cookie level. They then discarded the control group and re-estimated the effect using the observational methods advertisers actually use: exact matching, propensity scores, stratification, regression, inverse probability weighting and difference in differences, with four increasing sets of covariates.
Their conclusion: “the point estimates in 7 of the 14 studies with a checkout-conversion outcome are consistently off by more than a factor of three.”
In one study the true lift was 72.8 percent and the naive exposed-versus-unexposed comparison reported 316 percent. In another the true lift was 2.4 percent and the best observational estimate was 1,306 percent.
The follow-up, four years later, is larger and no more encouraging: 663 experiments, 7.9 billion observations, more than 5,000 user-level variables, and the median experiment covering 7.3 million users. Median true lifts by funnel stage were 29, 18 and 5 percent. The best observational methods reported 83, 58 and 24 percent, and a propensity method reported 173, 176 and 64 percent. The authors concluded that “despite having access to large-scale experiments and rich user-level data, we are unable to reliably estimate an ad campaign’s causal effect”, and called it a data problem rather than a model problem.
And the cleanest demonstration of why is older and simpler. In a one-day experiment, 35 million people saw a banner and 1.8 million saw an unrelated public service announcement instead. The true effect on brand searching was plus 5.4 percent. The exposed-versus-unexposed comparison reported plus 1,198 percent, and adding a week of behavioral controls only brought it down to 872 percent.
The reason is that the control group, which saw nothing relevant at all, showed the same spike. People who are browsing actively do more of everything, including searching and buying. The authors named it activity bias and summarized it memorably: “one could erroneously ‘show’ that nearly any browsing behavior is caused by nearly any other browsing behavior.”
The obvious reply is to run your own experiment. The arithmetic makes that harder than it sounds, and someone has published exactly how hard.
A study of 25 large field experiments, each with more than 500,000 unique users and collectively 2.8 million dollars of advertising spend, computed what precision those experiments actually achieved and what would be needed for a decision.
To reliably distinguish a highly profitable campaign, at plus 50 percent return, from one that merely breaks even, the median campaign would need to be 9 times larger.
To resolve a 10 percentage point difference in return, which is an ordinary threshold for an investment decision, the median campaign would need to be 62 times larger, which the authors call “nearly impossible for a campaign of any realistic size”.
Their summary: “informative advertising experiments can easily require more than 10 million person-weeks”.
The reason is variance. Individual purchase behavior has a coefficient of variation around 10. A campaign delivering a 25 percent return had to raise average spending per person by 35 cents, on a variable with a mean of 7 dollars and a standard deviation of 75. The resulting explanatory power is on the order of five millionths.
The median standard error on return on investment across those retail experiments was 26 percent, giving a confidence interval roughly 100 percentage points wide.
So for most companies the honest position is not that the test failed. It is that the test cannot be run at the size available, and that anyone reporting a precise return on advertising for a mid-size business is reporting attribution, not measurement.
What none of this proves about your store
Two honest limits, because they cut in your favor.
The eBay result is about a brand everyone already knows. The authors say so themselves, and the sentence deserves quoting because almost nobody who cites this study includes it: “This may not be true for small and new entities that have no brand recognition.” Their own mechanism predicts the opposite result for an unknown store. Advertising works when it informs, and nobody is searching for your brand name if they have never heard it.
And there is one large randomized experiment on small businesses. A study gave the standard advertising package of a local review platform, free, to 7,209 randomly selected restaurants out of 18,294, for three months. The results were positive and specific: plus 19 percent page views, plus 14 percent direction requests, plus 7 percent calls, plus 7 percent clicks to the restaurant’s own site, plus 5 percent reviews. Effects disappeared immediately when the advertising stopped. And, mirroring the eBay finding exactly, national chains gained less than independents with comparable attributes.
But read that study’s own limits before quoting its return figure. The effects on takeout orders and reservations were not statistically significant, and the authors say they were underpowered on those measures. The 88 percent return figure that circulates comes from an appendix exercise correlating revenue with page views on pre-experiment tax data matched to only 13 percent of the sample, and the authors describe it themselves as “a back-of-the-envelope calculation rather than as an accurate or precise estimate of the returns to advertising.”
So even the best small-business evidence measures attention and intent, not revenue.
Separate brand queries from everything else, and treat them as a different question. For an established name, brand ads mostly buy clicks that organic search would deliver. For a name nobody knows, there is nothing to cannibalize.
Judge non-brand search on new customers, not on total conversions. The experiment says the effect concentrates on people who have never bought and on people absent more than a year. Your reports probably mix those in with repeat buyers who arrived through a paid click out of habit. Separating the two in the reporting before a non-brand budget goes live is the groundwork of our B2B paid acquisition work, since the distinction cannot be recovered afterwards.
Read your attribution windows before arguing about the numbers. The default click window on one platform is 30 days. The view-through window on the other is one day and is triggered by an impression, not by anyone watching anything. Neither number is wrong; they are simply not the same measurement.
Do the one test that is affordable. Full incrementality measurement is out of reach at store scale, but switching one campaign off in some geographies and not others, for long enough, is not. Google publishes the method, and its own worked example is instructive: in one case each dollar bought one third of an incremental click, an incremental cost per click of three dollars against a reported cost of two dollars forty, understating the true cost by 20 percent.
And expect the honest answer to be uncomfortable. No public institution publishes a cost per click or a cost per acquisition for any industry, and neither platform publishes benchmarks by sector. Every average you have seen came from an agency describing the accounts it happens to manage. Your own before-and-after, with a geography held back, is worth more than all of them.
Frequently asked questions
Should I bid on my own brand name?
The one large experiment on this found that switching brand-keyword ads off cost 0.5 percent of total clicks, because 99.5 percent of the traffic came back through organic results. That was a marketplace with enormous brand recognition, and the authors say the result may not hold for small or new businesses.
Why does my ad dashboard show a good return when the experiments show a bad one?
Because platforms report attributed conversions, not incremental ones. On the same eBay data, a naive regression gave a return of positive 4,173 percent while the experimental estimate was negative 63 percent.
Are observational methods good enough if I have enough data?
A study of 663 experiments with 7.9 billion observations and over 5,000 user-level variables concluded the authors were unable to reliably estimate a campaign's causal effect, calling it a data problem rather than a model problem.
Can I run an incrementality test myself?
Probably not on a purchase outcome. A study of 25 large experiments found that the median campaign would need to be 62 times larger to distinguish a 10-point difference in return on investment, and that informative experiments can require more than 10 million person-weeks.
Does this mean advertising does not work for a small store?
No, and the eBay authors say so explicitly: their result may not apply to small or new entities with no brand recognition. Their own mechanism predicts advertising works when it informs, which is exactly the position an unknown store is in.