Attribution said the ads returned 4,173%. The experiment on the same spend said minus 63%.
eBay ran the experiment, published it, and the gap between the dashboard and the truth was about forty times. Then I sized a holdout for a normal advertiser and found most of them cannot run one.
- On the same eBay spend, the naive regression returned an ROI of 4,173% and the experiment returned minus 63%.
- Brand keyword ads had no measurable short term benefit at all, which is the cleanest result in the paper.
- The whole paid search regime moved eBay sales by 0.66%, with an interval running from minus 0.42% to plus 1.74%.
- A geo holdout across 210 regions and four weeks each side detects a 3.9% lift. Across 20 regions it detects 12.8%.
- Most advertisers cannot detect their own paid search effect, which is a different problem from not having one.
A client wanted to prove that their paid search was working. Sensible request. I gave them the answer I have given a hundred times, which is that the only honest way to know is to switch the channel off in some regions and watch what happens, and which is also the advice that everybody in measurement repeats and almost nobody bothers to size before repeating it.
So I sized it. They operate in about twenty regions. The smallest revenue change that test could detect was 12.8%.
Their paid search is roughly 6% of revenue. So even if every penny of it was incremental, and it almost certainly is not, the test would have come back flat and I would have had to stand in a room and explain what flat meant.
Here is my advice, up front, so nobody has to reach the end for it. You should size a holdout before you promise anyone an answer. Always compare the smallest detectable change against the share of revenue the channel could possibly be worth. Do not run the test when those two numbers do not overlap. Make sure the minimum detectable effect goes on the readout in the same size type as the result. And never report last click return as if it were a return, because there is a published experiment showing exactly how far apart those two numbers get, and the answer is roughly forty times.
The forty times gap
In 2015 Blake, Nosko and Tadelis published a series of field experiments run inside eBay, in Econometrica. It is the cleanest thing anyone has put in public on this question, and it is worth reading rather than citing, which I say as somebody who cited it for years before reading it.
They took the same spend and measured its return two ways, once with the method your reporting stack uses every morning and once with an experiment that actually removed the ads from a random set of regions, which is why this paper has survived eleven years of people wishing it would go away. That is the whole design. It is also why it is unanswerable.
The naive way first. Regress sales on ad spend, the way every dashboard on the market quietly does. Their words: a “ROI of over 4,100% without time and geographic controls”. The table in the paper puts it at 4,173%. Then they added time and geographic controls, which is the sort of careful thing I would have done and would have felt quite pleased about, and the figure came down to 1,632%. Still enormous. Still wrong.
Then the experiment. Ads off in a random set of regions, difference in differences against the rest. I want to be precise here because the precision is the point. The experimental ROI was minus 63%, with a 95% confidence interval running from minus 124% to minus 3%.
Same company. Same spend. Same period. The dashboard said the channel returned forty one times its cost, the experiment said it lost money, and the confidence interval around the experimental figure did not include zero, which means this was not a case of the truth sitting somewhere comfortably in the middle. I have shown that pair of numbers to a lot of people. Nobody has ever guessed the size of the gap correctly.
Why the gap is that size
The mechanism is not complicated. It is also not fraud. Nobody at eBay was lying to anybody.
Clicks and purchase intent are correlated. The people most likely to click are the people already looking for you. They would have arrived anyway. The organic result sits directly underneath the ad. So the ad collects the credit, and no analytics tool on the market can tell those two journeys apart, because from the outside they look identical.
Brand terms are the extreme case. The paper says so in those words. Brand keyword ads had “no measurable short-term benefits”. Somebody types your name into Google. You pay to appear above the free listing of yourself. The dashboard records a conversion. Everyone goes home happy, and the only way anybody would ever find out that the money bought nothing is if somebody switched the campaign off for long enough to watch the organic listing quietly absorb every one of those visits, which is precisely the experiment almost nobody is willing to run on their largest line item.
For non brand terms the picture was more interesting. New and infrequent users were positively influenced. Frequent users were not, and frequent users absorbed most of the money. So the average came out negative while a real effect sat inside it, hidden behind the people who were buying regardless.
That last part is the bit most summaries drop, including the summary I used to give. Paid search was not useless at eBay. It was mispriced, aimed at the wrong people, and reported through a method that could not see any of that.
There is one more number I keep. Across the whole experiment, the entire regime of paid search added 0.66% to sales, with a 95% interval running from minus 0.42% to plus 1.74%.
That is eBay. Fifty one million dollars of spend. Two hundred and ten designated market areas, with ads suspended in roughly 30% of them. Four months of data. The honest answer they could publish was a band about two points wide, sitting across zero. I think about that band whenever somebody asks me for a single number.
Now size your own holdout
Here is where I stopped feeling smug about attribution and started worrying about the alternative.
A geo holdout is a two sample comparison. For each region you take the log ratio of average daily revenue after the switch to average daily revenue before it. Then you compare that ratio between the switched off regions and the rest. Your resolution depends on one thing only, which is how noisy a single region is from day to day.
I ran it at a daily revenue variation of 0.35, which is an ordinary mid sized retailer, and a 30% holdout to match eBay’s design.
| Regions | 14 days each side | 28 days each side | 56 days each side |
|---|---|---|---|
| 20 | 18.1% | 12.8% | 9.0% |
| 50 | 11.4% | 8.1% | 5.7% |
| 100 | 8.1% | 5.7% | 4.0% |
| 210 | 5.6% | 3.9% | 2.8% |
Read the bottom right cell first. Two hundred and ten regions, eight weeks of data, and you can see a 2.8% change in total revenue. Now read the top left. Twenty regions and two weeks. Nothing under 18% exists for you there.
Most advertisers are in the top half of that table. Most paid search programmes are worth somewhere between two and eight percent of revenue. Those two facts do not overlap.
Method, so you can rerun it
Estimator is the difference in log ratios across regions, tested two sample, alpha 0.05 and power 0.80. The minimum detectable effect is (z of alpha over two plus z of beta) times the standard deviation of the log ratio times the root of one over the treated count plus one over the control count. The standard deviation of the log ratio is the daily coefficient of variation divided by the root of the days, doubled for two periods. Holdout share 30%. I checked the formula against 4,000 simulated experiments with lognormal region sizes and lognormal daily noise, seed 20260813. At 210 regions and 28 days each side the formula says 3.95% and the simulation detected that lift in 76.8% of runs, which is close enough to the 80% target for me to trust the table. At a quieter business, coefficient of variation 0.20, twenty regions gets you to 7.3%. At a noisier one, 0.50, it is 18.3%. I would rather publish the whole sensitivity than pick the flattering column.
What a flat result actually means
The word flat is doing an enormous amount of quiet damage in this industry, so it is worth pulling apart.
When a holdout comes back flat, three completely different things could be true. The channel might have no effect. The channel might have an effect that is smaller than your test could possibly see. Or the test was broken, and a broken test also produces flat.
Those three deserve different decisions. They arrive looking identical.
So I now write the minimum detectable effect into the readout, in the same size type as the result itself. Not in an appendix nobody opens. If the finding is no significant change, then it goes on the page as “no change larger than 12.8% detected”, which is a sentence people argue with, and the arguing is the entire point of writing it that way.
Somebody on a client team read one of those and said “so we learned nothing”. He was right, and I told him so. It was also the most useful meeting I had that quarter, because the very next question in the room was how many regions it would take to learn something, and unlike most questions asked in that room it has an actual answer with a number in it.
The asymmetry matters too. A flat result on branded search is close to confirmation, because the published expectation is already no effect and you are checking rather than exploring. A flat result on a channel nobody has ever measured tells you almost nothing. Same word on the slide. Two completely different amounts of knowledge sitting behind it.
The awkward middle
So attribution is wrong by roughly forty times, and the honest alternative needs more regions than most companies have. That is a genuinely uncomfortable place to end up, and I have watched it land badly in enough rooms now to know that the discomfort is the useful part rather than the problem, because the alternative is a number everybody trusts and nobody has checked. A founder at a home goods brand told me she would rather have no figure than a wrong one. She is the only person who has ever said that to me.
A colleague put it better than I did. He said “so the good measurement is unaffordable and the affordable measurement is fiction”, and then we both sat there for a while.
An aside with nothing to do with any of this. The eBay paper carries a footnote saying they kept buying brand ads nationally until roughly six weeks into the geographic test, and only then halted them everywhere. Six weeks of a live experiment went by before the obvious thing got switched off, in a building full of economists, with academics watching and a very large budget riding on the answer. I find that oddly comforting every time I read it. Anyway. Back to it.
What I actually recommend now
Switch off branded search and watch. This is the one place where the expected answer is known, the spend is usually large, and the downside is small. eBay found no measurable short term benefit. You will probably find the same. If you find otherwise, that is a real finding and worth having.
Size before you promise. Put your region count and your daily variation into the table above before anyone books a meeting about results. If the smallest detectable change is bigger than the channel could plausibly be worth, say so on day one rather than in week nine.
Pool your regions upward. Twenty regions is a hard place to work. Weekly aggregation, longer windows and matched pairing help, and none of them turn twenty regions into two hundred.
Stop calling last click ROI a return. Call it attributed revenue. Put it next to spend, without dividing one number by the other. The division is the step that smuggles in the claim of causation, and the eBay figure is a reasonable estimate of how wrong that claim gets when nobody checks it. I have started deleting the ratio column from reports before anyone sees them. It annoys people for about a week.
Budget the measurement like the media. If a channel takes a seven figure budget, the experiment that checks it is not overhead, it is the only thing standing between you and a 4,173% number that somebody is about to put in a board deck.
What I still do not know
The 0.35 daily variation is the weakest number in this article. I picked it from the retailers I have data for, which is a small and unrepresentative sample, and the whole table moves with it. At 0.20 the picture is much friendlier and at 0.50 it is hopeless. Anyone who publishes a distribution of that figure across real businesses would be doing the industry a favour, and I have not found one.
The other thing that still bothers me is the age of the eBay result. It is from 2012 and it is still the best public experiment of its kind, which tells you something about how badly the companies with the data want this question answered. I keep waiting for somebody large to publish a replication. Fourteen years is a long time to wait for a second data point.
Questions people ask about this
What is incrementality, in one sentence?
The sales you would lose if you switched the channel off, which is not the same thing as the sales the channel gets credit for.
Why is attribution wrong by so much rather than a little?
Because clicks and purchase intent are correlated. People who were already going to buy are the most likely to click, so the channel collects credit for demand it did not create, and the error compounds on brand terms where intent is highest.
Does this mean paid search never works?
No. eBay found new and infrequent users were positively influenced by non brand ads. The problem was that frequent users, who were buying anyway, absorbed most of the spend, so the average return came out negative.
How many geos do I need for a holdout?
In my calculation, 210 geos and four weeks on each side detects a 3.9% change in total revenue at 95% confidence and 80% power. Fifty geos detects 8.1%. Twenty geos detects 12.8%. If your true effect is smaller than that, the test will come back flat whether or not the channel works.
What should I do if I cannot run a holdout?
Run the switch off test on the one thing where the answer is known and the cost of being wrong is low, which is branded search. Then stop reporting last click ROI as if it were a return.
Size the test first. The tool on our homepage gives you the sample per variant and how long your traffic needs to produce it, and it tells you when the answer is that the test cannot work.
Open the sample size tool → See what we do