Advertising Waste Diagnostic · System 03 · the one test that settles the last bucket
Where the Next Dollar Goes
System 02 leaves one bucket deliberately unsized: spend that may be buying customers who were already on their way. No report settles it, because every report is built from the touchpoints the platform saw. The only way to find out is to stop advertising to some of them and watch what your own records do.
The idea in one line: attribution, the platform’s rule for handing out credit, asks whether one of our ads came before the sale. A holdout, where you stop showing ads to one comparable group and keep showing them to another, asks whether the people who saw the ads bought more than the people who did not. The second question is the one a budget decision rests on, and it is the one measurement no platform can hand you, because answering it means withholding their product from someone.
The difference between credit and cause
Both can be correct at the same time. They are answers to different questions, and mistaking one for the other is what makes a budget move feel safe when it is a coin flip.
Take the illustration in Meta’s own documentation, which confuses most people the first time because the numbers look like they cannot both be true. A campaign reports 1,000 purchases: the people who clicked an ad and bought within seven days. Now run the same campaign as a holdout. Everyone eligible to see the ads is split into two groups. In the group that saw them, 10,000 people bought during the test, through every route: they clicked an ad, or searched for the brand, or walked in, or were going to buy anyway. In the group the ads were withheld from, scaled to the same size, 9,700 bought. The ads caused the difference: about 300 purchases. So of the 1,000 the report credited to the campaign, roughly 700 were people who clicked an ad on their way to a purchase they would have made regardless. Nothing was measured wrong. The first number counts purchases that followed an ad inside a chosen window; the second counts the purchases that would not have happened without it, which is what the trade calls incremental.
Meta warns against comparing the two totals directly, and the warning generalizes. Once you have seen the gap, the practical rule follows: a platform’s conversion count is the right input for steering a campaign day to day, and the wrong input for deciding whether the channel deserves more money. Those are separate decisions made at separate cadences, and they need separate numbers.
What attribution answers
- Which of my ads should get credit for this, under my window and my rules? Useful for bidding, creative, and daily pacing, where the alternative is flying blind.
What a holdout answers
- If this spend stopped, what would the business lose? The only question a budget increase or a channel cut turns on.
What the platforms will run for you
Both platforms offer to run a holdout for you, under the name lift. Access is the part worth knowing before you plan around one.
Which leaves most accounts below the line, and that is a fine place to be. The test that follows costs nothing except the discipline of leaving it alone while it runs.
How to run one on a small budget
Turn the spend off somewhere, leave it on somewhere comparable, and read the difference in your own records. These decisions make it a test rather than an anecdote.
Reading the result
Four results, and the third and fourth are the ones that get misread.
Outcomes hold where the ads stopped
- The spend was doing less than the dashboard claimed. Before moving it, check where the demand went: if branded search, organic, or direct traffic rose to meet it, the channel was harvesting demand that already existed. That is worth knowing precisely, because the same demand may still need catching somewhere cheaper.
Outcomes fall about as predicted
- The spend is doing work. Turn it back on, write down what it is worth, and take the budget question elsewhere. This result also earns the platform’s conversion count a little more trust for that slice.
Outcomes fall harder than predicted
- You were under-crediting the channel, which happens most often where the effect shows up in a place nobody was watching: walk-ins, repeat customers, referrals, the phone. Widen the outcome measure and keep the spend.
Too noisy to say
- The most common result at a small budget, and the one that gets quietly rounded into “no effect.” No answer is a different finding from no effect. Record the volume the test would have needed, and leave the spend where it was until something changes.
Then close the loop, because a result that never gets checked again is a story. A finding becomes knowledge when you can write all five of these down: the leakage you suspected, the change you made, the business outcome you expected, the change you observed, and whether the saving was still there a quarter later.
That last one catches the two mistakes that look like wins. Pausing a branded campaign can make reported return fall while revenue barely moves, which reads as a loss and is closer to a saving. Expanding a non-brand campaign can lift reported conversions without adding a single profitable customer, which reads as growth and is closer to a leak. Neither is visible from inside the dashboard, and both show up plainly in a record of what the business took in.
Deciding where the next dollar goes
In order. Each step changes the evidence the next one runs on, which is why doing them out of order wastes the work.
Underneath all of it sits a three-layer habit worth keeping permanently, long after the audit is done.
The matching, the arithmetic, and the honest reading of a noisy result are all work an assistant can carry, on the same condition as before:
These five messages span the whole test, so it takes more than one sitting. Open ChatGPT, Claude, or whatever you use, attach your own outcomes by market, and send 1 to 3 before you switch anything off. Run the test. Weeks later, come back to the same chat and send 4 and 5 with the results, so the prediction you wrote at the start is still in front of it.
- “Here is the slice I suspect and the outcomes my records produced by market for the last six months. Propose a holdout: which markets to hold out, which to keep as control, why those are comparable, and how long to run it.”
- “Work out how many outcomes I need in each group to detect a difference worth acting on, using roughly 16 divided by the square of the fractional difference as the count per group. If my volume cannot support the test, say so and tell me what would.”
- “Write the prediction as one sentence I can be wrong about, before the test starts.”
- “Here are the results. Tell me whether the difference is bigger than the normal week-to-week variation in this data, and whether anything else changed in the period that could explain it.”
- “If the answer is that the test could not settle it, say that plainly instead of picking the more interesting reading.”
That completes the method: reconcile what the platform counts against what your records show, sort what is left into four kinds of waste, and settle the last kind with one controlled test. The complete system assembles all three into something you can run start to finish, with the evidence checklist, the finding format, and a master prompt you paste into any assistant along with your own exports.
