· 5 min read
Should you test one message properly or five badly?
Splitting a small budget across five variants feels thorough and answers nothing. The arithmetic of how many messages a budget can genuinely separate at all.
Asking whether to run one message or many sounds like a matter of taste, and it is not — it is division. A budget buys a fixed number of clicks, splitting it splits the clicks, and below a certain number per variant nothing you observe can be distinguished from chance. The number of messages a budget can genuinely separate is usually one, occasionally two, and almost never the five that a test-everything instinct reaches for. (Disclosure: published by BidSurvivor, which sells slots you could use to test messages, including for nothing.)
The division that decides it
Take $300 and a channel at $1 a click. That is 300 clicks.
- One message: 300 clicks on one line.
- Three messages: 100 each.
- Five messages: 60 each.
Now apply the rule of thumb: a rate needs roughly 15 events before it is worth quoting, and rather more before it is worth comparing to another rate. At a 3% conversion rate, 60 clicks produce 1.8 conversions on average. Two of your five variants will show zero and one will show four, and none of that is about the copy.
At 300 clicks on one line you get 9 conversions. Still not many — but it is enough to distinguish a page converting at 1% from one converting at 5%, which is the difference that actually changes what you do next.
So the five-way test does not give you five answers. It gives you five numbers, none of which is an answer, and a strong feeling that the one at the top is the winner. That feeling is the expensive part.
What the five-way test is really measuring
Worth naming the failure precisely: with 60 clicks per variant you are measuring which variant got lucky, and luck is not repeatable. Run the identical five lines again next week and a different one will lead. People who do this routinely discover that their "winner" does not replicate, and usually conclude that something changed in the market.
Nothing changed in the market. The first result was noise and so is the second.
This is the same mechanism as reading a conversion rate off too few clicks, except worse, because comparing two noisy numbers requires far more data than measuring one — Evan Miller's sample size calculator will price that difference for your own figures in about a minute.
When many is right
There is a real case for testing several at once, and it is not the one people think.
When the effects are enormous. If one line gets 40 clicks per thousand impressions and another gets 3, you do not need a calculator. Small budgets should be hunting for differences of that size — a different audience, a different offer, a different framing — not 15% improvements in wording.
When you are testing higher up the funnel. Click-through rate has roughly thirty times the event volume of a 3% conversion rate, so it reaches a usable sample thirty times faster. You can honestly compare four ad creatives on click-through with a few hundred impressions each. You cannot compare them on sign-ups. Test the thing you have the data for, and label it as the proxy it is.
When it costs nothing. This is the case worth building around, and it is where our own board sits. BidSurvivor runs two twelve-hour slots a day; the first brand into an empty one pays $0, and every visitor starts with $100 of house credit, so contested slots (from a $1.00 floor) cost nothing out of pocket at current prices. Different slots on different days can carry different lines, and each slot's card opens and click-throughs are published separately.
The honest limit: the volumes are small, so this eliminates rather than crowns. A line that gets nothing cold, twice, is not a line worth putting a budget behind — that is a real finding and it is free. A line that does slightly better than another is not distinguishable at these numbers, and the board will not pretend otherwise.
The order that wastes least
- Write more lines than you will test. Eight, deliberately. The eighth is often the good one, because by then the obvious framings are used up. One-line pitch examples has the patterns.
- Eliminate for free. Five-second tests on strangers, then free slots. Expect to lose most of them, and let the failures be unanimous rather than marginal.
- Get to two. Not one — two, so the paid test has something to compare against.
- Spend the whole budget on those two, one channel, enough clicks that each side clears the sample bar. If the budget cannot cover two, it covers one, and the honest move is to run one properly rather than two badly.
- Change one thing between rounds. Otherwise a difference tells you something moved without telling you what.
The uncomfortable version
Most small advertisers would learn more from spending their entire budget on a single message than from any split they are likely to design. That is an unsatisfying recommendation because it feels like less testing, and it is actually more: one readable result beats five unreadable ones, and five unreadable ones are what a split budget buys at this scale.
If the budget is small enough that even one message cannot clear the bar — and $200 usually cannot — then the answer is not to split it further. It is to stop buying clicks and go get the answer somewhere cheaper first.