Metrics and analysis
Creative testing that actually reaches a conclusion
Most creative tests end with no answer because they change two things at once or stop too early. How to design a test that genuinely decides something.
Also available in: Português · Español
Most of the creative testing that runs in paid media accounts produces no knowledge. It produces a feeling that something was tested, and a winner picked on a number that does not survive a second look.
The problem is rarely a lack of statistical rigor. It is four design mistakes that keep repeating.
The four mistakes
Changing two things at once. Variation B has a different image and a different headline. If it wins, you do not know what won — and you cannot repeat it.
Stopping on day three. Three days of data with 11 conversions decide nothing. People tend to stop when the result looks good, which is exactly when it is still luck.
Comparing by CTR. CTR measures attention. You are paying for sales. Those two frequently move in opposite directions.
Testing things that change nothing. Button color, font, frame. Even if one wins, the learning does not transfer to the next piece.
The design that works
One variable at a time, and make it big
Test the angle, not the finish. The angle is what the piece promises and who it talks to:
| variable | worth testing? |
|---|---|
| Angle (pain, proof, price, objection) | yes, it moves the most |
| Format (video, static, carousel) | yes |
| First second of the video | yes, it decides retention |
| The offer in the piece | yes |
| Color, font, frame | rarely |
| Logo placement | no |
If you need a magnifying glass to see the difference between variations, the test was born underpowered.
Everything equal except the variable
Same ad set, same audience, same budget, same time window. Separate campaigns look cleaner and are worse: each enters the auction under its own conditions, and you end up comparing auctions instead of creatives.
Volume before deadline
Do not stop by calendar, stop by volume. The question is: how many more conversions would the losing variation need to tie?
If the answer is "three", there is no conclusion. If it is "thirty", there is.
A practical rule for anyone who does not want to compute significance: at least 25 to 30 conversions per variation, and a gap of at least 20%. Below that, declare a draw and move on.
Seven-day cycles
Tuesday is not Saturday. Closing a test in multiples of seven days removes day-of-week distortion, which is larger than almost anyone assumes.
The deciding metric
Decide by the metric closest to the money that has enough volume within the test period.
| if you have | decide by |
|---|---|
| 30+ sales per variation | cost per sale |
| 30+ qualified leads | cost per qualified lead |
| only lead volume | CPL, knowing it is a weak proxy |
| none of those | the test is too short, extend it |
Never decide by CTR when any conversion metric is available. The price creative almost always has lower CTR and higher quality; declaring a winner on CTR is systematically choosing the worse one.
Write it down, or the test did not happen
An unrecorded test becomes a repeated test six months later, run by whoever arrived after you.
Four lines are enough:
TEST 14 Oct → 28 Oct
Hypothesis video testimonial beats static price on cost per sale
Variations A: static price · B: 20s video testimonial
Result A $740/sale (22 sales) · B $610/sale (31 sales)
Conclusion B wins. Becomes the default. Next: test a 10s testimonial.
The last line is the most valuable: every good test generates the next question.
When testing is not worth it
Volume too low. With 8 sales a month across the whole account, no creative test will conclude in useful time. Test offer and audience, which have bigger effects, or accumulate three months per variation.
A structural change mid-test. New budget, new optimization event, new landing page: any of them invalidates the comparison. Finish the test first, or accept that it died.
Pressure for immediate results. Testing costs money on the loser. If the account cannot afford to lose anything this month, it is not a month for testing — it is a month for running what works and planning the test for next month.
The summary in four rules
- One variable per test, and big enough to matter.
- Same ad set, same audience, same budget.
- Stop by volume, in seven-day cycles — never because the number looked good.
- Decide by the metric closest to the money that has volume.
One test that concludes per month is worth more than four that end in "B seems to have been better".
To decide whether the current piece is already tired before building the test, see creative fatigue. For the budget-per-piece criterion, see revenue per creative.
Frequently asked questions
How many creatives should I test at once?
Three to four variations in the same ad set. More than that splits delivery and none accumulates enough volume to be judged before the month ends.
How long should a test run?
Until each variation accumulates enough conversions for the difference not to be luck, or seven full days — whichever comes later. Weekly cycles avoid day-of-week effects.
Can I test in separate campaigns?
Prefer the same ad set. Separate campaigns with their own budgets introduce auction and audience differences that contaminate the comparison.
How do I know the difference is real?
If one variation needs only a few more conversions to tie with the other, there is no conclusion. Differences under 20% on small volume are almost always noise.