How Many Creatives to Test: A Platform Withdrew Its Number
Six or fewer per ad set was official guidance in a February 2025 capture. It is gone from the live page. What remains is a ceiling, a duration, and no budget.
Six or fewer creatives per ad set was official guidance, with a stated reason: beyond six, little marginal benefit. It was on the platform’s own page in an archived capture from February 2025. It is not on the live page now.
The current version advises decreasing ads per ad set while maintaining diverse creative assets, and mentions that one ad can contain up to ten creative assets. No number for ads. That change of direction is the most informative thing available on this question.
This page gives what each platform publishes today, what was removed, and how to work out your own number from volume you can actually measure.
The number that was published and then withdrawn
Worth recording carefully, because it will not be in the documentation much longer for anyone to check.
What the archived page said, verbatim. Use 6 or fewer creatives per ad set. The delivery system favors ads with more delivery because conversion predictions are more accurate. Once you’ve added more than 6 ads, there is little marginal benefit.
When that was captured. February 2025.
What the live page says now. Decrease ads per ad set, but maintain diverse creative assets per ad set. One ad can contain multiple, up to 10, creative assets.
What changed in substance. The direction is the same, fewer ads, but the threshold is gone and the emphasis has moved from separate ads to multiple assets inside one ad. That points at the platform’s automatic creative system rather than at manual testing.
Why the reason given mattered. The archived version explained the mechanism: more delivery per ad produces more accurate conversion predictions. That is an argument about statistical power, and it is the same argument that applies to your own testing.
What to do with a withdrawn figure. Treat it as a former platform position rather than current guidance. It tells you what the delivery system was optimising for in 2025, which is still useful, and it is no longer something you can cite as policy.
What the ceilings are, and why they are not advice
There are hard limits published, and they get quoted as recommendations. They are not.
The account-level ceiling. Active ads per page, scaling with spend: 250 ads for advertisers under $100,000 in their highest-spending month, 1,000 under $1M, 5,000 under $10M, and 20,000 at $10M or more.
What that is. A platform capacity limit, sized so that it never binds on anyone running a sensible account.
What it is not. Permission. If you are a small advertiser, the fact that you may run 250 ads says nothing about whether you should run more than five.
The relationship between the two numbers. The withdrawn guidance said six or fewer per ad set. The ceiling says 250 per page. Those are answers to different questions, and conflating them is how accounts end up with 80 near-identical ads competing for the same delivery.
The rule that connects them. Delivery is a fixed quantity divided among your ads. More ads is not more delivery; it is the same delivery, in smaller pieces, each learning more slowly.
Test duration. The largest social platform recommends a minimum of 7-day tests, permits a maximum of 30 days, and states that tests shorter than seven days may produce inconclusive results.
Test budget. Its guidance says only to set a budget that will produce enough results to confidently determine a winning strategy. No dollar amount. No minimum event count. The specific figures circulating online are agency heuristics.
How a winner is determined. The platform describes simulating possible outcomes tens of thousands of times to determine how often winning outcomes would have won, producing a confidence percentage representing the chance of similar results if the test were repeated.
On the search side, concrete counts. A responsive search ad takes a minimum of 3 headlines and up to 15, and a minimum of 2 descriptions and up to 4.
And an explicit recommendation. At least 2 responsive search ads with Good or Excellent strength per ad group. For display, 3 to 4 ads per ad group.
The only published performance link. The same platform reports that advertisers moving from one responsive search ad to two find a 6.6% increase in conversions at similar cost, and moving from two to three a 3.7% increase. Both figures are internal data, with no sample size, no period and no methodology published.
Note the shape of those two numbers anyway. 6.6% then 3.7%. Even in the platform’s own telling, the returns halve with each addition. That is the diminishing-returns pattern the withdrawn six-ad guidance was describing.
The platforms decline to give you a number because the number depends on something only you can see.
The arithmetic. Take your weekly conversions on the campaign. Divide by the number of creatives you intend to run. That is roughly what each creative gets per week.
What that number needs to be. Enough that a difference between two creatives is distinguishable from noise. On the evidence covered in our piece on button testing, detecting a modest difference reliably needs conversion events in the thousands per variant, which almost no B2B account has.
What follows for most B2B advertisers. You cannot statistically distinguish six creatives. You often cannot distinguish two. That is not a failure; it is the arithmetic of a business with 30 leads a month.
So what is creative variation actually for at that scale. Not measurement. Coverage and fatigue management: different propositions reaching different people, and enough rotation that the same person does not see one execution forty times.
Which changes the question. Not “how many can I test” but “how many genuinely different propositions do I have”. If the answer is two, run two. Producing six variations of one proposition gives you neither coverage nor a readable test.
The one thing worth testing at low volume. Big differences. A different offer, a different audience, a different format. Those produce effects large enough to see without thousands of events. Headline variations do not.
And the honest fallback. If volume is too low to test anything, judge creative on the two checks that need no statistics: is it attributable to you, and is the offer comprehensible. Then spend the effort on the offer instead.
Six steps, derived from the volume arithmetic rather than from a withdrawn threshold.
Count your genuinely distinct propositions. Different offers, different audiences, different problems solved. Not different headlines for the same thing.
Run one execution per proposition to start. You are buying coverage, not running an experiment.
Add a second execution per proposition only for fatigue. When frequency climbs and performance decays, rotation is the reason to add, not testing.
Keep the total low enough that delivery concentrates. The withdrawn guidance’s reason still holds: fewer ads means more delivery each, which means the system predicts better.
Test only large differences, and only one at a time. A different offer against your current one. Not four headline variants.
Run any test for at least seven days, since that is the one duration figure any platform publishes, and stop looking at it before then.
Six or fewer creatives per ad set was official guidance, with the reason that beyond six there is little marginal benefit. It has been removed from the live page.
The current page gives no figure, advising fewer ads but more assets within one ad, which points at the automatic creative system.
The account ceilings are capacity, not advice: 250 active ads under $100,000 monthly spend, rising to 20,000 at $10M.
The only published test duration is 7 days minimum, 30 maximum, with shorter tests described as possibly inconclusive.
No platform publishes a test budget. The dollar figures circulating are agency heuristics.
The search platform publishes counts: 3 to 15 headlines, 2 to 4 descriptions, at least 2 ads per group.
Its own performance figures halve with each addition: +6.6% from one ad to two, +3.7% from two to three, with no sample disclosed.
At B2B volumes you cannot separate six creatives, and often not two, so variation is for coverage and fatigue, not measurement.
The largest social platform used to say six or fewer, adding that beyond six there is little marginal benefit. That sentence has been removed from the live page, which now gives no figure at all.
Why did the guidance disappear?
The platform does not say. The current page advises decreasing ads per ad set while maintaining diverse creative assets, and notes one ad can contain up to 10 creative assets, which points toward its automatic creative system rather than toward separate ads.
Is there a limit on how many ads I can have?
Yes, but it is a capacity ceiling, not advice. It scales with spend: 250 active ads for advertisers under $100,000 in their biggest month, rising to 20,000 at $10 million or more.
How long should a creative test run?
The platform recommends a minimum of 7 days and allows a maximum of 30, and states that tests shorter than 7 days may produce inconclusive results.
What budget does a test need?
No figure is published. The guidance says only to set a budget that will produce enough results to confidently determine a winner. The dollar amounts circulating online are agency heuristics, not platform guidance.
What does the search platform recommend?
At least 2 responsive search ads of Good or Excellent strength per ad group, with 3 to 15 headlines and 2 to 4 descriptions each, and 3 to 4 ads per ad group for display.
Is there evidence that more creatives perform better?
Only the search platform's own internal figures: 6.6% more conversions moving from one responsive search ad to two, and 3.7% moving from two to three. No sample size or methodology is published.
So how many should I actually make?
As many genuinely different propositions as you have, not variations of one. Then check whether your weekly conversion volume can distinguish them at all, because usually it cannot.