Fifty results in a week, one to two conversion cycles, up to three weeks to calibrate. Those are published figures, and stacking them gives you a floor below which no judgement about performance is possible.
The floor is higher than most agency contracts assume, and the gap between the two is where a lot of relationships end badly. Not because anyone failed, but because the verdict arrived before the evidence could.
This page assembles the documented periods, then separates what you genuinely can assess early from what you cannot.
The clocks that have to run first
Four published figures, each of which delays a reliable verdict.
On the social platform, verbatim. Ad sets exit the learning phase as soon as they can deliver stably, which usually occurs after about 50 results in the week after the ad set’s last significant edit.
The failure state, verbatim. An ad set becomes learning limited when it is unlikely to receive about 50 optimization events in the week after the last significant edit.
What restarts that clock. Always: any change to targeting, any change to creative, any change to the optimization event, adding a new ad, pausing seven days or longer, changing bid strategy.
Read that list again in the context of a new agency. A new agency changes targeting, refreshes creative and adds ads. Every one of those is documented as always significant. Their first month of work restarts the clock by design.
On the search platform, verbatim. Smart Bidding will take some time to learn, one to two conversion cycles in most cases, after any changes to conversion goals or actions.
And on calibration, verbatim. It can take up to three weeks or one to two conversion cycles for the bid strategy to calibrate to a new objective.
The one change that is free, verbatim. Changing a target will not trigger a learning status and will not reset anything Smart Bidding has already learned.
What the stack implies. If a new agency restructures in weeks one and two, the learning clock starts in week three, runs one to two conversion cycles, and only then produces steady-state data. On a B2B account that is comfortably into the second month before the first honest number exists.
Some accounts cannot produce a verdict at any horizon, and it is worth knowing whether yours is one of them.
The threshold again. Roughly 50 optimization events per week, per ad set, to exit the learning phase.
The arithmetic. An account producing 40 leads a month across three ad sets gives each roughly three per week. That is not a slow exit from the learning phase. It is no exit.
What the platform recommends instead. Combining ad sets and campaigns, broadening the audience, raising budget, or changing the optimization event. Consolidation, in other words.
On the search side, the published eligibility thresholds. Target return on ad spend needs at least 15 conversions in 30 days on search and shopping, 30 in 30 days on video action, and 50 per week on some campaign types. Target cost per acquisition can start with no conversion history, but the platform recommends evaluating over 30 days including at least 30 conversions.
What that gives you as a test. Take your monthly conversions, divide by your number of ad sets or campaigns, and compare against those figures. If you are below them, more time will not help, and neither will a different agency.
Which reframes the whole question. If your volume is below the thresholds, the honest brief is not “improve performance”. It is “get us above the volume where optimisation becomes possible”, and that is a different piece of work with a different timeline.
And the uncomfortable implication for small B2B accounts. Many of them are permanently below these thresholds. For those, the agency’s value is in offer, targeting and creative judgement, not in algorithmic optimisation, and they should be judged accordingly.
What you can judge at thirty days
Process, not performance. All of it is observable immediately and all of it predicts the outcome.
Whether tracking is correct. Submit a test lead and confirm it appears in the platform, in analytics and in your CRM with the right source. If this is wrong at day thirty, nothing measured afterwards means anything.
Whether the account structure is explicable. Ask them to walk you through why campaigns are split the way they are. A structure that cannot be explained in five minutes was not designed.
Whether reporting carries its definitions. Attribution window, counting setting, model, and which recent days are provisional. Their absence at month one predicts an argument at month six.
Whether they told you what they changed and when. A dated change log. Without it, no future performance shift can be attributed to anything.
Whether they say what they do not know. The single most reliable early signal. A supplier who volunteers uncertainty at month one is a supplier who will tell you when something breaks.
And whether they resisted the urge to rebuild everything in week one. Given the documented relearning costs, an incoming agency that restructures immediately has spent your first quarter to satisfy its own preferences. Sometimes that is justified. It should be argued, not assumed.
The clock nobody puts in the contract: your sales cycle
Platform learning periods are measured in weeks. The measure that matters to the business is measured in months, and it is yours, not the agency’s.
The chain. A click becomes a form submission, a form submission becomes a qualified conversation, a qualified conversation becomes an opportunity, an opportunity becomes revenue. Each step has a delay, and the delays add. Each step also belongs to a different supplier in most setups, which is how a chain ends up with nobody answering for the whole of it, and it is why the ads, the page and the measurement are worth holding in one place.
Which means the honest question is not “how long before we can judge the agency”. It is “how long before the leads generated in month one have had time to become anything”, and only your own historical data answers that.
The number to pull before you sign, not after. Take your last thirty closed deals. For each one, the days between first touch and signature. The median of those is your minimum observation period for revenue-based judgement.
What that gives you. Two review dates instead of one. An early one on process and unit economics, and a later one on pipeline contribution, set at first-touch-to-close plus the relearning period.
And a warning about the middle. Between those two dates, cost per lead will move and it will tempt you to conclude something. Cost per lead is the metric most easily improved by attracting worse leads. Do not act on it alone before the qualification data catches up.
If the honest verdict date is later than a quarter, the contract should say so rather than pretend otherwise.
A stated stabilization window. A period at the start during which the incoming agency may restructure, and after which the account is held stable. The relearning clock starts at the end of it, not at signature.
Two named review dates. One on process, one on contribution, each with its criteria written down in advance. Criteria agreed before the numbers exist are criteria nobody argues about afterwards.
A change log obligation. Dated, listing what changed and why. This is what makes any later attribution possible at all.
A notice period that matches the clocks. A thirty-day notice on an engagement whose first honest verdict is at month five is a contract that will be terminated on noise.
And an exit clause covering assets. Given that some platform assets cannot be transferred, the contract should say what happens to them: which accounts you own, which datasets stay, what is handed over and in what format.
Five clauses that align the commercial agreement with the platforms' own published relearning periods. Source : Method (2026)
The verdict date, assembled
Put the published numbers end to end and the earliest honest date falls out of the arithmetic rather than out of convention.
Start of the relearning clock. Not signature. The end of the stabilization window, once the incoming agency stops making significant edits.
Add the relearning period. One to two conversion cycles on search, up to three weeks for a bidding strategy to calibrate to a new objective. Roughly a month before the account is in a steady state at all.
Add the volume needed to read anything. At the published thresholds, an account producing fifteen conversions a month reaches a thirty-conversion evaluation base in two months, not two weeks.
Add your sales cycle. The median first-touch-to-signature figure from your own closed deals.
And the result, for a typical B2B engagement. Process is readable at thirty days. Steady-state performance somewhere around month three. Pipeline contribution only after the sales cycle has run on leads generated post-stabilization, which for a four-month cycle lands past month six.
Which is not an argument for patience without conditions. It is an argument for judging the right thing at the right time: process early, and hard, on criteria that are readable immediately; results later, on criteria written before anyone knew what the numbers would say.
Four habits that push the verdict date further out
Most engagements take longer to become readable than the published periods require, and the reasons are usually client-side.
Changing the optimization event mid-flight. It is an always-significant edit on the social side and a relearning trigger for automated bidding on the search side. Every time the definition of a conversion moves, the clock restarts and the historical comparison breaks.
Asking for a new creative every fortnight. Adding an ad to an ad set is listed among the always-significant edits. A refresh cadence shorter than the learning period means the ad set never leaves it.
Splitting campaigns to get cleaner reporting. Every split divides the same volume across more ad sets, and the threshold is per ad set. Reporting granularity bought at the cost of never exiting the learning phase is a bad trade.
Pausing over a weekend to control spend. A pause of seven days or more counts as a significant edit. A cautious pause in week three costs you week four as well.
What follows from all four. The observation period is not something the agency alone controls. Half of the delay in a typical engagement is bought by the client, one reasonable-sounding request at a time, and the only defence is the stabilization window written into the contract.
What to do with this
Before you sign, pull the median first-touch-to-signature figure from your last thirty closed deals, and divide your monthly conversions by your number of ad sets. Those two numbers decide your review dates. Then write the dates and their criteria into the contract, along with the stabilization window and the change log obligation.
If you are already six weeks into an engagement and wondering, the answer is that you cannot judge the results yet but you can judge everything else. Ask for the change log, submit a test lead, and ask them to explain the account structure. What comes back tells you more than the cost per lead will for another three months.
If the account turns out to be permanently below the volume thresholds, the brief is different from the one you wrote. Read why B2B ads fail for what that changes, and what agency reporting must contain for the definitions to demand from month one.
Frequently asked questions
How long before I can judge a new agency?
Longer than a month. Platform relearning alone runs one to two conversion cycles and up to three weeks for calibration, and that clock only starts once the account stops being changed.
What does the learning phase actually require?
Meta states ad sets exit it after about 50 results in the week following the last significant edit. An ad set unlikely to reach that figure is described as learning limited.
So what if my account cannot produce 50 results a week?
Then your ad sets never leave the exploratory state, and no amount of time fixes it. That is a volume problem, not an agency problem, and consolidating ad sets is the documented remedy.
How long does Google's bidding take to relearn?
One to two conversion cycles after changes to conversion goals or actions, and up to three weeks to calibrate to a new objective. Changing only a target resets nothing.
Does a creative refresh restart the clock?
On Meta, yes. Adding a new ad to an ad set is always a significant edit, as are changes to targeting, creative and the optimization event. So an agency doing its job restarts the learning period.
What is the earliest honest verdict?
Roughly one quarter for leading indicators like cost per qualified lead, and one full sales cycle for anything about revenue. Anything sooner is reading noise.
What can I judge at 30 days then?
Process, not performance. Whether tracking is correct, whether the account is structured coherently, whether reporting arrives with definitions, and whether they tell you what they do not know.
Should I set a review date in the contract?
Yes, and make it the same period next quarter rather than next month, agreed in writing before anyone starts. Deciding the comparison point afterwards is how both sides end up arguing.