Fifty results in a week, one to two conversion cycles, up to three weeks to calibrate. Those are published figures, and stacking them gives you a floor below which no judgement about performance is possible.

The floor is higher than most agency contracts assume, and the gap between the two is where a lot of relationships end badly. Not because anyone failed, but because the verdict arrived before the evidence could.

This page assembles the documented periods, then separates what you genuinely can assess early from what you cannot.

The clocks that have to run first

Four published figures, each of which delays a reliable verdict.

On the social platform, verbatim. Ad sets exit the learning phase as soon as they can deliver stably, which usually occurs after about 50 results in the week after the ad set’s last significant edit.

The failure state, verbatim. An ad set becomes learning limited when it is unlikely to receive about 50 optimization events in the week after the last significant edit.

What restarts that clock. Always: any change to targeting, any change to creative, any change to the optimization event, adding a new ad, pausing seven days or longer, changing bid strategy.

Read that list again in the context of a new agency. A new agency changes targeting, refreshes creative and adds ads. Every one of those is documented as always significant. Their first month of work restarts the clock by design.

On the search platform, verbatim. Smart Bidding will take some time to learn, one to two conversion cycles in most cases, after any changes to conversion goals or actions.

And on calibration, verbatim. It can take up to three weeks or one to two conversion cycles for the bid strategy to calibrate to a new objective.

The one change that is free, verbatim. Changing a target will not trigger a learning status and will not reset anything Smart Bidding has already learned.

What the stack implies. If a new agency restructures in weeks one and two, the learning clock starts in week three, runs one to two conversion cycles, and only then produces steady-state data. On a B2B account that is comfortably into the second month before the first honest number exists.

Published learning and calibration periods that must elapse before advertising performance is readablePresentation of the published learning and calibration periods that must elapse before advertising performance can be reliably assessed, and the way a new agency’s initial work necessarily restarts them. The social platform states that advertisement sets exit the learning phase as soon as they can deliver stably, which usually occurs after approximately fifty results in the week following the advertisement set’s last significant edit, and that an advertisement set becomes learning limited when it is unlikely to receive about fifty optimization events in that week. The edits that always restart this clock comprise any change to targeting, any change to advertisement creative, any change to the optimization event, the addition of a new advertisement to the advertisement set, pausing the advertisement set for seven days or longer, and changing the bid strategy. The significance for agency transitions is that a new agency necessarily changes targeting, refreshes creative and adds advertisements, each of which is documented as always significant, so the agency’s first month of work restarts the learning clock by design rather than by error. On the search platform, automated bidding requires some time to learn, in most cases one to two conversion cycles, following any changes to conversion goals or actions, and calibration to a new objective can take up to three weeks or one to two conversion cycles. One change is explicitly free: changing a target will not trigger a learning status and will not reset anything the bidding system has already learned. Stacking these constraints, an agency that restructures during weeks one and two starts the learning clock in week three, which then runs for one to two conversion cycles before steady-state data exists, placing the first honest measurement comfortably within the second month on a business-to-business account.The clocks, stackedPublished periodWhat it delays~50 results in a week, per ad setExiting the learning phase at all1 to 2 conversion cyclesRelearning after conversion changesUp to 3 weeksCalibrating to a new objectiveAnd the new agency restarts it by doing the jobAlways significant: any change to targeting, any change to creative, any change to the optimizationevent, adding a new ad. Those are the first four things any incoming agency does.The one free changeAdjusting a target resets nothing already learned.So the first honest numberArrives in month two at the earliest.Restructure in weeks 1 and 2, clock starts week 3, then one to two conversion cycles on top.
A new agency's first month restarts the clock by definition, because doing the work is a significant edit. Source : Platform learning phase and bidding documentation (2026)

The volume problem that time does not solve

Some accounts cannot produce a verdict at any horizon, and it is worth knowing whether yours is one of them.

The threshold again. Roughly 50 optimization events per week, per ad set, to exit the learning phase.

The arithmetic. An account producing 40 leads a month across three ad sets gives each roughly three per week. That is not a slow exit from the learning phase. It is no exit.

What the platform recommends instead. Combining ad sets and campaigns, broadening the audience, raising budget, or changing the optimization event. Consolidation, in other words.

On the search side, the published eligibility thresholds. Target return on ad spend needs at least 15 conversions in 30 days on search and shopping, 30 in 30 days on video action, and 50 per week on some campaign types. Target cost per acquisition can start with no conversion history, but the platform recommends evaluating over 30 days including at least 30 conversions.

What that gives you as a test. Take your monthly conversions, divide by your number of ad sets or campaigns, and compare against those figures. If you are below them, more time will not help, and neither will a different agency.

Which reframes the whole question. If your volume is below the thresholds, the honest brief is not “improve performance”. It is “get us above the volume where optimisation becomes possible”, and that is a different piece of work with a different timeline.

And the uncomfortable implication for small B2B accounts. Many of them are permanently below these thresholds. For those, the agency’s value is in offer, targeting and creative judgement, not in algorithmic optimisation, and they should be judged accordingly.

What you can judge at thirty days

Process, not performance. All of it is observable immediately and all of it predicts the outcome.

Whether tracking is correct. Submit a test lead and confirm it appears in the platform, in analytics and in your CRM with the right source. If this is wrong at day thirty, nothing measured afterwards means anything.

Whether the account structure is explicable. Ask them to walk you through why campaigns are split the way they are. A structure that cannot be explained in five minutes was not designed.

Whether reporting carries its definitions. Attribution window, counting setting, model, and which recent days are provisional. Their absence at month one predicts an argument at month six.

Whether they told you what they changed and when. A dated change log. Without it, no future performance shift can be attributed to anything.

Whether they say what they do not know. The single most reliable early signal. A supplier who volunteers uncertainty at month one is a supplier who will tell you when something breaks.

And whether they resisted the urge to rebuild everything in week one. Given the documented relearning costs, an incoming agency that restructures immediately has spent your first quarter to satisfy its own preferences. Sometimes that is justified. It should be argued, not assumed.

What can and cannot be assessed about an advertising agency at thirty daysSeparation of what can legitimately be assessed about an advertising agency at thirty days from what cannot. What can be assessed at thirty days is process rather than performance, and all of it is immediately observable while also predicting the eventual outcome. First, whether tracking is correct, tested by submitting a test enquiry and confirming it appears in the advertising platform, in the analytics property and in the customer relationship management system with the correct source attribution, since if this is incorrect at day thirty then nothing measured subsequently carries meaning. Second, whether the account structure is explicable, tested by asking the agency to explain why campaigns are divided as they are, on the principle that a structure which cannot be explained within five minutes was not deliberately designed. Third, whether reporting carries its own definitions, specifically the attribution window, the conversion counting setting, the attribution model and which recent days remain provisional, the absence of which at month one reliably predicts a dispute at month six. Fourth, whether the agency recorded what it changed and when in a dated change log, without which no subsequent performance movement can be attributed to any action. Fifth, whether the agency states what it does not know, which is the single most reliable early signal, since a supplier volunteering uncertainty at month one is a supplier who will report a problem when one occurs. Sixth, whether the agency resisted rebuilding everything during the first week, since given the documented relearning costs an incoming agency restructuring immediately has consumed the client’s first quarter to satisfy its own preferences, which is sometimes justified but should be argued rather than assumed.At thirty days: process, not performanceFully readable at month oneIs tracking correct? Send a test lead.Can they explain the structure in five minutes?Does reporting carry its definitions?Is there a dated change log?Do they say what they do not know?Did they resist rebuilding everything?All observable now. All of it predicts the outcome.Not readable at month oneCost per qualified leadCost per meeting heldPipeline contributionWhether the creative direction worksWhether the targeting is rightBecause the learning clocks have not finished running,and the agency’s own work restarted them.The most reliable early signal of allWhether they volunteer what they do not know. That predicts whether they will tell you when something breaks.If tracking is wrong at day thirty, nothing you measure afterwards means anything. Test it first.
Process is fully observable at month one and predicts the outcome. Performance is not readable until the clocks have run. Source : Method, applied against the published learning periods (2026)

The clock nobody puts in the contract: your sales cycle

Platform learning periods are measured in weeks. The measure that matters to the business is measured in months, and it is yours, not the agency’s.

The chain. A click becomes a form submission, a form submission becomes a qualified conversation, a qualified conversation becomes an opportunity, an opportunity becomes revenue. Each step has a delay, and the delays add. Each step also belongs to a different supplier in most setups, which is how a chain ends up with nobody answering for the whole of it, and it is why the ads, the page and the measurement are worth holding in one place.

Which means the honest question is not “how long before we can judge the agency”. It is “how long before the leads generated in month one have had time to become anything”, and only your own historical data answers that.

The number to pull before you sign, not after. Take your last thirty closed deals. For each one, the days between first touch and signature. The median of those is your minimum observation period for revenue-based judgement.

What that gives you. Two review dates instead of one. An early one on process and unit economics, and a later one on pipeline contribution, set at first-touch-to-close plus the relearning period.

And a warning about the middle. Between those two dates, cost per lead will move and it will tempt you to conclude something. Cost per lead is the metric most easily improved by attracting worse leads. Do not act on it alone before the qualification data catches up.

Three sequential clocks that determine when an agency’s work becomes judgeableThree clocks that run in sequence rather than in parallel, and therefore add rather than overlap, when determining the earliest date at which an advertising agency’s work can be judged on business outcomes. The first clock is relearning, which begins at the moment the incoming agency stops making significant changes to the account rather than at the moment the contract starts. On the social platform this requires approximately fifty optimization events per advertisement set within seven days of the last significant edit, and on the search platform automated bidding requires one to two conversion cycles after a change of conversion goal or action, with calibration to a new objective taking up to three weeks. The second clock is volume accumulation, which is the time required to gather enough conversions for the published eligibility thresholds to be met, at least fifteen conversions in thirty days for target return on advertising spend on search and shopping campaigns and thirty conversions in thirty days recommended for evaluating target cost per acquisition, meaning an account producing few conversions per campaign takes proportionally longer or never arrives. The third clock is the client’s own sales cycle, measured as the median number of days between first touch and signature across the last thirty closed deals, which is the delay before leads generated during the first two clocks can have become revenue. Because the second clock cannot begin accumulating meaningful data until the first has finished and the third cannot begin until leads exist, the earliest honest verdict date on business outcomes is the sum of the three rather than the longest of them, which is why a review scheduled at ninety days for a business with a four month sales cycle measures process rather than results.They run in sequence, so they add1. RelearningStarts when they stop changingthings, not when you sign.50 results / 7 days per ad set2. Volume accumulationEnough conversions to reachthe published thresholds.15 to 30 conversions / 30 days3. Your sales cycleBefore a lead can have becomeanything measurable in revenue.Median first touch to signatureThe pull that settles it, before you signYour last thirty closed deals. Days from first touch to signature. Take the median. That is clock three, and it is notnegotiable by any supplier.Review one: processAround thirty days. Tracking, structure,change log, honesty about unknowns.Review two: contributionRelearning plus accumulation plus thesales cycle. Set the date in the contract.Cost per lead is the metric most easily improved by attracting worse leads. Do not judge on it alone in between.
The verdict date is not the longest of the three clocks. It is the sum, because they run in sequence rather than in parallel. Source : Method, over the platforms' published learning periods and thresholds (2026)

What a fair contract says

If the honest verdict date is later than a quarter, the contract should say so rather than pretend otherwise.

A stated stabilization window. A period at the start during which the incoming agency may restructure, and after which the account is held stable. The relearning clock starts at the end of it, not at signature.

Two named review dates. One on process, one on contribution, each with its criteria written down in advance. Criteria agreed before the numbers exist are criteria nobody argues about afterwards.

A change log obligation. Dated, listing what changed and why. This is what makes any later attribution possible at all.

A notice period that matches the clocks. A thirty-day notice on an engagement whose first honest verdict is at month five is a contract that will be terminated on noise.

And an exit clause covering assets. Given that some platform assets cannot be transferred, the contract should say what happens to them: which accounts you own, which datasets stay, what is handed over and in what format.

Five contract clauses that align an agency engagement with documented platform relearning periodsFive clauses that bring an advertising agency contract into line with the platforms’ documented relearning periods and with the client’s own sales cycle. First, a stated stabilization window, being a defined period at the start of the engagement during which the incoming agency may restructure the account, after which the account is held stable, so that the relearning clock starts at the end of that window rather than at contract signature. Second, two named review dates rather than one, the first assessing process and the second assessing pipeline contribution, each with its assessment criteria written down in advance, on the principle that criteria agreed before the numbers exist are criteria nobody disputes afterwards. Third, a change log obligation requiring the agency to maintain a dated record of what was changed and why, which is the precondition for attributing any later performance movement to any action. Fourth, a notice period matching the documented clocks, since a thirty day notice period on an engagement whose first honest verdict falls at month five is a contract that will be terminated on statistical noise. Fifth, an exit clause covering platform assets, which matters because the social platform states that an advertisement account created within a business portfolio can never be deleted from or transferred out of that portfolio and the same applies to datasets, so the contract must specify which accounts the client owns, which datasets remain in place, and what is handed over in what format at the end of the engagement.Five clauses that match the documented clocks1Stabilization windowRestructuring allowed inside it. The relearning clock starts when it closes.2Two review dates, criteria written firstProcess early, contribution later. Agreed before the numbers exist.3Dated change logWithout it, no later movement can be attributed to anything at all.4Notice period matched to the clocksThirty days against a month five verdict guarantees a termination on noise.5Exit clause covering platform assetsSome accounts and datasets can never be transferred out. Say who owns what.A contract whose review dates predate the platforms’ own learning periods measures noise, not work.
Five clauses that align the commercial agreement with the platforms' own published relearning periods. Source : Method (2026)

The verdict date, assembled

Put the published numbers end to end and the earliest honest date falls out of the arithmetic rather than out of convention.

Start of the relearning clock. Not signature. The end of the stabilization window, once the incoming agency stops making significant edits.

Add the relearning period. One to two conversion cycles on search, up to three weeks for a bidding strategy to calibrate to a new objective. Roughly a month before the account is in a steady state at all.

Add the volume needed to read anything. At the published thresholds, an account producing fifteen conversions a month reaches a thirty-conversion evaluation base in two months, not two weeks.

Add your sales cycle. The median first-touch-to-signature figure from your own closed deals.

And the result, for a typical B2B engagement. Process is readable at thirty days. Steady-state performance somewhere around month three. Pipeline contribution only after the sales cycle has run on leads generated post-stabilization, which for a four-month cycle lands past month six.

Which is not an argument for patience without conditions. It is an argument for judging the right thing at the right time: process early, and hard, on criteria that are readable immediately; results later, on criteria written before anyone knew what the numbers would say.

Assembled timeline from contract signature to the earliest honest verdict on pipeline contributionAssembled timeline showing the earliest honest verdict dates for an advertising agency engagement at a business with a four month median sales cycle, built by laying the platforms’ published periods end to end rather than by convention. The timeline begins at contract signature, followed by a stabilization window during which the incoming agency may restructure the account and during which no performance reading is meaningful because significant edits reset the learning period each time they are made. At approximately thirty days from signature the first review becomes possible, but only on process, meaning whether tracking is correct, whether the account structure can be explained, whether reporting carries its definitions, whether a dated change log exists and whether the agency states what it does not know. The relearning period then runs from the close of the stabilization window, requiring one to two conversion cycles for automated bidding and up to three weeks for calibration to a new objective, placing the account in a steady state at approximately month three. Volume accumulation runs concurrently with and beyond that, since an account producing fifteen conversions per month requires two months to reach a thirty conversion evaluation base. Finally the sales cycle runs on leads generated after stabilization, and at a four month median from first touch to signature this places the earliest honest verdict on pipeline contribution beyond month six from the start of the engagement. The conclusion is not that patience should be unconditional but that process should be judged early and strictly on criteria readable immediately, while results should be judged later on criteria written before anybody knew what the numbers would say.Signature to verdict, for a four month sales cycleSignatureDay 0Review 1: process~30 daysSteady state reached~month 3Review 2: contributionpast month 6Stabilization windowRelearning and accumulationSales cycle runs on post-stabilization leadsJudge early, and hardOn process. Tracking, structure, definitions,change log, stated unknowns. All of it isreadable at day thirty and predicts the rest.Judge later, on written criteriaOn contribution. With the criteria set downbefore anyone knew what the numberswould say.The dates come out of the published periods. Convention is what puts the review at ninety days regardless.
The arithmetic of the published periods, laid end to end for a business with a four month sales cycle. Source : Method, over Meta and Google published learning periods (2026)

Four habits that push the verdict date further out

Most engagements take longer to become readable than the published periods require, and the reasons are usually client-side.

Changing the optimization event mid-flight. It is an always-significant edit on the social side and a relearning trigger for automated bidding on the search side. Every time the definition of a conversion moves, the clock restarts and the historical comparison breaks.

Asking for a new creative every fortnight. Adding an ad to an ad set is listed among the always-significant edits. A refresh cadence shorter than the learning period means the ad set never leaves it.

Splitting campaigns to get cleaner reporting. Every split divides the same volume across more ad sets, and the threshold is per ad set. Reporting granularity bought at the cost of never exiting the learning phase is a bad trade.

Pausing over a weekend to control spend. A pause of seven days or more counts as a significant edit. A cautious pause in week three costs you week four as well.

What follows from all four. The observation period is not something the agency alone controls. Half of the delay in a typical engagement is bought by the client, one reasonable-sounding request at a time, and the only defence is the stabilization window written into the contract.

What to do with this

Before you sign, pull the median first-touch-to-signature figure from your last thirty closed deals, and divide your monthly conversions by your number of ad sets. Those two numbers decide your review dates. Then write the dates and their criteria into the contract, along with the stabilization window and the change log obligation.

If you are already six weeks into an engagement and wondering, the answer is that you cannot judge the results yet but you can judge everything else. Ask for the change log, submit a test lead, and ask them to explain the account structure. What comes back tells you more than the cost per lead will for another three months.

If the account turns out to be permanently below the volume thresholds, the brief is different from the one you wrote. Read why B2B ads fail for what that changes, and what agency reporting must contain for the definitions to demand from month one.