Picking a landing page structure is the easy half. The hard half is knowing whether the one you picked is better than the one you replaced, and the published experimentation data suggests most teams answer that question wrong.

The people who ran a major search engine’s testing platform report that “only one third of the ideas tested at Microsoft improved the metric(s) they were designed to improve.” They also report that where genuine breakthroughs are rare, most statistically significant results are false.

That is not an argument against testing. It is an argument for choosing the structure on reasoning, testing only what is worth the traffic, and knowing what your significance threshold actually buys you.

Structure follows the objective, and there are four

The structural question has a small number of answers, and they differ by what the visitor is being asked to do rather than by industry.

Capture, where the ask is an email for something. The page carries one offer, one form, minimal navigation, and enough about the deliverable to make the exchange legible. Length is short because the commitment is small.

Qualification, where the ask is a conversation. Longer, because the visitor is deciding whether to spend their time. It carries proof, scope, who it is not for, and an indication of price or engagement shape. The exclusion is doing real work here.

Comparison, where the visitor already knows the category. Structured around the decision criteria rather than around your features, with the honest limitations included. A page that cannot say what it is worse at is not a comparison page.

And transaction, where the ask is a purchase or a booking. The structure collapses toward the action: fewest fields, clearest total, no navigation competing with the button.

What determines which one you need is the traffic, not the taste. Search traffic for a category term arrives at a different stage from paid traffic against a competitor name, which arrives differently again from a link in a nurture email.

And one structural rule holds across all four. The page must answer the question the visitor arrived with before it introduces the one you want to ask. That ordering is what most B2B landing pages invert, and putting it back is the first move when we rebuild a B2B site section by section around the job each one has to do.

Four landing page structures distinguished by the action requested and the visitor’s prior knowledgeFour distinct landing page structures, distinguished not by industry but by what the visitor is being asked to do and by how much they already know when they arrive. The capture structure applies where the request is an email address in exchange for something, and consists of one offer, one form, minimal navigation and sufficient description of the deliverable to make the exchange legible, with the page kept short because the commitment being asked for is small. The qualification structure applies where the request is a conversation, and is longer because the visitor is deciding whether to spend their own time, carrying proof, an indication of scope, an explicit statement of who the offer is not for, and an indication of price or of the shape of a typical engagement, with the exclusion statement performing genuine work by removing unsuitable enquiries. The comparison structure applies where the visitor already understands the category, and is organised around the decision criteria rather than around the seller’s feature list, including honest limitations, on the principle that a page unable to state what it is worse at is not functioning as a comparison. The transaction structure applies where the request is a purchase or a booking, and collapses toward the action itself with the fewest possible fields, the clearest possible total, and no navigation competing with the primary control. The determinant of which structure is required is the traffic source rather than preference, since search traffic arriving on a category term is at a different stage from paid traffic arriving against a competitor’s name, which differs again from a click within a nurture email sequence. One structural rule holds across all four, namely that the page must answer the question the visitor arrived with before it introduces the question the seller wishes to ask, an ordering which most business to business landing pages invert.Four structures, decided by the trafficCaptureAsk: an email for something.One offer, one form, minimal navigation.Short, because the commitment is small.QualificationAsk: a conversation.Proof, scope, who it is not for, price shape.The exclusion is doing real work.ComparisonVisitor already knows the category.Built on their criteria, not your features.If it cannot say what you are worse at, it is not one.TransactionAsk: a purchase or a booking.Fewest fields, clearest total, no competing nav.And the rule that holds across all fourThe page answers the question the visitor arrived with before it introduces the one you want to ask.That ordering is what most B2B landing pages invert.
The structure follows what the visitor already knows and what they are being asked to give up. Source : Method (2026)

Now the part that decides whether you learned anything

Choosing well matters more than testing, because testing is far weaker than its reputation.

The base rate. “Only one third of the ideas tested at Microsoft improved the metric(s) they were designed to improve”, and “success is even harder to find in well-optimized domains.”

Corroborated elsewhere in the same paper. One organisation reports “about 10 percent of these [experiments] leading to business changes”. Another “considers 90% of what they try to be wrong”. A third practitioner: “I can only ‘guess’ the outcome of a test about 33% of the time!”

The effect sizes are small. Among those that succeed, most “improve key metrics by 0.1% to 1.0%, once diluted to overall impact”. Genuine breakthroughs run at “perhaps one in 500 experiments”.

Which sets up the finding that should change how you read a test result. Using the conventional thresholds, “if we have a prior probability of success of 1/3 … then the posterior probability for a true positive result given a statistically significant experiment is 89%. However, if breakthrough results … are one in 500, then the posterior probability drops to 3.1%.”

Read that twice. In a well-optimised area, where real winners are rare, roughly 97 percent of your significant results are noise. The p-value did not change. The base rate did.

And there is a second, avoidable inflation. Watching a dashboard and stopping when it turns green is documented as costly: “even with 10,000 samples … the false positive probability can easily be inflated by 5-10x. That means that, throughout the industry, users have been drawing inferences that are not supported by their data.”

How the base rate of successful ideas determines whether a statistically significant test result is trueHow the underlying rate at which ideas actually succeed determines the probability that a statistically significant experiment result reflects a real effect, drawn from published research by the team operating a major search engine’s experimentation platform. The base rate itself is reported directly, with only one third of ideas tested at that company improving the metrics they were designed to improve, and success described as even harder to find in well optimized domains. Corroborating figures from other organizations are collected in the same paper, including approximately ten percent of controlled experiments at one company leading to business changes, ninety percent of attempts considered wrong at another, and a practitioner reporting an ability to guess the outcome of a test about a third of the time. Effect sizes among successful experiments are small, most improving key metrics by between one tenth of one percent and one percent once diluted to overall impact, with genuine breakthroughs running at approximately one experiment in five hundred. Applying a significance threshold of zero point zero five and a power of eighty percent, the posterior probability that a statistically significant result is a true positive is eighty nine percent where the prior probability of success is one in three, but falls to three point one percent where breakthroughs occur at one in five hundred. The consequence is that in a well optimized area where real winners are rare, approximately ninety seven percent of statistically significant results are noise, without the significance threshold itself having changed, because it is the base rate rather than the threshold that moved. A second and avoidable source of inflation is adaptive stopping, since stopping experiments through continuous monitoring of a dashboard severely and favourably biases the selection of experiments deemed significant, with a separate paper reporting that even with ten thousand samples the false positive probability can easily be inflated by five to ten times, meaning inferences drawn throughout the industry are not supported by the underlying data.Same p-value, two very different answersIf 1 idea in 3 really works89%of significant results are trueThe reported base rate across manyexperiments at one large company.If 1 idea in 500 really works3.1%of significant results are trueThe reported rate of genuinebreakthroughs in a well-optimised area.The threshold did not move. The base rate did.Both figures use the same conventional significance level and the same power. Only the prior probability differs.And the inflation you controlStopping when the dashboard turns green: “the false positive probability can easily be inflated by 5-10x.”Which is why choosing the structure well matters more than testing variations of a bad one.
The same p-value means something different depending on how often ideas actually work. In a well-optimised area, most wins are noise. Source : Kohavi, Deng, Longbotham and Xu, Seven Rules of Thumb for Web Site Experimenters, KDD 2014 (2014)

How much traffic a real test needs

The published guidance replaces the folklore numbers with a formula that depends on your metric.

The rule of thumb. “Our rule of thumb for the minimum number of independent and identically distributed observations needed for the mean to have a normal distribution is 355 × s² for each variant, where s is the skewness coefficient”, recommended “when the |skewness| > 1.”

A worked example from the same source. A revenue-per-user metric with a skewness of 17.9 requires 114,000 users per variant, and that only detects a 4.4 percent change at 80 percent power.

And the authors dismiss the generic numbers directly. “Our advice in previous articles is that you need ‘thousands’ of users in an experiment … but the guidance should be refined to the metrics of interest.”

Their own experiments run far larger than most B2B sites can. “Sample sizes for the experiments are at least in the hundreds of thousands of users, with most experiments involving millions.”

With a duration floor. “Each experiment ran for at least a week”, and for novelty effects “we recommend running experiments for two weeks.”

Which produces an uncomfortable but useful conclusion for a B2B site. If your landing page sees a few thousand visits a month and converts in the low single digits, you cannot detect the effect sizes that experiments typically produce. That is not a reason to stop improving the page. It is a reason to stop pretending the improvement was measured.

And one corollary about null results. A published critique of a well-known “no effect” finding notes “a more likely hypothesis is that the experiment did not have sufficient statistical power to detect the differences.” No effect found is not the same as no effect.

Published sample size guidance for web experiments and its implication for low traffic sitesPublished guidance on the sample size and duration required for a valid web experiment, together with its implication for sites with modest traffic. The stated rule of thumb for the minimum number of independent and identically distributed observations needed for the mean to follow a normal distribution is three hundred and fifty five multiplied by the square of the skewness coefficient, calculated separately for each variant, and recommended for use where the absolute value of the skewness exceeds one. A worked example from the same source concerns a revenue per user metric carrying a skewness of seventeen point nine, which under the rule requires one hundred and fourteen thousand users per variant, and which at that sample size permits detection of a four point four percent change at eighty percent power. The authors explicitly dismiss generic visitor thresholds circulating in the industry, noting that their own previous advice referred to needing thousands of users and that another commentator suggests ten thousand monthly visitors, while stating that the guidance should instead be refined to the metrics of interest. Their own experiments are described as involving at least hundreds of thousands of users after removal of automated traffic, with most involving millions, providing statistical power to detect small differences with high statistical significance. A duration floor is specified, each experiment running for at least a week, with a recommendation of two weeks where novelty or primacy effects are a concern. The implication for a business to business landing page receiving a few thousand visits per month and converting in the low single digits is that it cannot detect the effect sizes experiments typically produce, which are between one tenth of one percent and one percent once diluted. This is not a reason to cease improving the page but a reason to cease describing such improvements as measured. A corollary concerns null results, since a published critique of a well known finding of no effect observes that a more likely hypothesis is that the experiment lacked sufficient statistical power to detect the differences, meaning that no effect found is not equivalent to no effect.What a real test costs in trafficThe published rule, per variant355 × s²where s is the skewness coefficient, recommended when |skewness| > 1Their worked exampleRevenue per user, skewness 17.9114,000 users per variantAnd that detects a 4.4% change at 80% power.And the generic numbers”The guidance should be refined to themetrics of interest.”Their own runs: hundreds of thousands, mostly millions.Duration floor: “each experiment ran for at least a week”, and two weeks where novelty effects are a concern.And on null results: a published critique of a famous “no effect” finding suggests “the experiment did not havesufficient statistical power to detect the differences.” No effect found is not no effect.
The formula depends on your metric's skew. The generic visitor thresholds are explicitly dismissed by the source. Source : Kohavi, Deng, Longbotham and Xu, Seven Rules of Thumb, KDD 2014 (2014)

Two rules you can stop applying

Both are widely repeated and both have been measured.

The three-click rule does not survive contact with data. A study of 44 users completing 620 tasks across more than 8,000 clicks reported: “our analysis showed that there wasn’t any more likelihood of a user quitting after three clicks than after 12 clicks.”

And the follow-on findings are more useful than the headline. “It isn’t until 15 clicks that we see 80% of our tasks completed.” Also: “successful clickstreams have the same distribution as unsuccessful clickstreams, the number of clicks doesn’t predict task success or failure.” And on satisfaction, it varied only “between 46% and 61%” across clickstream lengths, so “fewer clicks do not make more satisfied users.”

The seven-item menu limit is a misapplied memory result. Researchers who set out to confirm it built three sites with 512 documents each, varying only the link structure: three levels of eight links, or 16 then 32, or 32 then 16. The outcome: “users performed reliably slowest and were most lost when the number of links was held to 8.”

Their explanation is the transferable part. “When picking between web links subject’s short-term memory does not appear to be an important factor. Instead, it is far more important to ensure that the labeling is good enough to tell people what information will be available if they follow the link.”

Which reframes the navigation decision on a landing page. The count of options is not the problem. Labels that do not tell you what is behind them are.

One honesty note. The participant count for that study is not stated in the authors’ own later account, and the original paper is paywalled. The three conditions and the conclusion are verified; the sample size is not.

Two widely repeated navigation rules tested empirically and the results obtainedTwo widely repeated rules about web navigation which were tested empirically and which the resulting data did not support. The first is the rule that a user should be able to reach any destination within three clicks. It was tested in a study involving forty four users, six hundred and twenty tasks and more than eight thousand counted clicks. The analysis showed no greater likelihood of a user quitting after three clicks than after twelve clicks. Further findings from the same study were that it is not until fifteen clicks that eighty percent of tasks are completed, that successful clickstreams have the same distribution as unsuccessful clickstreams so that the number of clicks does not predict task success or failure, and that reported satisfaction varied only between forty six and sixty one percent across clickstream lengths, meaning fewer clicks do not produce more satisfied users. The second is the rule that a navigation menu should contain no more than approximately seven items, derived by misapplication from a psychology paper concerning short term memory for chunks whose original experiment involved matching a new tone against a set of existing tones. It was tested by researchers who constructed three websites each containing five hundred and twelve documents of similar content, varying only the link structure, with one site presenting three levels of eight links per level, a second presenting sixteen links at the top level and thirty two at the second, and a third presenting thirty two at the top level and sixteen at the second. Users performed reliably slowest and were most lost on the site where the number of links was held to eight. The authors’ conclusion was that short term memory does not appear to be an important factor when choosing between web links, and that it is far more important to ensure the labelling is good enough to tell people what information will be available if they follow the link. The participant count for that study is not stated in the authors’ later account and the original paper is not freely accessible, so the conditions and conclusion are verified while the sample size is not.Two rules, both measured, both wrongThe three-click rule44 users, 620 tasks, 8,000+ clicks”There wasn’t any more likelihood of a user quitting after three clicks than after 12 clicks.""Successful clickstreams have the same distribution as unsuccessful clickstreams.”Satisfaction varied only “between 46% and 61%”, so “fewer clicks do not make more satisfied users.”The seven-item menu limitThree sites, 512 documents eachEight links per level, or 16 then 32, or 32 then 16. Result:“Users performed reliably slowest and were most lost when the number of links was held to 8.”Their conclusion: what matters is “the labeling is good enough to tell people what information will be available”.Which reframes the navigation decision entirely: the count of options is not the problem.Labels that do not say what is behind them are.
Both were tested with real tasks. One found no relationship at all; the other found the opposite of its hypothesis. Source : Porter, Testing the Three-Click Rule, 2003; Czerwinski and Larson, Cognition and the Web (2003)

What to do instead of testing everything

The published data suggests a different allocation of effort.

Choose the structure from the traffic source, not from a template gallery. Capture, qualification, comparison or transaction, decided by where the visitor came from and what they already know.

Fix the things that are not tests. A missing postal address, an unlabelled field, a target under 24 pixels, a form asking for the same thing twice. These are defects, not variants, and they do not need traffic to justify.

Reserve experiments for changes worth the sample. If the honest sample calculation says you cannot detect the effect, the experiment produces a number rather than an answer.

Never stop a test because it turned green. Fix the duration in advance, at a minimum of one week and preferably two, and read it once.

And write down what you expected before you look. The base rate is what turns a significant result into a believable one, and you can only estimate your base rate if you recorded your predictions.

One thing worth measuring that is not a conversion rate. Whether the page answers the question people arrive with. Six people reading it aloud will tell you that in an afternoon, and no sample size calculation is required.

Separation of page defects from testable hypotheses and the different treatment each requiresThe separation of proposed page changes into defects and hypotheses, and the different treatment each category requires, which is the mechanism that makes a limited testing budget affordable. Defects are changes correcting something that is objectively wrong rather than uncertain, and they include a missing physical postal address in a commercial message, an unlabelled form field, a pointer target smaller than the applicable minimum size, a form requesting the same information twice within one process, an error message conveyed only by colour rather than in text, and a focus indicator hidden behind a sticky header. These require no traffic and no measurement to justify, because a broken label is not a hypothesis, and shipping them immediately costs nothing in statistical terms. Hypotheses are genuine uncertainties about which alternative performs better, such as a change of headline, a change of page order, the addition or removal of a section, or a change to the offer itself. These require an honest sample size calculation performed before the variant is built, using the published rule of three hundred and fifty five multiplied by the square of the skewness coefficient per variant where absolute skewness exceeds one, and a fixed duration set in advance of at least one week and preferably two. Where the calculation shows that available traffic cannot detect the effect sizes experiments typically produce, which are between one tenth of one percent and one percent once diluted, the question should be decided by reasoning and qualitative research rather than by an underpowered experiment, and that decision should be described accurately rather than presented as measured. Two procedural rules apply to the hypotheses that do proceed, namely that the test duration is fixed in advance and the dashboard is not consulted in between, since adaptive stopping is documented to inflate the false positive probability by five to ten times, and that the expected outcome is recorded before the result is examined, since the base rate of successful predictions is what converts a statistically significant result into a believable one.Sort the backlog before you test itDefects: ship now, measure nothingA missing postal addressAn unlabelled fieldA target under 24 pixelsThe same question asked twiceAn error shown only in redHypotheses: cost a sampleA different headlineA different page orderA section added or removedA different offerRun the arithmetic before building.And two rules for whatever survives the sortFix the duration in advance and do not look. Write down what you expect before you read the result.If the traffic cannot detect it, say soDeciding by reasoning is legitimate. Presenting an underpowered test as evidence is not, and it is the morecommon of the two.
A broken label is not a variant. Sorting the backlog this way is what makes the remaining tests affordable. Source : Method, over the published experimentation guidance (2026)

What to do with this

Sort your planned page changes into two piles: defects and hypotheses. Ship every defect this week without measuring anything, because a broken label is not a hypothesis.

For what remains, run the sample size arithmetic honestly before you build the variant. If your traffic cannot detect a one percent change, decide the question by reasoning and research instead, and say so.

Set test durations in advance and stop looking at the dashboard in between. The documented cost of adaptive stopping is a five to ten times inflation in false positives, and it is entirely self-inflicted.

And when a test does come back significant, ask what your base rate is before you believe it. In a page you have already optimised twice, the honest prior is low, and the honest reading of a significant result is that it probably is not one.

The related pieces are what a B2B homepage has to do and from website to booked meeting.