Picking a landing page structure is the easy half. The hard half is knowing whether the one you picked is better than the one you replaced, and the published experimentation data suggests most teams answer that question wrong.
The people who ran a major search engine’s testing platform report that “only one third of the ideas tested at Microsoft improved the metric(s) they were designed to improve.” They also report that where genuine breakthroughs are rare, most statistically significant results are false.
That is not an argument against testing. It is an argument for choosing the structure on reasoning, testing only what is worth the traffic, and knowing what your significance threshold actually buys you.
Structure follows the objective, and there are four
The structural question has a small number of answers, and they differ by what the visitor is being asked to do rather than by industry.
Capture, where the ask is an email for something. The page carries one offer, one form, minimal navigation, and enough about the deliverable to make the exchange legible. Length is short because the commitment is small.
Qualification, where the ask is a conversation. Longer, because the visitor is deciding whether to spend their time. It carries proof, scope, who it is not for, and an indication of price or engagement shape. The exclusion is doing real work here.
Comparison, where the visitor already knows the category. Structured around the decision criteria rather than around your features, with the honest limitations included. A page that cannot say what it is worse at is not a comparison page.
And transaction, where the ask is a purchase or a booking. The structure collapses toward the action: fewest fields, clearest total, no navigation competing with the button.
What determines which one you need is the traffic, not the taste. Search traffic for a category term arrives at a different stage from paid traffic against a competitor name, which arrives differently again from a link in a nurture email.
And one structural rule holds across all four. The page must answer the question the visitor arrived with before it introduces the one you want to ask. That ordering is what most B2B landing pages invert, and putting it back is the first move when we rebuild a B2B site section by section around the job each one has to do.
The structure follows what the visitor already knows and what they are being asked to give up. Source : Method (2026)
Now the part that decides whether you learned anything
Choosing well matters more than testing, because testing is far weaker than its reputation.
The base rate. “Only one third of the ideas tested at Microsoft improved the metric(s) they were designed to improve”, and “success is even harder to find in well-optimized domains.”
Corroborated elsewhere in the same paper. One organisation reports “about 10 percent of these [experiments] leading to business changes”. Another “considers 90% of what they try to be wrong”. A third practitioner: “I can only ‘guess’ the outcome of a test about 33% of the time!”
The effect sizes are small. Among those that succeed, most “improve key metrics by 0.1% to 1.0%, once diluted to overall impact”. Genuine breakthroughs run at “perhaps one in 500 experiments”.
Which sets up the finding that should change how you read a test result. Using the conventional thresholds, “if we have a prior probability of success of 1/3 … then the posterior probability for a true positive result given a statistically significant experiment is 89%. However, if breakthrough results … are one in 500, then the posterior probability drops to 3.1%.”
Read that twice. In a well-optimised area, where real winners are rare, roughly 97 percent of your significant results are noise. The p-value did not change. The base rate did.
And there is a second, avoidable inflation. Watching a dashboard and stopping when it turns green is documented as costly: “even with 10,000 samples … the false positive probability can easily be inflated by 5-10x. That means that, throughout the industry, users have been drawing inferences that are not supported by their data.”
The published guidance replaces the folklore numbers with a formula that depends on your metric.
The rule of thumb. “Our rule of thumb for the minimum number of independent and identically distributed observations needed for the mean to have a normal distribution is 355 × s² for each variant, where s is the skewness coefficient”, recommended “when the |skewness| > 1.”
A worked example from the same source. A revenue-per-user metric with a skewness of 17.9 requires 114,000 users per variant, and that only detects a 4.4 percent change at 80 percent power.
And the authors dismiss the generic numbers directly. “Our advice in previous articles is that you need ‘thousands’ of users in an experiment … but the guidance should be refined to the metrics of interest.”
Their own experiments run far larger than most B2B sites can. “Sample sizes for the experiments are at least in the hundreds of thousands of users, with most experiments involving millions.”
With a duration floor. “Each experiment ran for at least a week”, and for novelty effects “we recommend running experiments for two weeks.”
Which produces an uncomfortable but useful conclusion for a B2B site. If your landing page sees a few thousand visits a month and converts in the low single digits, you cannot detect the effect sizes that experiments typically produce. That is not a reason to stop improving the page. It is a reason to stop pretending the improvement was measured.
And one corollary about null results. A published critique of a well-known “no effect” finding notes “a more likely hypothesis is that the experiment did not have sufficient statistical power to detect the differences.” No effect found is not the same as no effect.
Both are widely repeated and both have been measured.
The three-click rule does not survive contact with data. A study of 44 users completing 620 tasks across more than 8,000 clicks reported: “our analysis showed that there wasn’t any more likelihood of a user quitting after three clicks than after 12 clicks.”
And the follow-on findings are more useful than the headline. “It isn’t until 15 clicks that we see 80% of our tasks completed.” Also: “successful clickstreams have the same distribution as unsuccessful clickstreams, the number of clicks doesn’t predict task success or failure.” And on satisfaction, it varied only “between 46% and 61%” across clickstream lengths, so “fewer clicks do not make more satisfied users.”
The seven-item menu limit is a misapplied memory result. Researchers who set out to confirm it built three sites with 512 documents each, varying only the link structure: three levels of eight links, or 16 then 32, or 32 then 16. The outcome: “users performed reliably slowest and were most lost when the number of links was held to 8.”
Their explanation is the transferable part. “When picking between web links subject’s short-term memory does not appear to be an important factor. Instead, it is far more important to ensure that the labeling is good enough to tell people what information will be available if they follow the link.”
Which reframes the navigation decision on a landing page. The count of options is not the problem. Labels that do not tell you what is behind them are.
One honesty note. The participant count for that study is not stated in the authors’ own later account, and the original paper is paywalled. The three conditions and the conclusion are verified; the sample size is not.
The published data suggests a different allocation of effort.
Choose the structure from the traffic source, not from a template gallery. Capture, qualification, comparison or transaction, decided by where the visitor came from and what they already know.
Fix the things that are not tests. A missing postal address, an unlabelled field, a target under 24 pixels, a form asking for the same thing twice. These are defects, not variants, and they do not need traffic to justify.
Reserve experiments for changes worth the sample. If the honest sample calculation says you cannot detect the effect, the experiment produces a number rather than an answer.
Never stop a test because it turned green. Fix the duration in advance, at a minimum of one week and preferably two, and read it once.
And write down what you expected before you look. The base rate is what turns a significant result into a believable one, and you can only estimate your base rate if you recorded your predictions.
One thing worth measuring that is not a conversion rate. Whether the page answers the question people arrive with. Six people reading it aloud will tell you that in an afternoon, and no sample size calculation is required.
Sort your planned page changes into two piles: defects and hypotheses. Ship every defect this week without measuring anything, because a broken label is not a hypothesis.
For what remains, run the sample size arithmetic honestly before you build the variant. If your traffic cannot detect a one percent change, decide the question by reasoning and research instead, and say so.
Set test durations in advance and stop looking at the dashboard in between. The documented cost of adaptive stopping is a five to ten times inflation in false positives, and it is entirely self-inflicted.
And when a test does come back significant, ask what your base rate is before you believe it. In a page you have already optimised twice, the honest prior is low, and the honest reading of a significant result is that it probably is not one.
Published data from a major experimentation platform reports that only one third of ideas tested improved the metric they were designed to improve, and that success is harder in already-optimised areas.
If my test is significant, is the change real?
It depends on your base rate. At p < 0.05 with 80 percent power, if one idea in three succeeds the result is a true positive 89 percent of the time. If one in 500 succeeds, that falls to 3.1 percent.
Can I stop a test when it turns green?
Not without cost. One paper reports that adaptive stopping through continuous monitoring can inflate the false positive probability by five to ten times, even with 10,000 samples.
Is the three-click rule real?
No. A study of 44 users, 620 tasks and more than 8,000 clicks found 'there wasn't any more likelihood of a user quitting after three clicks than after 12 clicks'.