The four experiments everyone cites about buttons have, between them, no published per-variant sample size and no published confidence interval. One of them is not an experiment at all. Three of the ten results in the most reproduced collection are negative, which the summaries tend to drop.

Meanwhile the same button is governed by a specification with exact numbers in it, which you can verify in an afternoon without a single visitor.

This page separates the two, because the industry spends its attention almost entirely on the half it cannot verify.

The red button test, in full

It is the most cited result in conversion optimisation. Here is everything its source actually reports.

The finding. Red outperformed green by 21%, meaning 21% more people clicked the red button.

The sample. More than 2,000 visits in total, across both variants combined. No per-variant split is published. No conversion counts are published.

The duration. A few days of traffic. No dates, no window.

The statistics. No confidence interval. No p-value. The article makes a general remark that A/B testing often requires thousands of trials, but publishes no figure for this test.

What the author himself wrote, and this is the part that never travels with the statistic: that the results cannot be generalized to all situations, that the most that can be said is that they hold for the conditions in which they occurred, in that page design, on that site, with that audience, and therefore that readers should not go out and blindly switch their green buttons to red without testing first.

What to take from it. A real result, honestly reported, with an explicit warning attached that the industry removed before repeating it fifteen thousand times.

The $300 million button was not a button test

The second most cited item in this field is a case study, and its headline number is an extrapolation.

What changed. On an unnamed retail site, a “Register” button that forced account creation before purchase was replaced with a “Continue as Guest” option.

What was measured. Roughly $6 million in additional revenue in the first week, and password reset requests falling by about 80%.

Where the $300 million came from. The author states it plainly: some arithmetic let them calculate the revenue being abandoned in carts by everyone who could not authenticate, and that is where the annual figure came from. It was computed, not observed.

The diagnostic numbers behind it. About 40% of visitors to the authentication page were requesting a password reset. Fewer than 25% of those requests were completed. Of those who completed one, fewer than 20% went on to purchase.

Why it still matters. Because the intervention was not cosmetic. It removed a mandatory account creation step from a purchase flow. The lesson is about forced friction, not about buttons, and certainly not about colour.

The measurement caveat the author raised himself. The authentication pages had never been instrumented before the analysis, which is why the problem had gone unnoticed.

Published evidentiary detail of the canonical call-to-action button experimentsTable recording what each of the three canonical call-to-action button experiments actually published, retrieved from the original sources including web archives where the original domains have lapsed. The colour experiment comparing a red button with a green button reports a twenty-one percent relative advantage for red, based on more than two thousand visits in total across both variants combined, run over a few days of traffic; it publishes no per-variant sample split, no conversion counts and no confidence interval or probability value, and its author wrote explicitly that the results cannot be generalized to all situations, that the most that can be said is that they hold for the conditions in which they occurred in that page design on that site with that audience, and that readers should not blindly switch green buttons to red without testing first. The three hundred million dollar button is a case study rather than an experiment: an unnamed retail site replaced a register button forcing account creation with a continue as guest option, the measured outcome being approximately six million dollars of additional revenue in the first week and a roughly eighty percent reduction in password reset requests, while the three hundred million dollar annual figure was calculated from abandoned cart revenue rather than observed, as the author states directly. The collection of ten button wording and design experiments reports percentage changes ranging from a positive three hundred and four percent to a negative eighteen percent, but publishes no sample size, no duration and no confidence level for any individual case, the author having explained in his own comment thread that most clients would not permit their data to be disclosed to competitors and offering only an aggregate statement afterwards that samples were typically above two thousand visitors at a declared ninety-eight percent confidence level. Three of those ten published results are negative, a detail generally omitted by the summaries that reproduce the collection.What the famous tests actually publishedTestHeadlineSample publishedConfidence publishedRed vs green+21%2,000+ visits, both armsNoneNo per-variant split. No conversion counts. “A few days of traffic.”The $300M button$300M/yrNot an experimentCase studyMeasured: ~$6M in week one. The $300M was calculated from abandoned carts.The ten-case collection+304% to −18%None per caseNone per caseAuthor, in his own comments: clients would not allow the data to be shownThe warning that never travels with the red button statistic”we cannot generalize these results to all situations… do not go out and blindly switch yourgreen buttons to red without testing first.”And three of the ten results in that collection are negativeChanging the button made things worse. The summaries reproducing the list tend to drop those.
Between them: no per-variant sample, no confidence interval, and one that is not an experiment. Source : Original publications, several retrieved from web archives (2013)

The collection of ten, including the three that went backwards

The most reproduced set of button experiments is worth reading in the original rather than in the listicles that quote it.

What it contains. Ten before-and-after cases from 2013: wording changes, a colour change, a font colour change, a size change, an arrow removal, a placement change.

The famous one. Changing “Start your free 30 day trial” to “Start my free 30 day trial” produced a reported 90% increase in trial signups, over three weeks.

The range. From +304% for moving the call to action below the fold, down to -18% for a font colour change.

Three of the ten are negative. The font colour change lost 18%. A bigger button lost 10.6%. Removing a green arrow lost 12.3%. Those three almost never appear in the summaries.

What is not published, for any of the ten. Sample size. Duration. Confidence level. Not one case carries them.

Why not, in the author’s own words. Challenged on exactly this point in his own comment thread, he explained that most clients are not happy about showing their data to competitors, and that he includes as much as they will allow. He then offered an aggregate: most cases based on samples above 2,000 visitors, a declared 98% confidence level.

The honest reading. A practitioner reporting real client work under a confidentiality constraint, which is a reasonable thing to do, and which is also not evidence anyone else can act on.

All ten reported outcomes from the most widely reproduced collection of button experimentsChart displaying all ten reported outcomes from the most widely reproduced collection of call-to-action button experiments, published in two thousand and thirteen and retrieved from a web archive because the original domain has since lapsed. The reported changes are as follows: moving the call to action below the fold produced a three hundred and four percent increase; changing get membership to find your gym and get membership produced a two hundred and thirteen point one six percent increase; changing start your free thirty day trial to start my free thirty day trial produced a ninety percent increase in trial signups over three weeks; changing order information and prices to get information and prices produced a fourteen point seven nine percent increase, and a thirty-eight point two six percent increase when replicated on a sister site in another language; a button colour change on a porcelain retailer produced a thirty-five point eight one percent increase in sales; adding the phrase get started produced a thirty-one point zero three percent increase; adding a view bundle text link produced a seventeen point one eight percent increase. Three results were negative: a font colour change lost eighteen point zero one percent, a larger button lost ten point five six percent, and removing a green arrow lost twelve point two nine percent. No individual case in the collection publishes a sample size, a duration or a confidence level, the author having explained in his own comment thread that most clients would not permit their data to be shown to competitors, and offering only an aggregate statement afterwards that most cases rested on samples above two thousand visitors at a declared ninety-eight percent confidence level. The three negative results are generally omitted from the summaries that reproduce this collection, despite being the most informative entries, since they demonstrate that changing a button can and does make performance worse.All ten, including the three nobody quotesNo sample size, duration or confidence level published for any individual case.worsebetterCTA below the fold+304%“Find Your Gym & Get Membership”+213%“your” changed to “my”+90%Same test, sister site, other language+38.3%Button colour, porcelain retailer+35.8%Added “Get started”+31.0%Added “View Bundle” text link+17.2%“Order” changed to “Get”+14.8%Font colour change−18.0%Removed a green arrow−12.3%Made the button bigger−10.6%Changing a button made things worse three times out of ten. That is the half worth knowing.
The listicles keep the +304% and drop the losses. The losses are the informative half. Source : Aagaard, ContentVerve, retrieved from the Internet Archive (2013)

Why 2,000 visits is not a sample

There is a serious quantitative treatment of this problem, and it changes how you read every number above.

The power requirement. A published analysis by a statistician working inside an A/B testing company puts the requirement near 6,000 conversion events per group to detect a 5% uplift at 80% power, and around 1,600 per group to detect a 10% uplift.

Note the unit. Conversion events, not visits. The famous button tests report visits in the low thousands, which is a different and much smaller quantity.

What under-powering does. The same analysis states that shortening a test that needs two months to two weeks makes it under-powered, and that almost two-thirds of winning tests will be completely bogus.

What peeking does, which is worse. Checking a test daily and stopping as soon as the p-value drops below 5% produced a positive result 41% of the time in a simulation where the two variants were identical. The conclusion drawn is that at least 80% of winning results obtained that way are worthless.

The base rate that frames all of it. The same paper cites data from a large search company suggesting a new variant is generally only 10% likely to cause a true uplift, with 90% having no effect or degrading performance.

And the multiple comparisons problem. Run twenty tests and you will on average get one winner even if none of the variants does anything.

The winner’s curse. Tests that do win tend to overstate the size of the effect, and the effect is stronger the smaller the sample. So the published 21% and the published 90% are, if anything, the inflated versions.

Statistical power requirements for split testing compared with what the canonical button tests reportedDiagram presenting the statistical power requirements for split testing established by a published quantitative analysis, alongside what the canonical button experiments actually reported, and the documented consequences of failing to meet those requirements. The analysis, produced by a statistician working within an experimentation company, establishes that detecting a five percent uplift at eighty percent power requires approximately six thousand conversion events per group, and detecting a ten percent uplift requires approximately one thousand six hundred conversion events per group. The unit is conversion events rather than visits, which matters because the canonical button experiments report visit counts in the low thousands, a different and considerably smaller quantity. The analysis further establishes three consequences of inadequate practice. First, shortening a test that requires two months of data collection to two weeks renders it under-powered, with the result that almost two-thirds of winning tests will be completely bogus. Second, the practice of checking a test daily and stopping as soon as the probability value falls below five percent produced a positive result forty-one percent of the time in a simulation where the two variants were identical, from which the analysis concludes that at least eighty percent of winning results obtained this way are worthless. Third, running twenty tests will on average produce one winning test even when none of the variants generates any uplift. The analysis also cites data from a large search company indicating that a new variant of a website is generally only ten percent likely to cause a true uplift, with the remaining ninety percent having no effect or degrading performance, and describes a winner’s curse whereby tests that do win systematically overstate the size of the effect, an effect that intensifies as sample size decreases.What a test needs, and what these reportedRequired, at 80% power~6,000 conversion events per groupto detect a 5% uplift~1,600 per group for a 10% upliftReported by the famous tests”over 2,000 visits” across both armsand no conversion count at allVisits are not conversion events.What happens when the requirement is not metCutting a two-month test to two weeks~2 in 3 winners are bogusChecking daily, stopping at p < 0.0541% “wins” on identical variantsThe winners you get that way≥ 80% worthlessRunning twenty testsOne winner, even if none worksThe base rate underneath all of itData from a large search company suggests a new variant is only about 10% likely to produce a trueuplift. The other 90% do nothing, or make things worse.
Conversion events per group, not visits. And stopping early manufactures winners on tests where nothing is happening. Source : Goodson, Most Winning A/B Test Results Are Illusory, Qubit (2014)

What is actually specified, and it is exact

Switch to the accessibility specification and the vagueness disappears. These are numbers, with conformance levels, that you can check without any traffic at all.

Text contrast, level AA. The visual presentation of text and images of text has a contrast ratio of at least 4.5:1. Large text may go to 3:1, where large means at least 18 point, or 14 point bold.

Text contrast, level AAA. 7:1, with large text at 4.5:1.

The button itself, level AA. Non-text contrast requires a ratio of at least 3:1 against adjacent colours for user interface components and graphical objects. This is the one people miss: it governs the button’s own border or fill, separately from the text sitting on it. Inactive components are exempt.

Target size, level AA. New in the current version: the target for pointer inputs is at least 24 by 24 CSS pixels, with five exceptions: spacing, an equivalent control elsewhere on the page, inline targets constrained by line height, targets whose size the browser controls, and cases where a particular presentation is essential or legally required.

Target size, level AAA. 44 by 44 CSS pixels, with only three exceptions: inline, user agent control, and essential.

Colour alone, level A. Colour is not used as the only visual means of conveying information, indicating an action, prompting a response, or distinguishing a visual element. Note what this does and does not say: it forbids relying on colour alone. It expresses no preference between colours.

The accessible name, level A. For components with labels that include text, the name contains the text presented visually. A button reading “Send” must not carry a completely different accessible name.

The normative accessibility requirements applying to a call-to-action button with their thresholds and conformance levelsTable of the accessibility success criteria that apply to a call-to-action button, with the exact threshold and conformance level of each, drawn from the published specification. Criterion one point four point three, contrast minimum, at conformance level double A, requires that the visual presentation of text and images of text has a contrast ratio of at least four point five to one, with large text permitted at three to one, where large text means at least eighteen point or fourteen point bold. Criterion one point four point six, contrast enhanced, at conformance level triple A, raises those figures to seven to one and four point five to one respectively. Criterion one point four point eleven, non-text contrast, at conformance level double A, requires a contrast ratio of at least three to one against adjacent colours for user interface components and graphical objects, which governs the button’s own border or fill separately from the text placed upon it, with inactive components exempted. Criterion two point five point eight, target size minimum, at conformance level double A and newly introduced in the current version of the specification, requires the target for pointer inputs to be at least twenty-four by twenty-four CSS pixels, subject to five exceptions covering spacing, an equivalent control elsewhere on the page, inline targets constrained by line height, targets whose size the user agent controls, and cases where a particular presentation is essential or legally required. Criterion two point five point five, target size enhanced, at conformance level triple A, requires forty-four by forty-four CSS pixels with only three exceptions. Criterion one point four point one, use of colour, at conformance level A, requires that colour is not used as the only visual means of conveying information, indicating an action, prompting a response or distinguishing a visual element, which forbids reliance on colour alone while expressing no preference between colours. Criterion two point five point three, label in name, at conformance level A, requires that for components with labels including text, the accessible name contains the text presented visually.What is actually specified for a buttonExact numbers, conformance levels, and no traffic required to check any of them.CriterionLevelThreshold1.4.3 Contrast (Minimum)AA4.5:1 text, 3:1 large textLarge = 18pt, or 14pt bold1.4.6 Contrast (Enhanced)AAA7:1 text, 4.5:1 large1.4.11 Non-text ContrastAA3:1 against adjacent coloursThe button’s own border or fill, separately from the text on it. The one people miss.2.5.8 Target Size (Minimum)AA24 x 24 CSS pixelsNew in 2.2. Five exceptions: spacing, equivalent, inline, user agent control, essential2.5.5 Target Size (Enhanced)AAA44 x 44 CSS pixels1.4.1 Use of ColorANever colour aloneNote what 1.4.1 does not say: it forbids relying on colour alone. It prefers no colour over another.
These are checkable in an afternoon, with no traffic and no statistics. The colour question is not among them. Source : W3C, Web Content Accessibility Guidelines 2.2 (2023)

Nobody has ever specified a converting colour

Worth stating plainly, because the search volume on this question is enormous and the answer is short.

No standards body says it. The accessibility specification governs contrast ratios, never hues. There is no criterion anywhere that prefers one colour to another.

No ergonomics standard says it. The relevant international standards on visual display cover contrast and legibility, not commercial performance.

No platform says it. None of the advertising platforms publishes guidance on button colour as a conversion lever.

And the best-known practitioner in the field says it does not exist. The author of the ten-case collection wrote that there are no set rules for which colours work, and that rules like “never use red, it’s a stop colour” or “green is always best, it’s a positive colour” are, in his words, plain stupid.

The likeliest explanation for the famous result. Red won on a site whose brand palette was green. That is a contrast effect, which is a property of the relationship between the button and its surroundings, not of the colour red.

What that means for your page. The useful question is not which colour converts. It is whether your primary action is the highest-contrast element on the page. That question has an answer, and you can measure it.

What to do instead

Six steps, ordered by how much they cost and how reliably they pay.

Meet the specified thresholds first. 4.5:1 on the text, 3:1 on the button itself, 24 by 24 pixels minimum, no meaning carried by colour alone, accessible name containing the visible label. This costs an afternoon, needs no traffic, and is the only part of this subject with defensible numbers.

Make the primary action the highest-contrast element on the page. One primary action, visually dominant. If two things compete for that role, neither is winning.

Write the button as the completion of a sentence. The visitor is finishing the thought “I want to…”. Buttons that read as an action the visitor is taking generally beat buttons that read as an instruction from you. This is a copywriting principle, not a tested law, and I am labelling it as such.

Stop testing buttons and start testing offers. Your traffic almost certainly cannot detect a 5% button effect, since that needs thousands of conversion events per group. It can detect the difference between two genuinely different offers, because that difference is large.

If you do test, decide the sample size and the end date before you start. Then do not look until you get there. Peeking is what turns a null result into a 41% chance of a false winner.

And record what you did. Most organisations rerun the same inconclusive test every eighteen months because nobody wrote down that it was inconclusive the first time.

Recommended allocation of effort between verifiable button requirements and unverifiable button experimentsDiagram allocating effort between the aspects of a call-to-action button that can be verified and those that cannot, in order of cost and reliability of return. The first and cheapest action is to meet the specified accessibility thresholds: a text contrast ratio of four point five to one, a non-text contrast ratio of three to one for the button itself, a minimum pointer target of twenty-four by twenty-four CSS pixels, no information conveyed by colour alone, and an accessible name containing the visible label; this work requires an afternoon, needs no traffic whatsoever, and constitutes the only portion of the subject supported by defensible numbers. The second action is to ensure the primary action is the highest-contrast element on the page, with a single visually dominant primary action, since two elements competing for that role means neither achieves it. The third action is to write the button text as the completion of the visitor’s own sentence beginning I want to, on the basis that buttons reading as an action the visitor takes generally outperform buttons reading as an instruction from the advertiser, a copywriting principle rather than a tested law. The fourth action is to stop testing buttons and begin testing offers, since detecting a five percent button effect requires thousands of conversion events per group whereas the difference between two genuinely distinct offers is large enough for ordinary traffic to detect. The fifth action, applicable when a test is run, is to fix the sample size and end date before beginning and to avoid inspecting results until that point, since interim checking converts a null result into a forty-one percent probability of a false winner. The sixth action is to record what was done, because organisations otherwise repeat the same inconclusive test approximately every eighteen months, nobody having recorded that it was inconclusive on the previous occasion.Spend the effort on the half you can verifyCosts an afternoon. Needs zero traffic. Has exact numbers.Meet the thresholds4.5:1, 3:1, 24x24, name contains label.One dominant actionHighest contrast element on the page.Costs a cycle. Your traffic can actually resolve it.Test offers, not buttonsThe difference is large enough to detect.Fix the end date in advanceThen do not look until you reach it.Costs months. Cannot resolve at your traffic. Skip it.Button colour. Button size. Arrow or no arrow. A 5% effect needs thousands of conversions per arm.And write down that it was inconclusive, so nobody reruns it in eighteen months.
The specified requirements cost an afternoon. Button colour tests cost months and cannot resolve at your traffic. Source : W3C WCAG 2.2 and Goodson, Qubit (2023)

Where to go next

You are deciding how much the form should ask. How many form fields.

Your ad and your page do not say the same thing. Message match.

You are choosing where to send paid traffic. Dedicated landing page or website page.

You want the page anatomy in detail. Anatomy of a high-converting B2B landing page.

Your bounce rate looks wrong. GA4 bounce rate.

Your conversion rate is the number in dispute. Conversion rate and its denominator.

In short

  • The red button test reports 2,000+ visits across both arms, no per-variant split, no conversion counts and no confidence interval.
  • Its own author warned against generalising, and told readers not to switch their buttons without testing first. That warning never travels with the statistic.
  • The $300 million button was a case study, not a test. The measured result was about $6 million in week one; the annual figure was calculated from abandoned carts.
  • The ten-case collection publishes no sample size for any case, by the author’s own explanation, and three of its ten results are negative.
  • A 5% uplift needs roughly 6,000 conversion events per group to detect at 80% power. Visits are not conversion events.
  • Stopping a test when it looks significant produced a winner 41% of the time on identical variants, and at least 80% of such winners are worthless.
  • A new variant is only about 10% likely to produce a true uplift in the first place.
  • What is specified is exact: 4.5:1 text contrast, 3:1 for the button itself, 24 by 24 pixels at AA and 44 by 44 at AAA, never colour alone, accessible name containing the visible label.

Fix the verifiable half this week, then test offers rather than buttons. Book a diagnostic, or see how we approach B2B websites.