One controlled study on web forms is worth more than the rest of the literature combined, and it reports first-attempt success rising from 42 percent to 78 percent. Sixty-five participants, three real registration forms from live company websites, each tested in its original version and in a version rewritten to comply with twenty published guidelines.

That is the good news. The complication is that it tested the twenty guidelines together, so it cannot tell you which of them mattered. And when you go looking for evidence behind the individual guidelines, most of them do not have any.

This article separates the two, because a form that loses meetings is worth fixing on the basis of the study that exists rather than the fourteen that are quoted as if they did.

The study that carries the weight

Worth reading in detail, because its design is unusually clean and its limitations are unusually clear.

The method. An eye-tracking laboratory study, between subjects, 65 participants with a mean age of 27.5, all self-rated experienced internet users. Thirty-two were assigned to the original-forms condition and 33 to the improved condition, with groups matched on age, education and computer knowledge.

The material. Registration forms from three real company websites, each in its original state and in a version rewritten to satisfy twenty published guidelines.

The results, form by form, on first attempt. On the first form, 10 of 32 succeeded originally against 24 of 33 improved. On the second, 9 of 32 against 24 of 33. On the third, 22 of 33 against 28 of 32. All three differences were statistically significant.

Aggregated, that is 42 percent against 78 percent.

Completion time is the honest footnote. It improved significantly on two of the three forms and not on the third, where the difference did not reach significance. Reporting the aggregate without that detail overstates the finding.

And the limitation that matters most for practice. The study applied twenty guidelines together. It establishes that following published form guidance produces a large improvement. It does not establish that any particular guideline is responsible, and it cannot be cited for one.

Results of a controlled study comparing original and guideline compliant versions of three web formsResults of a controlled eye tracking laboratory study comparing original registration forms taken from three real company websites against versions of the same forms rewritten to comply with twenty published form design guidelines. The study used a between subjects design with sixty five participants, forty two of whom were female, with a mean age of twenty seven point five years and a standard deviation of nine point seven, ranging from eighteen to sixty seven, all self rating as experienced internet users. Thirty two participants were assigned to the original forms condition and thirty three to the improved condition, with the groups matched on age, education, computer and web knowledge and internet usage. On first attempt submission success, the first form recorded ten successes from thirty two participants in the original condition against twenty four from thirty three in the improved condition, a difference significant at a chi squared value of eleven point two zero with a probability below zero point zero zero one. The second form recorded nine from thirty two against twenty four from thirty three, significant at a chi squared value of twelve point nine three with a probability below zero point zero zero one. The third form recorded twenty two from thirty three against twenty eight from thirty two, significant at a chi squared value of three point two nine with a probability of zero point zero three five. Aggregated across all three forms, forty one successes from ninety seven attempts in the original condition and seventy six from ninety eight in the improved condition, that is forty two point three percent against seventy seven point six percent. Completion time improved significantly on two of the three forms with effect sizes of one point zero zero and zero point nine three, but the difference on the third form did not reach significance with an effect size of zero point three eight, a detail that must accompany any citation of the aggregate. The principal limitation for practical application is that the study applied twenty guidelines simultaneously and therefore establishes that following published form guidance produces a large improvement without establishing which individual guideline is responsible.First-attempt success, original against improvedForm 1, original10 / 32Form 1, improved24 / 33Form 2, original9 / 32Form 2, improved24 / 33Form 3, original22 / 33Form 3, improved28 / 32Aggregated42% to 78%on first attempt, 65 participantsThe limitation to carry with itTwenty guidelines applied together. It cannot tellyou which one did the work.Completion time improved significantly on two forms of three. On the third it did not reach significance.Quoting the aggregate without that detail overstates the result.
Sixty-five participants, three live forms, twenty guidelines applied together. The effect is large and the attribution is not available. Source : Seckler, Heinz, Bargas-Avila, Opwis and Tuch, Designing Usable Web Forms, CHI 2014 (2014)

Which individual guidelines have evidence

This is the uncomfortable inventory, and it is worth doing because it tells you where to spend attention.

Label placement, the most repeated advice, has the weakest source. The article everyone cites reports saccade times of 500 milliseconds for left-aligned labels, 170 to 240 for right-aligned, and 50 for labels above the field. It states no participant count and identifies no equipment.

And a peer-reviewed study reached the opposite conclusion. An eye-tracking study published in 2008 found right-aligned labels produced the shortest completion times, at least for multi-column forms. Its sample size is not publicly available either, so the honest summary is that the question is open rather than settled.

Field count has a benchmark but not a causal result. A checkout study reports “the average checkout flow in 2024 is 5.1 steps long and contains 11.3 form fields”, against a stated view that “most sites need only 8 form fields in total”, with the trend falling from 12.7 in 2019. That is descriptive. It does not say that removing fields raises conversion.

Inline validation has a real study, and it is small. Twenty-two participants across six form variations, with the best-performing variant showing a 22 percent higher success rate, 22 percent fewer errors and 42 percent faster completion. No significance tests are reported, and it is industry research.

And several familiar guidelines have nothing behind them. Published guidance on placeholders, on marking required fields, on single-column layout and on error message specificity carries no cited study or sample size in its source. That does not make the advice wrong. It makes it advice.

One guideline does have a hard number, from an unusual source. A dataset of 2,130 conference participants found “99.9% of city names were 19 characters or shorter, making 19 characters a reasonable width for a city field.”

Evidence status of commonly repeated web form design guidelinesThe evidence status of commonly repeated guidelines for web form design, graded by the strength of the published source behind each. The strongest evidence supports the general proposition that following published form guidance improves outcomes, resting on a controlled between subjects eye tracking study with sixty five participants comparing three real forms against guideline compliant rewrites, which found first attempt success rising from forty two to seventy eight percent with all three individual comparisons statistically significant, while noting that twenty guidelines were applied together so that no individual guideline can be isolated. Moderate evidence supports inline validation, resting on an industry study of twenty two participants across six form variations in which the best performing variant produced a twenty two percent higher success rate, twenty two percent fewer errors, thirty one percent higher satisfaction, forty two percent faster completion and forty seven percent fewer eye fixations, with no significance tests reported and a small sample per condition. Descriptive benchmark data exists on field counts, an electronic commerce study reporting an average checkout flow of five point one steps containing eleven point three form fields in 2024, down from eleven point eight in 2021 and twelve point seven in 2019, against a stated view that most such flows require only eight fields, this being a count rather than a causal conversion finding. Contested evidence surrounds label placement, since the most widely cited source reports saccade times of five hundred milliseconds for left aligned labels, one hundred and seventy to two hundred and forty milliseconds for right aligned labels and fifty milliseconds for labels above the field, while stating no participant count and identifying no equipment, and a separate peer reviewed eye tracking study published in 2008 found right aligned labels produced the shortest completion times for multi column forms. Absent evidence characterises several familiar guidelines, since published guidance on placeholder text, on marking required fields, on single column layout and on error message specificity carries no cited study or sample size in its source, which does not make the advice wrong but does make it advice rather than research.What is evidence, and what is conventionGuidelineWhat is behind itFollow published form guidance generallyControlled study, 65 participants, 3 real forms42% to 78% first-attempt successInline validation22 participants, 6 variations, industry-runNo significance tests reportedFewer fieldsA benchmark: 11.3 fields average, 8 said to sufficeDescriptive, not causalLabels above the fieldMost-cited source states no sample and no equipmentA peer-reviewed study found the oppositePlaceholders, required marks, single columnNo cited study or sample size in the sourceNone of this makes the advice wrong. It makes it advice, which is a different thing to cite and a different thing tospend three weeks arguing about.
One controlled study, one small industry study, one benchmark, and a set of guidelines with nothing published behind them. Source : Seckler et al. CHI 2014; Wroblewski 2009; Baymard 2024; and the guidance sources checked (2024)

The requirements that are not advice

Between the study and the folklore sits a set of rules with actual force, and they cover most of what goes wrong on a booking flow.

Every field that asks about the person must declare its purpose. At AA, the input purpose must be programmatically determinable, which means correct autocomplete tokens: name, organization, organization-title, email, tel.

Nothing may be asked twice in one process. At Level A: “Information previously entered by or provided to the user that is required to be entered again in the same process is either: auto-populated, or available for the user to select.”

Errors must be text. At Level A, the item in error must be identified “and the error is described to the user in text.” A red outline is not a description.

And a correction must be offered where one is known. At AA, if suggestions for correction “are known, then the suggestions are provided to the user.”

The placeholder question is settled by the HTML specification rather than by usability opinion. “The placeholder attribute should not be used as an alternative to a label.”

Targets have a minimum. 24 by 24 CSS pixels at AA, which catches small radio buttons, tight date pickers and the close control on a booking modal.

One that specifically affects scheduling widgets. Any function using a dragging movement must be achievable “by a single pointer without dragging”, which reaches time-range selectors and drag-to-pick calendars.

Form construction requirements carrying normative force as distinct from advisory guidanceThe requirements governing form construction that carry normative force under published standards, as distinct from the advisory guidance that dominates discussion of form design, together with the conformance level of each. At double A, the purpose of each input field collecting information about the user must be programmatically determinable, which in practice requires correct automatic completion tokens such as those for name, organisation, job title, email address and telephone number. At single A, information previously entered or provided by the user which is required to be entered again within the same process must be either automatically populated or available for the user to select, subject to exceptions where re entry is essential, where security requires it, or where the earlier information is no longer valid. At single A, where an input error is automatically detected the item in error must be identified and the error described to the user in text, so that a coloured outline alone does not satisfy the requirement because an outline is not a description. At double A, where an input error is detected and suggestions for correction are known, those suggestions must be provided unless doing so would jeopardise the security or purpose of the content. Under the markup living standard rather than the accessibility guidelines, the placeholder attribute should not be used as an alternative to a label, this being normative prose in a different document from the one usually cited for it. At double A, the size of a target for pointer input must be at least twenty four by twenty four stylesheet pixels subject to five exceptions, which catches small radio controls, tightly spaced date pickers and the close control on a booking modal. At double A, all functionality using a dragging movement must be achievable by a single pointer without dragging unless dragging is essential, which reaches time range selectors and drag to select calendars in scheduling widgets.Seven rules that are not opinionsRequirementSource and levelEvery field about the person declares its purpose (autocomplete)WCAG, AANothing asked twice in the same processWCAG, AErrors described in text, not by colourWCAG, ACorrection suggestions offered where knownWCAG, AAPlaceholder is not a labelHTML StandardTargets at least 24 by 24 CSS pixelsWCAG, AAAny dragging has a single-pointer alternativeWCAG, AAThe last one reaches scheduling widgets: time-range selectors and drag-to-pick calendars both engage it.
Seven requirements from two standards. None of them needs a test, and most booking flows fail at least three. Source : W3C WCAG 2.2 and the WHATWG HTML Standard (2026)

What actually breaks a booking flow

The failures that lose meetings tend to be structural rather than cosmetic, which is why they belong to the design of the path from first visit to booked meeting rather than to a later round of copy changes.

Asking for qualification information before the person has decided to talk. Company size, budget and timeline on the first screen convert a booking into an application.

Requiring an account. The most common self-inflicted step, and the one furthest from the visitor’s goal.

Hiding the calendar behind a form. The visitor came to see whether a time exists that works. Making them earn the sight of it inverts the order.

Not saying what happens next. Duration, who will be on the call, and what to bring. Absent, the visitor supplies the worst assumption.

Losing the entry when a session expires. Covered by an AAA criterion for a reason: “when an authenticated session expires, the user can continue the activity without loss of data after re-authenticating.”

And the one nobody notices until it costs a deal. No confirmation that survives the browser. An email with the time, the link and a way to reschedule, sent immediately, is the whole of it.

Structural failures in a meeting booking flow and the expectation each one invertsSix structural failures commonly found in the flow from a website visit to a booked meeting, none of which concerns colour, copy or layout, and each of which inverts an order the visitor expects. The first is requesting qualification information before the visitor has decided to hold a conversation, so that questions about company size, budget and timeline placed on the first screen convert a booking into an application and are answered by fewer people. The second is requiring the creation of an account, which is the most common self inflicted additional step and the step furthest removed from the visitor’s actual goal. The third is placing the calendar behind a form, when the visitor’s purpose in arriving was to establish whether a workable time exists, so that requiring them to earn sight of the availability reverses the sequence of the decision. The fourth is omitting a statement of what happens next, specifically the duration of the meeting, who will attend from the seller’s side and what the visitor should bring or prepare, in the absence of which the visitor supplies the least favourable assumption. The fifth is losing the entered data when a session expires, which the accessibility standard addresses at its highest conformance level by requiring that when an authenticated session expires the user can continue the activity without loss of data after re authenticating. The sixth is the absence of a confirmation that survives the closing of the browser, the remedy being an immediate email containing the time, the joining link and a means of rescheduling. Alongside these structural failures sit a set of requirements with actual force rather than advisory status, comprising correct automatic completion tokens declaring the purpose of each field about the person, a prohibition on requesting the same information twice within one process, a requirement that errors be described in text rather than indicated by colour alone, a requirement that correction suggestions be offered where known, a minimum pointer target size of twenty four by twenty four stylesheet pixels, and a requirement that any dragging interaction such as a time range selector be achievable by a single pointer without dragging.Six failures that lose meetingsQualification questions before the decision to talkCompany size, budget and timeline on screen one turn a booking into an application.Requiring an accountThe step furthest from what the visitor came to do.The calendar hidden behind the formThey came to see whether a time exists. Make them earn it and the order is inverted.No statement of what happens nextDuration, who attends, what to bring. Absent, the visitor assumes the worst version.Losing the entry when the session dies”The user can continue the activity without loss of data after re-authenticating.”No confirmation that survives the browserAn email with the time, the link and a way to reschedule, sent immediately. That is the whole of it.
Six failures, none of which is a colour or a headline. Each inverts the order the visitor expects. Source : Method, over the form criteria and the controlled study cited above (2026)
Comparison of the sequence a visitor expects in a booking flow with the sequence commonly implementedComparison between the sequence of steps a visitor expects when booking a meeting and the sequence most commonly implemented on business websites, which reverses two of them at material cost. The sequence a visitor expects begins with establishing whether a suitable time exists, since determining availability is the reason they arrived at the booking page at all. It continues with understanding what the meeting involves, meaning its duration, who will attend from the seller’s side and what preparation is expected. It then proceeds to selecting a time, then to providing the minimum contact details required to hold that time, and concludes with receiving a confirmation that persists outside the browser session in the form of an email containing the time, the joining link and a means of rescheduling. The sequence commonly implemented reverses the first two stages against the last three, beginning instead with a form requesting contact details, frequently accompanied by qualification questions covering company size, budget and timeline, and in some cases by a requirement to create an account, before any availability is displayed. The consequence is that the visitor is asked to commit personal information and to answer questions appropriate to a later stage of the relationship before they have been able to establish the single fact that brought them to the page. The remedy requires no experimentation and no additional development beyond reordering, namely displaying availability first, stating what the meeting involves alongside it, collecting only the minimum contact details necessary at the point of selection, moving all qualification questions to after the booking is confirmed where they will still be answered but by people who have already committed to the conversation, and sending an immediate confirmation email that survives the closing of the browser.Two steps, in the wrong orderWhat the visitor expects1. Does a workable time exist?2. What does the meeting involve?3. Pick a time4. Give the minimum details5. Get a confirmation that persistsWhat most flows do1. Fill in a form2. Answer qualification questions3. Sometimes: create an account4. Now see whether a time exists5. Hope for an emailThe fix costs nothing to build: show availability first, and move qualification to after the booking is confirmed.You still get the answers, from people who have already decided to talk.This is a reordering, not a hypothesis. It does not need traffic to justify it.
Reversing two steps costs nothing to build and removes the most common reason people leave. Source : Method (2026)

What to do with this

Take your booking flow and count the steps between arriving and seeing an available time. If the answer is more than one, that is the change to make first, and it does not need a test.

Then apply the requirements that are not opinions: correct autocomplete tokens, nothing asked twice, errors written as text, targets at 24 pixels, and no drag-only interactions in the scheduler.

Move every qualification question after the booking is confirmed. You will still get the answers, from people who have already committed to the conversation rather than from people deciding whether to.

And treat the rest as convention rather than evidence. Label alignment, placeholder style and single-column layout are all worth a decision and none is worth a fortnight, because the published sources behind them do not carry the weight the argument assumes.

The related pieces are landing page structures by objective and B2B website usability that is testable.