Your identity is decided. Your image is observed. That distinction sounds like a slogan until you try to observe the image, at which point it becomes an expensive methodological problem that most brand tracking quietly ignores.

The useful thing is that the profession which sells this measurement publishes its own standards, and they are strict, public and free. Read against them, the typical brand perception exercise fails at the first item. Not because anyone is dishonest, but because the failure modes are counterintuitive: the sample that feels representative is not, the bigger survey is often the worse one, and the number you print with a plus or minus sign usually should not have one.

Here is what those standards actually require, and what a company of ordinary size can honestly say afterward.

Eleven things you have to disclose

The American Association for Public Opinion Research code, revised April 2021, requires that eleven items be disclosed at the moment results are released, publicly or to a client. Not on request. At release.

They are: the data collection strategy; who sponsored and who conducted the research; the measurement instruments; the population under study; the method used to generate and recruit the sample; the methods and modes of data collection; the dates of data collection; sample sizes with any discussion of precision; how the data were weighted; how the data were processed and the procedures used to ensure quality; and a general statement acknowledging limitations.

Two of those deserve quoting in full. On sampling:

“Explicitly state whether the sample comes from a frame selected using a probability-based methodology (meaning selecting potential participants with a known non-zero probability from a known frame) or if the sample was selected using non-probability methods (potential participants from opt-in, volunteer, or other sources).”

On limitations:

“All research has limitations and researchers will include a general statement acknowledging the unmeasured error associated with all forms of public opinion research.”

And the part that closes the usual escape route, from the association’s own guidance: this applies “whether or not the individuals are AAPOR members”, and “All research firms, organizations and individuals are obligated to make the minimally required disclosures, upon request, for any publicly released survey results.” Proprietary methodology is not an exemption.

If you commission brand research, this list is your acceptance criteria. Ask for the eleven items before the fieldwork, not after the presentation.

The eleven items requiring immediate disclosure when survey results are releasedThe eleven items that the American Association for Public Opinion Research code of professional ethics and practice, as revised in April twenty twenty one, requires to be disclosed at the time results are released, whether publicly or to a research client. The items are, in order: the data collection strategy; who sponsored the research and who conducted it; the measurement tools and instruments; the population under study; the method used to generate and recruit the sample; the methods and modes of data collection; the dates of data collection; the sample sizes broken down by sampling frame if more than one frame was used, together with any discussion of the precision of the results; how the data were weighted; how the data were processed and the procedures used to ensure data quality; and a general statement acknowledging limitations of the design and data collection. The item on sample generation requires the researcher to explicitly state whether the sample comes from a frame selected using a probability based methodology, meaning selecting potential participants with a known non zero probability from a known frame, or whether the sample was selected using non probability methods drawing participants from opt in, volunteer or other sources. The item on limitations states that all research has limitations and that researchers will include a general statement acknowledging the unmeasured error associated with all forms of public opinion research. The obligation applies to researchers and research companies whether or not the individuals concerned are members of the association, and all research firms, organizations and individuals are obligated to make the minimally required disclosures upon request for any publicly released survey results, so proprietary methodology is not an exemption.Disclosed at release, not on request1Data collection strategy2Who sponsored the research and who conducted it3Measurement tools and instruments4Population under study5Method used to generate and recruit the sample, stating probability or non-probability6Methods and modes of data collection7Dates of data collection8Sample sizes by frame, and any discussion of precision9How the data were weighted10How the data were processed, and quality assurance procedures11A general statement acknowledging limitations
Not on request. At release. Use this list as acceptance criteria before commissioning brand research. Source : AAPOR Code of Professional Ethics and Practice, Section III, revised April 2021 (2021)

”The response rate is X” is not an acceptable sentence

This one is quotable verbatim and settles a lot of arguments. From the tenth edition of Standard Definitions, 2023:

“In calculating and reporting outcome rates according to the rules and formulas below, researchers must precisely define the rates used. For example, a statement that ‘the response rate is X’ is unacceptable. One must report exactly which rate was used, such as ‘Response Rate 2 was X.’”

There are six numbered response rates. They differ on exactly two things: whether partial interviews count as respondents, and how cases of unknown eligibility are handled in the denominator. Response Rate 1 is described as “the minimum response rate”. Response Rate 6 “represents the maximum response rate”. Response Rate 3, which estimates what share of unknown cases were eligible, is “the most-common AAPOR response rate in reporting”.

Same fieldwork, six defensible numbers, and one of them is always the flattering one. Which is why the document also states, on the estimation step: “One must not select a proportion to boost the response rate. The basis for the estimate must be explicitly stated and detailed.”

If a vendor gives you a response rate without a number after the words “Response Rate”, you have not been told the response rate.

You probably cannot publish a margin of error

This is the point that most often surprises people who commission brand research. If your respondents opted in, volunteered, came from a panel, answered a link in your newsletter, or were recruited from your own customer list, you do not have a probability sample. And the association’s position on that is unambiguous:

“Currently, it is impossible to develop statistically valid margins of sampling error from nonprobability surveys, such as opt-in, online polls.”

Its explainer document adds: “the margin of sampling error only applies to probability-based surveys where participants have a known and non-zero chance of being included in the sample. It does not apply to opt-in online surveys and other non-probability based polls.” And a sentence that applies to every survey ever run: “There is no such thing as a measurable overall margin of error for a poll.”

That does not license the opposite error. The 2013 task force report is explicit: “Treating estimates as though they had no error at all is not a reasonable option.”

So what is permitted? The code answers precisely, and it is a high bar:

“Reports of non-probability sample surveys will only provide measures of precision if they are defined and accompanied by a detailed description of how the underlying model was specified, its assumptions validated, and the measure(s) calculated.”

You may publish a plus or minus figure if you also publish the model behind it and the validation of its assumptions. Almost nobody does, which is why almost nobody should be printing the figure.

Two further notes worth holding. For non-probability samples, the correct term for the share who answered is a participation rate, not a response rate, and the standards document explains why: “a high response rate does not necessarily mean the risk of bias is reduced.” And the association’s own condemned practices list includes “Representing the results of a self-selected ‘poll’ as if they were the outcome of legitimate research.”

What may and may not be reported as a measure of precision from a non probability surveyWhat may and may not be reported as a measure of precision from a non probability survey, according to three published positions of the American Association for Public Opinion Research. The first position, from a statement issued in twenty twelve, is that it is currently impossible to develop statistically valid margins of sampling error from non probability surveys such as opt in online polls, and a companion explainer adds that the margin of sampling error only applies to probability based surveys where participants have a known and non zero chance of being included in the sample and does not apply to opt in online surveys and other non probability based polls, while also noting that there is no such thing as a measurable overall margin of error for any poll since surveys are subject to other errors that cannot be measured. The second position, from the twenty thirteen task force report on non probability sampling, is that treating estimates as though they had no error at all is not a reasonable option, so the correct response to the first position is not to publish a bare number with no uncertainty attached. The third position, from section three A of the association code as revised in twenty twenty one, resolves the tension by stating that reports of non probability sample surveys will only provide measures of precision if they are defined and accompanied by a detailed description of how the underlying model was specified, its assumptions validated and the measures calculated. In practice this means a company may publish a plus or minus figure on an opt in sample only if it simultaneously publishes the statistical model behind that figure and the validation of the model’s assumptions. Two related points follow. For a non probability sample the correct term for the share of invitees who answered is a participation rate rather than a response rate, since a high response rate does not necessarily mean the risk of bias is reduced. And the association’s condemned practices list includes representing the results of a self selected poll as if they were the outcome of legitimate research.Three published positions, one workable conclusion2012 statement”it is impossible to develop statistically valid margins of sampling error from nonprobability surveys”2013 task force”Treating estimates as though they had no error at all is not a reasonable option.”2021 code, Section III.A”will only provide measures of precision if they are defined and accompanied by a detaileddescription of how the underlying model was specified, its assumptions validated”
Three positions, all published by the same body. The middle one is the one that gets skipped. Source : AAPOR statement on credibility intervals, 2012; AAPOR Task Force on Non-Probability Sampling, 2013; AAPOR Code Section III.A, 2021 (2013)

More responses do not fix a bad sample

The most useful single study on this was published in Nature in December 2021, and it settles an argument that comes up in every brand tracking discussion.

Two very large surveys estimated first-dose COVID-19 vaccine uptake among United States adults. One collected about 250,000 responses per week. The other about 75,000 every two weeks. A third panel collected about 1,000 per week using conventional survey research practice.

Against the benchmark, the 250,000-response survey overestimated uptake by 17 percentage points. The 75,000-response survey by 14 points. The 1,000-response panel produced reliable estimates.

The authors call the mechanism the Big Data Paradox, and their summary is worth quoting at length because every clause matters:

“Increasing data size shrinks confidence intervals but magnifies the effect of survey bias… Moreover, their large sample sizes led to miniscule margins of error on the incorrect estimates… We show how a survey of 250,000 respondents can produce an estimate of the population mean that is no more accurate than an estimate from a simple random sample of size 10. Our central message is that data quality matters more than data quantity, and that compensating the former with the latter is a mathematically provable losing proposition.”

And the formulation that belongs on the wall of anyone who buys brand research: “when biased samples are large, they are doubly misleading: they produce confidence intervals with incorrect centres and substantially underestimated widths.”

For scale on the ordinary arithmetic, the published sampling error table, at 95 percent confidence and assuming a genuine random sample: 100 responses gives about plus or minus 9.8 points; 500 gives 4.4; 1,000 gives 3.1; 2,000 gives about 2.0. As the association notes, “doubling the sample size from 1,000 to 2,000 only reduces the margin of sampling error by about a single percentage point.”

The subgroup warning is the one that bites in practice. A 1,000-person survey carries plus or minus 3.1 overall, but a 200-person segment inside it carries plus or minus 6.9. And nobody commissions a brand tracker to look at the total.

Sampling error by sample size, and evidence that sample size does not correct selection biasSampling error by sample size at ninety five percent confidence assuming a genuine simple random sample, together with evidence that increasing sample size does not correct selection bias. The published sampling error figures are as follows. Fifty responses give plus or minus thirteen point eight percentage points. One hundred responses give plus or minus nine point eight points. Two hundred give plus or minus six point nine. Four hundred give plus or minus four point nine. Five hundred give plus or minus four point four. Seven hundred give plus or minus three point seven. One thousand give plus or minus three point one. Fifteen hundred give plus or minus two point five. Two thousand five hundred give plus or minus two point zero. Five thousand give plus or minus one point four. The association notes that doubling the sample size from one thousand to two thousand only reduces the margin of sampling error by about a single percentage point, and warns that the overall sampling error applies to the total sample and not to subgroups, so that a two hundred person segment inside a one thousand person survey carries plus or minus six point nine points rather than plus or minus three point one. A peer reviewed article published in Nature in December twenty twenty one demonstrates that none of this corrects bias. Two large surveys estimating first dose vaccine uptake among United States adults, one collecting about two hundred and fifty thousand responses per week and another about seventy five thousand every two weeks, overestimated uptake by seventeen and fourteen percentage points respectively, while a panel of about one thousand responses per week following conventional survey research practice produced reliable estimates. The authors state that increasing data size shrinks confidence intervals but magnifies the effect of survey bias, that the large sample sizes led to minuscule margins of error on the incorrect estimates, and that a survey of two hundred and fifty thousand respondents can produce an estimate no more accurate than a simple random sample of size ten.Precision flattens. Bias does not.Margin of sampling error at 95 percent confidence, assuming a true random samplen = 5013.8 ptsn = 1009.8n = 2006.9n = 4004.9n = 5004.4n = 1,0003.1n = 2,5002.0n = 5,0001.4And on the other axis, from a 2021 Nature articleA survey of 250,000 responses overestimated vaccine uptake by 17 points. One of 75,000 by 14 points.A panel of about 1,000 done properly was reliable. “a survey of 250,000 respondents can produce anestimate of the population mean that is no more accurate than an estimate from a simple random sample of size 10.”
Precision improves sharply up to about 1,000 and then flattens. None of it corrects a sample that was wrong to begin with. Source : AAPOR, Margin of Sampling Error explainer; Bradley et al., Nature, 8 December 2021 (2021)

The question you ask is part of the answer

The Pew Research Center runs split-sample experiments: two versions of a question, respondents randomly assigned to one or the other, so any difference in the answers is attributable to the wording. The published results are larger than most people expect.

  • Asked whether they favored military action in Iraq to end Saddam Hussein’s rule, 68 percent were in favor. Adding the clause “even if it meant that U.S. forces might suffer thousands of casualties” moved that to 43 percent. A 25 point swing from one clause.
  • “Give terminally ill patients the means to end their lives” drew 51 percent support. “Assist terminally ill patients in committing suicide” drew 44 percent. 7 points from two synonyms.
  • Asked whether “jobs” were available in their community, 60 percent said yes. Asked about “good jobs”, 48 percent. 12 points from one adjective.

Question order does the same work. Support for legal agreements for same-sex couples was 45 percent when asked after the marriage question and 37 percent when asked before, an 8 point difference from position alone. In another experiment, agreement that Republican leaders should work with the incoming president was 81 percent when asked second and 66 percent when asked first, a 15 point gap.

The consequence for brand tracking is the one sentence to take away from all of it:

“The order questions are asked is of particular importance when tracking trends over time. As a result, care should be taken to ensure that the context is similar each time a question is asked. Modifying the context of the question could call into question any observed changes over time.”

Two further design points, both sourced. Agree or disagree statements invite acquiescence: “less educated and less informed respondents have a greater tendency to agree with such statements”, and the effect is stronger with an interviewer present, so “a better practice is to offer respondents a choice between alternative statements.” And the international standard on brand evaluation notes about its own recommended instrument that “the advantage of a rating scale lies in its ease of use, however, respondents often tend to assume extreme or middle positions.”

Measured effects of question wording and question order in randomized split sample experimentsMeasured effects of question wording and question order on survey answers, from randomized split sample experiments in which two versions of a question are fielded to randomly assigned halves of the same sample so that any difference is attributable to the version. On wording, asking whether respondents favored or opposed taking military action in Iraq to end Saddam Hussein’s rule produced sixty eight percent in favor, while adding the clause even if it meant that United States forces might suffer thousands of casualties produced forty three percent in favor, a difference of twenty five percentage points. Asking about giving terminally ill patients the means to end their lives produced fifty one percent in favor, while asking about assisting terminally ill patients in committing suicide produced forty four percent, a difference of seven points from two near synonyms. Asking whether jobs were available in the respondent’s community produced sixty percent, while asking about good jobs produced forty eight percent, a difference of twelve points from a single adjective. On order, support for legal agreements for same sex couples was forty five percent when the question followed the same sex marriage question and thirty seven percent when it preceded it, a difference of eight points from position alone. Agreement that Republican leaders should work with the incoming president was eighty one percent when asked after the mirror question about Democratic leaders and sixty six percent when asked first, a difference of fifteen points. The published guidance states that the order questions are asked is of particular importance when tracking trends over time, that care should be taken to ensure the context is similar each time a question is asked, and that modifying the context of the question could call into question any observed changes over time.The same people, two wordingsMilitary action in Iraqwith the casualties clause added68%43%25 pts”jobs” available in your community”good jobs” available60%48%12 pts”the means to end their lives""assist in committing suicide”51%44%7 ptsThe same wording, two positions”Republican leaders should work withthe president”, asked secondthe same question, asked first81%66%
All figures from randomized split-sample experiments on the same population. The wording is not a neutral container. Source : Pew Research Center, Writing Survey Questions (2019)

Aided and unaided awareness are not substitutes

A widely repeated claim needs correcting, because it leads companies to switch measures when the numbers look thin.

A 1995 article in Marketing Science showed that aided, spontaneous and top-of-mind awareness scores across more than 39 categories are related by a logistic transformation, and argued that they behave like one latent trait measured at three difficulty levels. That was read to mean the three are interchangeable.

A 2004 replication reproduced the fit and then rejected the inference:

“while there is a good category level fit, modelling a single brand over time is less successful. Indeed, Laurent et al.’s excellent cross-sectional fit appears due to substantially different levels of salience between larger and smaller brands. This suggests that while the different types of awareness tend to vary with a brand’s overall level of salience, this does not mean that the different measures simply reflect a single underlying construct. Further, our finding challenges the previous authors’ claim that knowing the score for one measure allows the estimation of the score for another measure.”

The tidy relationship across a category comes from big brands scoring higher on everything. It does not hold for one brand tracked over time. Pick a measure, and keep it.

For a sense of how large the prompted versus unprompted gap can be, a Pew experiment offers a clean illustration on a different subject. Asked in closed form what issue mattered most in their vote, 58 percent chose the economy from a list of five. Asked the same question open-ended, only 35 percent volunteered it, and 43 percent gave an answer that was not on the list at all.

What a company of ordinary size can honestly claim

Nonresponse is where most brand surveys quietly fail, and its shape is instructive. In a benchmarking exercise on telephone surveys running at about a 9 percent response rate, the average absolute error across 13 demographic, lifestyle and health questions was 2.7 percentage points. On one civic participation question, the error was 38 points: 8 percent in the government benchmark against 46 percent in the survey.

Nonresponse bias is not a property of the response rate. It is a property of the correlation between the topic and the willingness to answer. Which is exactly the problem when you survey your own customer list about your own brand: the people who answer are the people who care about you.

So here is the defensible position, and it is narrower than what most brand trackers claim but genuinely useful.

You can measure change in your own measured population over time, provided the method is frozen. Same wording, same order, same mode, same recruitment source, same season. The instrument does not have to be representative to detect movement. It has to be identical. It helps to know what ought to be moving, since recognition, perceived consistency and clarity of message are the things a brand platform and a visual identity are built to change.

You cannot produce a population estimate. Your number does not describe your market. It describes the people who answered you.

You cannot attach a margin of error, unless you publish the model and its validation.

You cannot compare your number to an industry benchmark computed by a different method, on a different frame, with different wording.

Claims a mid-size company may and may not make from its own brand perception surveyClaims that a company of ordinary size may and may not legitimately make on the basis of a brand perception survey run on its own contacts or on an opt in panel. The one available claim is a measurement of change within its own measured population over time, and it holds only where the method is frozen, meaning the same question wording, the same question order, the same mode of data collection, the same recruitment source and the same time of year at every wave, because published randomized experiments show wording effects of up to twenty five percentage points and order effects of up to fifteen points on the same population, and the published guidance warns that modifying the context of a question could call into question any observed changes over time. The instrument does not need to be representative in order to detect movement, but it does need to be identical. Three claims are not available. The company cannot produce a population estimate, because the number describes the people who answered rather than the market, and nonresponse bias is a property of the correlation between the topic and the willingness to answer, which is why surveying your own customer list about your own brand oversamples the people who care about you. The company cannot attach a margin of sampling error, because the association responsible for these standards states that it is impossible to develop statistically valid margins of sampling error from non probability surveys, and its code permits a measure of precision only when accompanied by a detailed description of how the underlying model was specified, its assumptions validated and the measure calculated. And the company cannot compare its figure to an industry benchmark computed by a different method on a different frame with different wording.One claim survivesChange in your own measured population over timeAvailable, on one condition: the method is frozen. Same wording, same order, same mode,same recruitment source, same season. Identical, not representative.A population estimateYour number describes thepeople who answered, notyour marketA margin of errorNot on a non-probabilitysample, unless you publishthe model and its validationA benchmark comparisonDifferent method, differentframe, different wording,different numberWording effects of up to 25 points and order effects of up to 15 points are why “frozen method” is the whole condition.
Four claims, one of which is available. The instrument does not have to be representative to detect movement. It has to be identical each time. Source : AAPOR Code Section III.A and Task Force on Non-Probability Sampling; Pew Research Center, Writing Survey Questions (2021)

And two things you can do cheaply that are endorsed in the published best practice. Pretest the questionnaire through cognitive interviews with people resembling your respondents, before fielding it. And where you are unsure how to word something, run the two versions as a randomized experiment inside your own fieldwork, which is the same technique that produced every number in the section above.

The honest summary is that a mid-size company does not own a measurement of its market. It owns a consistent instrument pointed at a slice of it. That is worth having, and it is worth describing accurately, because the first time the brand score and the pipeline disagree, the accuracy of your original claim is what determines whether anyone still believes the number.