Your identity is decided. Your image is observed. That distinction sounds like a slogan until you try to observe the image, at which point it becomes an expensive methodological problem that most brand tracking quietly ignores.
The useful thing is that the profession which sells this measurement publishes its own standards, and they are strict, public and free. Read against them, the typical brand perception exercise fails at the first item. Not because anyone is dishonest, but because the failure modes are counterintuitive: the sample that feels representative is not, the bigger survey is often the worse one, and the number you print with a plus or minus sign usually should not have one.
Here is what those standards actually require, and what a company of ordinary size can honestly say afterward.
Eleven things you have to disclose
The American Association for Public Opinion Research code, revised April 2021, requires that eleven items be disclosed at the moment results are released, publicly or to a client. Not on request. At release.
They are: the data collection strategy; who sponsored and who conducted the research; the measurement instruments; the population under study; the method used to generate and recruit the sample; the methods and modes of data collection; the dates of data collection; sample sizes with any discussion of precision; how the data were weighted; how the data were processed and the procedures used to ensure quality; and a general statement acknowledging limitations.
Two of those deserve quoting in full. On sampling:
“Explicitly state whether the sample comes from a frame selected using a probability-based methodology (meaning selecting potential participants with a known non-zero probability from a known frame) or if the sample was selected using non-probability methods (potential participants from opt-in, volunteer, or other sources).”
On limitations:
“All research has limitations and researchers will include a general statement acknowledging the unmeasured error associated with all forms of public opinion research.”
And the part that closes the usual escape route, from the association’s own guidance: this applies “whether or not the individuals are AAPOR members”, and “All research firms, organizations and individuals are obligated to make the minimally required disclosures, upon request, for any publicly released survey results.” Proprietary methodology is not an exemption.
If you commission brand research, this list is your acceptance criteria. Ask for the eleven items before the fieldwork, not after the presentation.
”The response rate is X” is not an acceptable sentence
This one is quotable verbatim and settles a lot of arguments. From the tenth edition of Standard Definitions, 2023:
“In calculating and reporting outcome rates according to the rules and formulas below, researchers must precisely define the rates used. For example, a statement that ‘the response rate is X’ is unacceptable. One must report exactly which rate was used, such as ‘Response Rate 2 was X.’”
There are six numbered response rates. They differ on exactly two things: whether partial interviews count as respondents, and how cases of unknown eligibility are handled in the denominator. Response Rate 1 is described as “the minimum response rate”. Response Rate 6 “represents the maximum response rate”. Response Rate 3, which estimates what share of unknown cases were eligible, is “the most-common AAPOR response rate in reporting”.
Same fieldwork, six defensible numbers, and one of them is always the flattering one. Which is why the document also states, on the estimation step: “One must not select a proportion to boost the response rate. The basis for the estimate must be explicitly stated and detailed.”
If a vendor gives you a response rate without a number after the words “Response Rate”, you have not been told the response rate.
You probably cannot publish a margin of error
This is the point that most often surprises people who commission brand research. If your respondents opted in, volunteered, came from a panel, answered a link in your newsletter, or were recruited from your own customer list, you do not have a probability sample. And the association’s position on that is unambiguous:
“Currently, it is impossible to develop statistically valid margins of sampling error from nonprobability surveys, such as opt-in, online polls.”
Its explainer document adds: “the margin of sampling error only applies to probability-based surveys where participants have a known and non-zero chance of being included in the sample. It does not apply to opt-in online surveys and other non-probability based polls.” And a sentence that applies to every survey ever run: “There is no such thing as a measurable overall margin of error for a poll.”
That does not license the opposite error. The 2013 task force report is explicit: “Treating estimates as though they had no error at all is not a reasonable option.”
So what is permitted? The code answers precisely, and it is a high bar:
“Reports of non-probability sample surveys will only provide measures of precision if they are defined and accompanied by a detailed description of how the underlying model was specified, its assumptions validated, and the measure(s) calculated.”
You may publish a plus or minus figure if you also publish the model behind it and the validation of its assumptions. Almost nobody does, which is why almost nobody should be printing the figure.
Two further notes worth holding. For non-probability samples, the correct term for the share who answered is a participation rate, not a response rate, and the standards document explains why: “a high response rate does not necessarily mean the risk of bias is reduced.” And the association’s own condemned practices list includes “Representing the results of a self-selected ‘poll’ as if they were the outcome of legitimate research.”
More responses do not fix a bad sample
The most useful single study on this was published in Nature in December 2021, and it settles an argument that comes up in every brand tracking discussion.
Two very large surveys estimated first-dose COVID-19 vaccine uptake among United States adults. One collected about 250,000 responses per week. The other about 75,000 every two weeks. A third panel collected about 1,000 per week using conventional survey research practice.
Against the benchmark, the 250,000-response survey overestimated uptake by 17 percentage points. The 75,000-response survey by 14 points. The 1,000-response panel produced reliable estimates.
The authors call the mechanism the Big Data Paradox, and their summary is worth quoting at length because every clause matters:
“Increasing data size shrinks confidence intervals but magnifies the effect of survey bias… Moreover, their large sample sizes led to miniscule margins of error on the incorrect estimates… We show how a survey of 250,000 respondents can produce an estimate of the population mean that is no more accurate than an estimate from a simple random sample of size 10. Our central message is that data quality matters more than data quantity, and that compensating the former with the latter is a mathematically provable losing proposition.”
And the formulation that belongs on the wall of anyone who buys brand research: “when biased samples are large, they are doubly misleading: they produce confidence intervals with incorrect centres and substantially underestimated widths.”
For scale on the ordinary arithmetic, the published sampling error table, at 95 percent confidence and assuming a genuine random sample: 100 responses gives about plus or minus 9.8 points; 500 gives 4.4; 1,000 gives 3.1; 2,000 gives about 2.0. As the association notes, “doubling the sample size from 1,000 to 2,000 only reduces the margin of sampling error by about a single percentage point.”
The subgroup warning is the one that bites in practice. A 1,000-person survey carries plus or minus 3.1 overall, but a 200-person segment inside it carries plus or minus 6.9. And nobody commissions a brand tracker to look at the total.
The question you ask is part of the answer
The Pew Research Center runs split-sample experiments: two versions of a question, respondents randomly assigned to one or the other, so any difference in the answers is attributable to the wording. The published results are larger than most people expect.
- Asked whether they favored military action in Iraq to end Saddam Hussein’s rule, 68 percent were in favor. Adding the clause “even if it meant that U.S. forces might suffer thousands of casualties” moved that to 43 percent. A 25 point swing from one clause.
- “Give terminally ill patients the means to end their lives” drew 51 percent support. “Assist terminally ill patients in committing suicide” drew 44 percent. 7 points from two synonyms.
- Asked whether “jobs” were available in their community, 60 percent said yes. Asked about “good jobs”, 48 percent. 12 points from one adjective.
Question order does the same work. Support for legal agreements for same-sex couples was 45 percent when asked after the marriage question and 37 percent when asked before, an 8 point difference from position alone. In another experiment, agreement that Republican leaders should work with the incoming president was 81 percent when asked second and 66 percent when asked first, a 15 point gap.
The consequence for brand tracking is the one sentence to take away from all of it:
“The order questions are asked is of particular importance when tracking trends over time. As a result, care should be taken to ensure that the context is similar each time a question is asked. Modifying the context of the question could call into question any observed changes over time.”
Two further design points, both sourced. Agree or disagree statements invite acquiescence: “less educated and less informed respondents have a greater tendency to agree with such statements”, and the effect is stronger with an interviewer present, so “a better practice is to offer respondents a choice between alternative statements.” And the international standard on brand evaluation notes about its own recommended instrument that “the advantage of a rating scale lies in its ease of use, however, respondents often tend to assume extreme or middle positions.”
Aided and unaided awareness are not substitutes
A widely repeated claim needs correcting, because it leads companies to switch measures when the numbers look thin.
A 1995 article in Marketing Science showed that aided, spontaneous and top-of-mind awareness scores across more than 39 categories are related by a logistic transformation, and argued that they behave like one latent trait measured at three difficulty levels. That was read to mean the three are interchangeable.
A 2004 replication reproduced the fit and then rejected the inference:
“while there is a good category level fit, modelling a single brand over time is less successful. Indeed, Laurent et al.’s excellent cross-sectional fit appears due to substantially different levels of salience between larger and smaller brands. This suggests that while the different types of awareness tend to vary with a brand’s overall level of salience, this does not mean that the different measures simply reflect a single underlying construct. Further, our finding challenges the previous authors’ claim that knowing the score for one measure allows the estimation of the score for another measure.”
The tidy relationship across a category comes from big brands scoring higher on everything. It does not hold for one brand tracked over time. Pick a measure, and keep it.
For a sense of how large the prompted versus unprompted gap can be, a Pew experiment offers a clean illustration on a different subject. Asked in closed form what issue mattered most in their vote, 58 percent chose the economy from a list of five. Asked the same question open-ended, only 35 percent volunteered it, and 43 percent gave an answer that was not on the list at all.
What a company of ordinary size can honestly claim
Nonresponse is where most brand surveys quietly fail, and its shape is instructive. In a benchmarking exercise on telephone surveys running at about a 9 percent response rate, the average absolute error across 13 demographic, lifestyle and health questions was 2.7 percentage points. On one civic participation question, the error was 38 points: 8 percent in the government benchmark against 46 percent in the survey.
Nonresponse bias is not a property of the response rate. It is a property of the correlation between the topic and the willingness to answer. Which is exactly the problem when you survey your own customer list about your own brand: the people who answer are the people who care about you.
So here is the defensible position, and it is narrower than what most brand trackers claim but genuinely useful.
You can measure change in your own measured population over time, provided the method is frozen. Same wording, same order, same mode, same recruitment source, same season. The instrument does not have to be representative to detect movement. It has to be identical. It helps to know what ought to be moving, since recognition, perceived consistency and clarity of message are the things a brand platform and a visual identity are built to change.
You cannot produce a population estimate. Your number does not describe your market. It describes the people who answered you.
You cannot attach a margin of error, unless you publish the model and its validation.
You cannot compare your number to an industry benchmark computed by a different method, on a different frame, with different wording.
And two things you can do cheaply that are endorsed in the published best practice. Pretest the questionnaire through cognitive interviews with people resembling your respondents, before fielding it. And where you are unsure how to word something, run the two versions as a randomized experiment inside your own fieldwork, which is the same technique that produced every number in the section above.
The honest summary is that a mid-size company does not own a measurement of its market. It owns a consistent instrument pointed at a slice of it. That is worth having, and it is worth describing accurately, because the first time the brand score and the pipeline disagree, the accuracy of your original claim is what determines whether anyone still believes the number.