Approving a Creative Direction Without Relying on Taste
Show one creative route and people judge it one way, show two and the judgment reverses. Asking them why they prefer one degrades the choice they make.
Show a decision-maker one creative direction and they judge it one way. Show them two and they judge the same one differently. This is not a failure of attention or seniority. It is a documented and replicable property of how preference judgments are formed, and the size of the shift is large enough to reverse a decision.
Which has a practical consequence, because most creative review is designed without any awareness of it. The room sees three routes on a screen, one person asks what everyone thinks, another explains why they prefer the second one, and a decision emerges from a process that the research says will produce a different answer if you rearrange it.
You cannot make the judgment objective. You can control the conditions under which it is made. Here is what the evidence supports, what it does not, and how to run the review.
The used dictionaries
The foundational demonstration is from 1996, and its numbers are worth memorizing because they are so clean.
116 students, randomly assigned to one of three questionnaires, were asked to imagine buying a used music dictionary in a bookstore, with a stated budget between $10 and $50. Two options existed:
Dictionary A: 1993, 10,000 entries, like new.
Dictionary B: 1993, 20,000 entries, torn cover.
Participants who saw both together offered $19 for A and $27 for B. Participants who saw only one offered $24 for A and $20 for B.
The preference reverses. Seen alone, the intact dictionary wins. Seen side by side, the damaged one with twice the content wins.
The explanation is that “10,000 entries” means nothing on its own. Nobody knows whether that is generous or thin for a music dictionary. But everybody knows what a torn cover is. Alone, you weight the attribute you can actually judge. Compared, the entry count becomes readable, and its weight rises.
One point of honesty that popular retellings drop. The gap in the separate condition, $24 against $20, is not significant at the conventional threshold. What is solidly established is the interaction: the difference changes sign depending on the mode. Say that preference reverses, not that people pay significantly more for A when they see it alone.
The dictionary result is famous. The follow-up studies in the same article are more useful, because they show the mechanism can be disabled.
Two televisions. 98 students judged sets differing on Clarity (90 against 40) and Warranty (9 against 18). When both numbers were meaningless, the hard and hard condition, no reversal occurred. When the warranty numbers were revealed to be months, making one attribute easy, the reversal appeared.
Two CD players. 202 students judged players differing on total harmonic distortion (.003 percent against .01 percent) and capacity (5 against 20 discs). In the hard and easy condition, the reversal appeared. Then the researchers simply added a sentence giving the reference range, telling participants that distortion runs from .002 percent at best to .012 percent at worst, and the reversal disappeared entirely.
One sentence of context removed the effect.
That is the actionable finding for creative review, and it does not require any change to how work is presented. The instability comes from attributes the viewer cannot evaluate in isolation. Supply the scale, and the instability drops.
In practice, the unreadable attributes in a design review are things like distinctiveness within the category, legibility at small sizes, production cost, and how a system behaves once extended. None of them can be judged from a single board. All of them can be made readable: show the competitive set, show the mark at 16 pixels, state the production constraint, show three more applications built from the same system.
The version of this that happens with real careers
Everything above uses students and hypothetical objects. In February 2026 a study appeared that does not.
Researchers analyzed 1,804 real promotion and tenure decisions across six universities, 906 made under joint evaluation, where multiple candidates were considered together, and 898 under separate evaluation, where each was considered alone. These are irreversible decisions about people’s careers, made by committees of professionals, in the field.
Under joint evaluation, Black and Hispanic faculty received on average 9 percent fewer negative departmental votes than under separate evaluation, controlling for research productivity, institution, gender, rank, discipline, department size and grant funding. Through moderated mediation, that corresponded to a 16.2 percent increase in the probability of promotion.
The follow-up is what makes this worth a paragraph in an article about design review. The researchers surveyed 289 professors who had actually served on promotion committees and asked what they expected. Only 17 percent thought joint evaluation would improve outcomes for under-represented candidates. 43 percent believed the opposite, that separate evaluation would help.
The people running these processes were, in the majority, wrong about the direction of the effect of the procedure they were using.
That is the general lesson, and it transfers cleanly. Nobody in a creative review has an intuition about how the format of the review shapes its outcome, and the intuitions people do have are unreliable. The format is a decision, and it is currently being made by accident.
The second body of evidence is more uncomfortable, because the instruction it argues against is the one every review meeting opens with.
The jams. 49 undergraduates tasted five strawberry jams that Consumer Reports had ranked using seven trained sensory panelists rating 16 characteristics across 45 products. The jams used ranked 1st, 11th, 24th, 32nd and 44th, so there was a real quality spread. Half the participants were first asked to “analyze why you feel the way you do”, and told the list would not be collected, with the experimenter visibly throwing it away. Final ratings went into a sealed box, collected by a second experimenter blind to condition.
Agreement with the trained panel, measured as a rank correlation:
Control group: 0.55, significantly above zero.
Reasons group: 0.11, not significantly different from zero.
Then the mechanism, which is the number to remember. The reasons people wrote were coded blind. The attitude implied by those reasons did not correlate with expert judgment. It correlated with the rating the person went on to give at a mean within-subject value of 0.92.
People generate mediocre reasons, and then obey them.
The authors ruled out the obvious alternative: analyzing reasons did not simply flatten the ratings toward the middle, since the within-subject range and standard deviation were essentially unchanged. It changed minds rather than producing hesitation.
The posters. The follow-up study is the one that matters for creative decisions, and its protocol is unusually careful. 43 undergraduates rated five posters: two art reproductions that pretested well, three humorous posters that pretested poorly. Half analyzed their reasons first, again with the sheet thrown away in front of them. The experimenter was blind to the hypotheses.
Then each participant was offered a poster to take home. The experimenter left the room, and the posters were rolled with the blank side out so nobody could see which had been taken, ruling out any desire to please. An average of 25 days later, a different experimenter, blind to condition, phoned to ask whether they still had it, whether it was on a wall, and how much they liked it now, verifying possession by asking them to read the manufacturer’s name off the bottom of the poster.
Choice: 95 percent of the control group took an art poster. 64 percent of the reasons group did.
Satisfaction 25 days later, on a standardized composite of three behavioral and two subjective measures: control +0.26, reasons negative 0.26.
Analyzing your reasons changed what people chose, and left them measurably less happy with it a month afterward.
The honest limit, and it is a real one. These experiments used students, jams and free posters. Nobody has tested reason-analysis on a high-stakes professional decision. The bridge from “explaining why you like a poster degrades your satisfaction” to “explaining why you like a creative route degrades the decision” is an extrapolation, not a result. It is a well-supported extrapolation, and it is still an extrapolation.
Expertise changes the criteria, not the reliability
The one study that transfers the evaluation-mode effect directly to aesthetic objects is worth knowing, including its weaknesses.
Chinese university students judged pairs of products differing on form and function. With USB drives, preference for the aesthetically stronger option rose from 3.65 in separate evaluation to 4.75 in joint evaluation, while preference for the functionally stronger option did not move at all. The mode shifted the aesthetic judgment specifically.
In a second experiment on basketball shoes, participants were split by an objective product-knowledge test. Novices preferred the aesthetic option (4.93 against 4.02). Experts preferred the functional one (4.17 against 4.83). The evaluation-mode effect disappeared among the high-knowledge group.
Two caveats the authors state or imply. They measured functional knowledge only, and explicitly flag as future work whether knowledge of form would produce the same immunity. And the samples were 80 and 101, on hypothetical products, in a second-tier journal. The only attempt to run this on actual design professionals returned a null result, and it was weak: evaluation mode was self-selected rather than randomized, and only 6 percent chose to see the options separately.
So the defensible reading is narrow. Expertise changes which attributes carry weight. It does not make anyone’s judgment correct in the abstract, and it does not mean the non-expert in the room should defer or the expert should overrule.
What does not replicate
An article like this is more credible for saying what fails, and two things fail.
Choice overload. The famous finding that more options reduce purchase and satisfaction does not hold up cleanly. A 2010 meta-analysis of 63 conditions across 50 experiments, 5,036 participants, found a mean effect of 0.02 with a confidence interval from negative 0.09 to 0.12. Trimming the extremes brought it to 0.001. Documented direct replication failures include the jam study itself, retried in a German supermarket with no negative effect. A 2015 meta-analysis of 99 observations reached the opposite conclusion and the authors acknowledged the contradiction. The defensible statement is that choice overload depends on conditions nobody has stabilized, and the average effect across conditions is near zero. Do not present three options because more would overwhelm anyone. Present three because three genuinely distinct routes exist.
Half of the evaluation-mode effect itself. A preregistered 2023 replication with 403 participants reproduced the “less is better in separate evaluation” side solidly. The “more is better in joint evaluation” side replicated weakly, and in one study the confidence interval included zero.
Present the options together. Joint evaluation is what the high-stakes field evidence supports, and it is what makes hard-to-judge attributes readable.
Supply the scale for every attribute that cannot be judged alone. The competitive set, the mark at 16 pixels, the production constraint, the system extended across three more applications. This is the single highest-leverage move, and one sentence was enough to eliminate the effect in the laboratory.
Do not open with “why do you like it”. Ask which one serves the brief, and require the answer to reference the brief. The written reasons people generate under pressure predict their own subsequent ratings almost perfectly and predict expert quality not at all.
Write the brief before anyone sees anything. A judgment needs a criterion outside itself. Without one, the review defaults to taste, and taste defaults to whoever has the most authority in the room. The brief that holds is the one drawn from a brand platform, where the position, the promise and the audience are settled before any route is drawn.
Make the options genuinely incompatible. If two routes could both be right, they are not a choice. The point of showing more than one is to make a decision visible and forcible.
Decide the format on purpose. That is the actual lesson of the promotion study. The format of an evaluation changes its outcome, most people running evaluations guess wrong about which way, and almost nobody chooses the format deliberately.
None of this makes an aesthetic decision objective. It makes it a decision, taken under conditions you chose, against a criterion you wrote down. That is all the research supports, and it is considerably more than most creative reviews currently have.
Frequently asked questions
Should I review creative directions one at a time or side by side?
Side by side, on the current evidence. Joint evaluation makes hard-to-judge attributes weigh more and, in a study of 1,804 real promotion decisions, reduced disparities that separate evaluation produced. Most decision-makers guess the opposite: only 17 percent of surveyed committee members expected joint evaluation to help.
Why does seeing two options change which one I prefer?
Because attributes differ in how easy they are to evaluate alone. Shown a single option, you cannot tell whether a number or a quality is good, so you weight what you can judge. Shown two, the comparison makes the hard attribute readable, and its weight rises.
Is it bad to explain why I like a direction?
In the published experiments, yes. Participants asked to analyze their reasons agreed less with trained experts, chose differently, and were less satisfied 25 days later. Note the honest limit: those studies used jams and posters, not professional decisions.
How many options should we present?
The evidence does not support a number. The choice overload effect fails to replicate: a meta-analysis of 63 conditions found a mean effect of 0.02 with a confidence interval spanning zero. Present enough options to make a real choice visible, which usually means two or three that are genuinely incompatible.
Does expertise fix any of this?
Partly. In the one direct study on aesthetic objects, the evaluation-mode effect disappeared among people with high product knowledge. But the authors measured functional knowledge only, so this does not establish that aesthetic expertise immunizes anyone.