Show a decision-maker one creative direction and they judge it one way. Show them two and they judge the same one differently. This is not a failure of attention or seniority. It is a documented and replicable property of how preference judgments are formed, and the size of the shift is large enough to reverse a decision.

Which has a practical consequence, because most creative review is designed without any awareness of it. The room sees three routes on a screen, one person asks what everyone thinks, another explains why they prefer the second one, and a decision emerges from a process that the research says will produce a different answer if you rearrange it.

You cannot make the judgment objective. You can control the conditions under which it is made. Here is what the evidence supports, what it does not, and how to run the review.

The used dictionaries

The foundational demonstration is from 1996, and its numbers are worth memorizing because they are so clean.

116 students, randomly assigned to one of three questionnaires, were asked to imagine buying a used music dictionary in a bookstore, with a stated budget between $10 and $50. Two options existed:

  • Dictionary A: 1993, 10,000 entries, like new.
  • Dictionary B: 1993, 20,000 entries, torn cover.

Participants who saw both together offered $19 for A and $27 for B. Participants who saw only one offered $24 for A and $20 for B.

The preference reverses. Seen alone, the intact dictionary wins. Seen side by side, the damaged one with twice the content wins.

The explanation is that “10,000 entries” means nothing on its own. Nobody knows whether that is generous or thin for a music dictionary. But everybody knows what a torn cover is. Alone, you weight the attribute you can actually judge. Compared, the entry count becomes readable, and its weight rises.

One point of honesty that popular retellings drop. The gap in the separate condition, $24 against $20, is not significant at the conventional threshold. What is solidly established is the interaction: the difference changes sign depending on the mode. Say that preference reverses, not that people pay significantly more for A when they see it alone.

Willingness to pay for two used dictionaries under joint and separate evaluationWillingness to pay for two used music dictionaries under joint and separate evaluation, from a nineteen ninety six experiment with one hundred and sixteen unpaid students at the University of Chicago and the University of Illinois at Chicago, randomly assigned to one of three questionnaires presenting either both options together or one option alone. Participants imagined buying a used music dictionary in a bookstore with a stated budget between ten and fifty dollars. Dictionary A was described as dated nineteen ninety three, containing ten thousand entries, and being like new. Dictionary B was described as dated nineteen ninety three, containing twenty thousand entries, and having a torn cover. Under joint evaluation, where both were seen together, participants offered nineteen dollars for Dictionary A and twenty seven dollars for Dictionary B, a difference significant at a t value of seven point one one with a probability below point zero zero one. Under separate evaluation, where each participant saw only one dictionary, participants offered twenty four dollars for Dictionary A and twenty dollars for Dictionary B, a difference with a t value of one point six nine and a probability of point one zero, which is not significant at the conventional five percent threshold. The reversal itself, meaning the change in the sign of the difference between modes, is significant at a t value of four point five six with a probability below point zero zero one. The explanation offered is that the number of entries carries no meaning in isolation, since nobody knows whether ten thousand entries is generous or thin for a music dictionary, whereas everybody can evaluate a torn cover, so in isolation people weight the attribute they can judge, while comparison makes the entry count readable and raises its weight.Same two options, opposite winnersSeen togetherA: 10,000 entries, like new$19B: 20,000 entries, torn cover$27B winsSeen one at a timeA: 10,000 entries, like new$24B: 20,000 entries, torn cover$20A wins”20,000 entries” means nothing on its own. A torn cover means something to everyone.Alone, you weight what you can judge. Compared, the unreadable attribute becomes readable.
116 students, randomly assigned. The reversal is the finding. The gap within the separate condition is not significant on its own. Source : Hsee, The evaluability hypothesis, Organizational Behavior and Human Decision Processes 67(3), 1996 (1996)

The effect has an off switch

The dictionary result is famous. The follow-up studies in the same article are more useful, because they show the mechanism can be disabled.

Two televisions. 98 students judged sets differing on Clarity (90 against 40) and Warranty (9 against 18). When both numbers were meaningless, the hard and hard condition, no reversal occurred. When the warranty numbers were revealed to be months, making one attribute easy, the reversal appeared.

Two CD players. 202 students judged players differing on total harmonic distortion (.003 percent against .01 percent) and capacity (5 against 20 discs). In the hard and easy condition, the reversal appeared. Then the researchers simply added a sentence giving the reference range, telling participants that distortion runs from .002 percent at best to .012 percent at worst, and the reversal disappeared entirely.

One sentence of context removed the effect.

That is the actionable finding for creative review, and it does not require any change to how work is presented. The instability comes from attributes the viewer cannot evaluate in isolation. Supply the scale, and the instability drops.

In practice, the unreadable attributes in a design review are things like distinctiveness within the category, legibility at small sizes, production cost, and how a system behaves once extended. None of them can be judged from a single board. All of them can be made readable: show the competitive set, show the mark at 16 pixels, state the production constraint, show three more applications built from the same system.

How supplying a reference scale eliminates the joint versus separate preference reversalHow supplying a reference scale eliminates the preference reversal between joint and separate evaluation, demonstrated in two follow up studies within the same nineteen ninety six article. In the television study, ninety eight students judged sets differing on a clarity index of ninety against forty and a warranty figure of nine against eighteen. When both numbers were meaningless to participants, described as the hard and hard condition, no reversal occurred at all. When the warranty numbers were revealed to represent months, making one attribute easy to evaluate while the other remained hard, the reversal appeared with a t value of three point four seven and a probability below point zero one. In the compact disc player study, two hundred and two students judged players differing on total harmonic distortion of point zero zero three percent against point zero one percent and on capacity of five against twenty discs. In the hard and easy condition the reversal appeared with a t value of three point three two and a probability below point zero one. The researchers then added a single sentence giving the reference range, telling participants that total harmonic distortion runs from point zero zero two percent at best to point zero one two percent at worst, which converted the hard attribute into an easy one, and the reversal disappeared entirely with a t value below one and no statistical significance. The practical implication for a creative review is that the instability arises from attributes the viewer cannot evaluate in isolation, such as distinctiveness within the competitive set, legibility at small sizes, production cost, and how a system behaves when extended, and that each of these can be made readable by showing the competitive set, showing the mark at sixteen pixels, stating the production constraint, and showing further applications built from the same system.Make the unreadable attribute readableBoth attributes hardno reversalTwo numbers nobodycan interpretOne hard, one easyreversalt = 3.47 and t = 3.32across the two studiesBoth made easyno reversalOne sentence gave thereference rangeThe unreadable attributes in a design review, and how to supply the scaleDistinctiveness: show the competitive set. Legibility: show the mark at 16 pixels.Production cost: state the constraint. Extensibility: show three more applications from the same system.None of these can be judged from one board. All of them can be made judgeable.
One sentence giving the range made a meaningless number readable, and the reversal disappeared. That is the lever in a design review. Source : Hsee, Organizational Behavior and Human Decision Processes 67(3), 1996, Studies 3 and 4 (1996)

The version of this that happens with real careers

Everything above uses students and hypothetical objects. In February 2026 a study appeared that does not.

Researchers analyzed 1,804 real promotion and tenure decisions across six universities, 906 made under joint evaluation, where multiple candidates were considered together, and 898 under separate evaluation, where each was considered alone. These are irreversible decisions about people’s careers, made by committees of professionals, in the field.

Under joint evaluation, Black and Hispanic faculty received on average 9 percent fewer negative departmental votes than under separate evaluation, controlling for research productivity, institution, gender, rank, discipline, department size and grant funding. Through moderated mediation, that corresponded to a 16.2 percent increase in the probability of promotion.

The follow-up is what makes this worth a paragraph in an article about design review. The researchers surveyed 289 professors who had actually served on promotion committees and asked what they expected. Only 17 percent thought joint evaluation would improve outcomes for under-represented candidates. 43 percent believed the opposite, that separate evaluation would help.

The people running these processes were, in the majority, wrong about the direction of the effect of the procedure they were using.

That is the general lesson, and it transfers cleanly. Nobody in a creative review has an intuition about how the format of the review shapes its outcome, and the intuitions people do have are unreliable. The format is a decision, and it is currently being made by accident.

Effect of joint versus separate evaluation across one thousand eight hundred and four real promotion decisionsEffect of joint versus separate evaluation across one thousand eight hundred and four real promotion and tenure decisions at six universities, published in Nature Communications in February twenty twenty six as a natural experiment on irreversible high stakes decisions made by professional committees in the field rather than in a laboratory. Nine hundred and six of the decisions were made under joint evaluation, in which multiple candidates were considered together, and eight hundred and ninety eight under separate evaluation, in which each candidate was considered alone. Under joint evaluation, Black and Hispanic faculty received on average nine percent fewer negative votes at the departmental level than under separate evaluation, controlling for research productivity, institution, gender, rank, discipline, department size and grant funding. Moderated mediation analyses indicated that this corresponded to a sixteen point two percent increase in the probability of promotion. The researchers additionally surveyed two hundred and eighty nine professors who had served on promotion committees and asked what effect they expected. Only seventeen percent anticipated that joint evaluation would improve outcomes for under represented minority candidates, while forty three percent believed the opposite, namely that separate evaluation would help. The majority of the people operating these processes were therefore wrong about the direction of the effect of the procedure they were using. The transferable lesson is that the format of an evaluation is itself a decision with measurable consequences, that intuitions about which format helps are unreliable, and that in most organizations this format is currently being chosen by accident rather than deliberately.Not a laboratory. Actual promotion committees.Decisions analyzed1,804906 joint, 898 separate, six universitiesFewer negative votes underjoint evaluation9%a 16.2% rise in promotion probabilityCommittee members whoexpected this result17%43% expected the oppositeWhy this belongs in an article about creative reviewThe format of an evaluation changes its outcome, and the people running it mostly do not know which way.Controls applied: research productivity, institution, gender, rank, discipline, department size, grant funding.Survey base: 289 professors who had actually served on promotion committees.
Real decisions, real careers, six institutions. And the committee members surveyed mostly expected the opposite result. Source : Masters-Waage et al., Evaluating multiple candidates simultaneously reduces racial disparities in promotion and tenure, Nature Communications 17, 3080, 23 February 2026 (2026)

Do not ask people why

The second body of evidence is more uncomfortable, because the instruction it argues against is the one every review meeting opens with.

The jams. 49 undergraduates tasted five strawberry jams that Consumer Reports had ranked using seven trained sensory panelists rating 16 characteristics across 45 products. The jams used ranked 1st, 11th, 24th, 32nd and 44th, so there was a real quality spread. Half the participants were first asked to “analyze why you feel the way you do”, and told the list would not be collected, with the experimenter visibly throwing it away. Final ratings went into a sealed box, collected by a second experimenter blind to condition.

Agreement with the trained panel, measured as a rank correlation:

  • Control group: 0.55, significantly above zero.
  • Reasons group: 0.11, not significantly different from zero.

Then the mechanism, which is the number to remember. The reasons people wrote were coded blind. The attitude implied by those reasons did not correlate with expert judgment. It correlated with the rating the person went on to give at a mean within-subject value of 0.92.

People generate mediocre reasons, and then obey them.

The authors ruled out the obvious alternative: analyzing reasons did not simply flatten the ratings toward the middle, since the within-subject range and standard deviation were essentially unchanged. It changed minds rather than producing hesitation.

The posters. The follow-up study is the one that matters for creative decisions, and its protocol is unusually careful. 43 undergraduates rated five posters: two art reproductions that pretested well, three humorous posters that pretested poorly. Half analyzed their reasons first, again with the sheet thrown away in front of them. The experimenter was blind to the hypotheses.

Then each participant was offered a poster to take home. The experimenter left the room, and the posters were rolled with the blank side out so nobody could see which had been taken, ruling out any desire to please. An average of 25 days later, a different experimenter, blind to condition, phoned to ask whether they still had it, whether it was on a wall, and how much they liked it now, verifying possession by asking them to read the manufacturer’s name off the bottom of the poster.

  • Choice: 95 percent of the control group took an art poster. 64 percent of the reasons group did.
  • Satisfaction 25 days later, on a standardized composite of three behavioral and two subjective measures: control +0.26, reasons negative 0.26.

Analyzing your reasons changed what people chose, and left them measurably less happy with it a month afterward.

The honest limit, and it is a real one. These experiments used students, jams and free posters. Nobody has tested reason-analysis on a high-stakes professional decision. The bridge from “explaining why you like a poster degrades your satisfaction” to “explaining why you like a creative route degrades the decision” is an extrapolation, not a result. It is a well-supported extrapolation, and it is still an extrapolation.

Effects of asking people to analyze their reasons before making a preference judgmentEffects of asking people to analyze their reasons before making a preference judgment, from two experiments. In the first, published in nineteen ninety one, forty nine undergraduates tasted five strawberry jams that a consumer magazine had ranked using seven trained sensory panelists who rated sixteen characteristics across forty five products, with the five jams used ranking first, eleventh, twenty fourth, thirty second and forty fourth. Half the participants were first instructed to analyze why they felt the way they did, were told the list would not be collected, and watched the experimenter throw it away, while final ratings went into a sealed box collected by a second experimenter blind to condition. Agreement with the trained panel, measured as a rank correlation, was zero point five five in the control group, significantly above zero, and zero point one one in the reasons group, not significantly different from zero. The written reasons were then coded blind, and the attitude implied by those reasons did not correlate with expert judgment but correlated with the rating the participant subsequently gave at a mean within subject value of zero point nine two, indicating that people generate reasons and then obey them. In the second experiment, published in nineteen ninety three, forty three undergraduates rated five posters consisting of two art reproductions that pretested well and three humorous posters that pretested poorly, with half analyzing their reasons first. Each participant was then offered a poster to take home, with the experimenter leaving the room and the posters rolled blank side out so the choice was unobserved. Ninety five percent of the control group took an art poster against sixty four percent of the reasons group. An average of twenty five days later a different experimenter, blind to condition, telephoned and verified possession by asking participants to read the manufacturer name from the bottom of the poster. On a standardized composite of three behavioral and two subjective satisfaction measures, the control group scored plus zero point two six and the reasons group minus zero point two six.Two experiments, one resultAgreement with a trained sensory panel, five strawberry jamsControlr = 0.55Asked to analyze reasonsr = 0.110.11 is not significantly different from zero. Their own written reasons predicted their ratings at r = 0.92.Poster taken home, and satisfaction 25 days laterControl: chose an art poster95%Reasons: chose an art poster64%Satisfaction composite after an average of 25 dayscontrol +0.26reasons -0.26Honest limit: students, jams and free posters. Nobody has tested this on a high-stakes professional decision.
Agreement with trained experts collapsed, and satisfaction a month later went negative. Both studies used low-stakes objects. Source : Wilson and Schooler, Journal of Personality and Social Psychology 60(2), 1991; Wilson, Lisle, Schooler, Hodges, Klaaren and LaFleur, Personality and Social Psychology Bulletin 19(3), 1993 (1993)

Expertise changes the criteria, not the reliability

The one study that transfers the evaluation-mode effect directly to aesthetic objects is worth knowing, including its weaknesses.

Chinese university students judged pairs of products differing on form and function. With USB drives, preference for the aesthetically stronger option rose from 3.65 in separate evaluation to 4.75 in joint evaluation, while preference for the functionally stronger option did not move at all. The mode shifted the aesthetic judgment specifically.

In a second experiment on basketball shoes, participants were split by an objective product-knowledge test. Novices preferred the aesthetic option (4.93 against 4.02). Experts preferred the functional one (4.17 against 4.83). The evaluation-mode effect disappeared among the high-knowledge group.

Two caveats the authors state or imply. They measured functional knowledge only, and explicitly flag as future work whether knowledge of form would produce the same immunity. And the samples were 80 and 101, on hypothetical products, in a second-tier journal. The only attempt to run this on actual design professionals returned a null result, and it was weak: evaluation mode was self-selected rather than randomized, and only 6 percent chose to see the options separately.

So the defensible reading is narrow. Expertise changes which attributes carry weight. It does not make anyone’s judgment correct in the abstract, and it does not mean the non-expert in the room should defer or the expert should overrule.

What does not replicate

An article like this is more credible for saying what fails, and two things fail.

Choice overload. The famous finding that more options reduce purchase and satisfaction does not hold up cleanly. A 2010 meta-analysis of 63 conditions across 50 experiments, 5,036 participants, found a mean effect of 0.02 with a confidence interval from negative 0.09 to 0.12. Trimming the extremes brought it to 0.001. Documented direct replication failures include the jam study itself, retried in a German supermarket with no negative effect. A 2015 meta-analysis of 99 observations reached the opposite conclusion and the authors acknowledged the contradiction. The defensible statement is that choice overload depends on conditions nobody has stabilized, and the average effect across conditions is near zero. Do not present three options because more would overwhelm anyone. Present three because three genuinely distinct routes exist.

Half of the evaluation-mode effect itself. A preregistered 2023 replication with 403 participants reproduced the “less is better in separate evaluation” side solidly. The “more is better in joint evaluation” side replicated weakly, and in one study the confidence interval included zero.

Findings in the preference literature that fail to replicate cleanlyFindings in the preference and choice literature that fail to replicate cleanly, and which should therefore not be cited as established. The first is choice overload, the widely repeated claim that offering more options reduces purchase and satisfaction. A meta analysis published in the Journal of Consumer Research in twenty ten covering sixty three conditions from fifty published and unpublished experiments with five thousand and thirty six participants found a mean effect of zero point zero two with a ninety five percent confidence interval running from negative zero point zero nine to zero point one two, and after trimming twenty percent of the most extreme studies the effect fell to zero point zero zero one with heterogeneity dropping from sixty eight percent to twenty two percent. Documented direct replication failures include an attempt to repeat the original jam tasting study in a German upmarket supermarket, which found no negative effect, a laboratory attempt with exotic chocolates that found no difference, and a further attempt using jelly beans that also failed. A competing meta analysis published in the Journal of Consumer Psychology in twenty fifteen, covering ninety nine observations from fifty three studies in twenty one articles with seven thousand two hundred and two participants, reached the opposite conclusion, and its authors explicitly acknowledged the contradiction and attributed it to the subset of studies retained and the conceptual model tested. The defensible statement is therefore that choice overload depends on conditions that have not been stabilized and that the average effect across conditions is near zero. The second finding that does not fully hold is one half of the evaluation mode effect itself, since a preregistered replication published in twenty twenty three with four hundred and three participants reproduced the less is better pattern in separate evaluation solidly, with effect sizes of zero point nine nine, zero point three two and zero point seven six across three studies, but reproduced the more is better pattern in joint evaluation only weakly, with one study returning an effect size of zero point zero nine whose confidence interval included zero.What not to cite as establishedChoice overloadMeta-analysis of 63 conditions, 50 experiments, 5,036 participants: mean effect 0.02,confidence interval from -0.09 to 0.12. Trimmed to 0.001.A competing 2015 meta-analysis reached the opposite conclusion and said so.Half of the evaluation-mode effectPreregistered replication, N = 403. “Less is better” in separate evaluation replicatessolidly at d = 0.99, 0.32 and 0.76 across three studies.”More is better” in joint evaluation replicates weakly, one study at d = 0.09 with aconfidence interval including zero.
Cite the parts that hold. The choice overload effect averages near zero, and half of the evaluation-mode effect replicates weakly. Source : Scheibehenne, Greifeneder and Todd, Journal of Consumer Research 37(3), 2010; Chernev, Bockenholt and Goodman, Journal of Consumer Psychology 25(2), 2015; Vonasch et al., Collabra: Psychology 9(1), 2023 (2023)

How to run the review

Six things follow, and all of them are free.

Present the options together. Joint evaluation is what the high-stakes field evidence supports, and it is what makes hard-to-judge attributes readable.

Supply the scale for every attribute that cannot be judged alone. The competitive set, the mark at 16 pixels, the production constraint, the system extended across three more applications. This is the single highest-leverage move, and one sentence was enough to eliminate the effect in the laboratory.

Do not open with “why do you like it”. Ask which one serves the brief, and require the answer to reference the brief. The written reasons people generate under pressure predict their own subsequent ratings almost perfectly and predict expert quality not at all.

Write the brief before anyone sees anything. A judgment needs a criterion outside itself. Without one, the review defaults to taste, and taste defaults to whoever has the most authority in the room. The brief that holds is the one drawn from a brand platform, where the position, the promise and the audience are settled before any route is drawn.

Make the options genuinely incompatible. If two routes could both be right, they are not a choice. The point of showing more than one is to make a decision visible and forcible.

Decide the format on purpose. That is the actual lesson of the promotion study. The format of an evaluation changes its outcome, most people running evaluations guess wrong about which way, and almost nobody chooses the format deliberately.

None of this makes an aesthetic decision objective. It makes it a decision, taken under conditions you chose, against a criterion you wrote down. That is all the research supports, and it is considerably more than most creative reviews currently have.