Ninety percent of voters preferred the concept we did not ship. That is not an argument against asking people what they think. It is an argument for knowing, before you ask, what their answer is allowed to decide.
The question of how to choose between website design concepts usually gets answered badly in one of two directions. One habit treats audience preference as noise for a designer's taste to override. The other treats a vote as a verdict and hands a branding decision to whoever happened to be scrolling that afternoon. Both skip the same step: naming what evidence you are collecting and what that evidence can settle.
A design poll is better at measuring immediate drama than long-term brand fit

A design poll is better at measuring immediate drama than long-term brand fit. Voters look at two frames for a few seconds each and pick the one that produces a reaction. Brand fit is a different property entirely. It asks whether a returning customer still recognises the company, whether the identity the business already owns survives contact with the new layout, and whether someone can put the site in front of a cautious buyer without a preamble.
None of that is visible in a feed. The voter has no relationship with the brand, no memory of the previous site, and no stake in what the company still has to look like in three years.
So the votes are not wrong. They are answers to a narrower question than the one being decided, and the narrower question happens to be the one where drama wins. Strong contrast and unexpected composition both read as quality in a four-second glance. That glance is real, and it matters on a first visit. It is simply not the same measurement as brand fit, and treating the two as interchangeable is where the mistake enters.
What does a design poll actually measure?

A poll measures attitude: how a design is perceived at a glance, not whether it helps a buyer trust the company.
Nielsen Norman Group draws the line the same way, splitting visual-design research into preference testing and behavioural testing and reserving preference work for brand alignment and first impressions. The same article names a failure mode worth internalising: the query effect, where the two options are not different enough for a non-designer to distinguish, so people manufacture a preference rather than report none. Ask a yes-or-no question and you will get a yes or a no whether or not one exists.
There is a second distortion underneath it. The classic 1995 study by Kurosu and Kashimura, covering 252 participants across 26 ATM interface variations and summarised in NN/g's write-up of the aesthetic-usability effect, found that beauty tracked perceived ease of use more closely than actual ease of use. Attractive work hides problems. That is useful in production and actively unhelpful in evaluation.
Was the sample too small to trust?

No: a 90 percent share of twenty-one votes is about nineteen to two, and the 95 percent Wilson interval runs from roughly 71 to 97 percent. Sauro's guidance on testing preference data with the one-sample binomial is the method; run it and the result holds.
Dismissing the poll as underpowered would have been the comfortable move. It also would have been wrong.
Who answered matters more than how many
The defect in a public design poll is construct validity. Wrong population, wrong question, right arithmetic.
A poll posted to a professional network reaches designers and developers who look at interfaces for a living. They are a fine audience for craft. They are not the relocation manager or the investor comparing three agencies before sending one enquiry. When the voting audience and the buying audience differ, a clean statistical result on the wrong population tells you what your peers admire, which is worth knowing and is not the thing you were trying to decide.
Here is how that played out on Elpida Solutions.
For Elpida Solutions, the more dramatic concept won 90 percent of 21 public votes. I still chose the quieter direction because it matched the existing dove-and-house identity and the trust the agency needed to communicate. The poll was useful evidence, but treating it as the decision would have produced the wrong site.
The Elpida Solutions real-estate website case study records what the quieter direction had to carry: company, service, contact and legal pages, a contact flow, and a five-language architecture generating Serbian, English, German, Greek and Russian routes from one shared structure. Five locales is the detail that settles a lot of visual arguments on its own. A layout that depends on a particular line length or a tight headline rhythm has to survive the same headline in five languages, and German or Serbian will not cooperate with an English-tuned grid.
Six criteria that outrank the vote
Once a poll stops being the verdict, it needs something to be evidence for. These are the criteria I weigh, and the honest version of this table admits that the louder concept wins some of them.
| Criterion | The question it answers | Favours the dramatic concept | Favours the quieter concept |
|---|---|---|---|
| Brand continuity | Does a returning customer still recognise the company? | Rarely, unless the brand is being deliberately reset | Usually, because it inherits existing marks and palette |
| Customer trust | Does the intended buyer read this as credible? | In creative and consumer categories | In legal, financial, medical and property categories |
| Content legibility | Can long copy and service detail sit here comfortably? | When the page is short and visual | When the page carries real explanatory text |
| Motion tolerance | Can effects be reduced without the design collapsing? | Seldom, since motion is often the concept | Usually, because motion is decoration rather than structure |
| Production cost | What does it cost to build and to extend? | Higher, especially for bespoke interaction | Lower, and cheaper to hand over |
| Maintainability | Can the client's next contractor keep it consistent? | Only with a documented system | Yes, with conventional components |
Two notes on reading it. Brand continuity and maintainability both improve when the concept is expressed as reusable components with named tokens rather than as one beautiful page; this is the practical argument for a real design system rather than a style guide, and it applies at four pages as much as at forty. Content legibility is the criterion clients underrate most, because a concept is presented with placeholder copy and lives with the real thing — the same trap that shows up when a marketing site has to convert rather than just impress.
Novelty has a ceiling, and it is measurable
Preference rises with novelty and then falls, so the most unusual concept in a set is frequently past the point where unfamiliarity starts costing more than it earns.
Hung and Chen put numbers on this in the International Journal of Design. In their 2012 study of novelty and aesthetic preference, 60 participants rated 88 chair designs and the relationship came out as an inverted U: moderately novel chairs were rated most beautiful, while both the most conventional and the most radical scored lower. It is the empirical shape of Raymond Loewy's older rule that a design should be as advanced as possible while staying acceptable.
That gives the poll a legitimate job. A vote is decent evidence about where a concept sits on the novelty axis. It is poor evidence about where the ceiling is for a specific audience, because the ceiling moves with category. A property agency's buyers sit closer to the typical end than a studio's do.
Motion is no longer only a question of taste
For any service offered into the EU, a concept built on heavy motion now carries a compliance obligation as well as an aesthetic one.
Article 31 of the European Accessibility Act set 28 June 2025 as the date from which member states apply its measures, and the European Commission's scope list includes e-commerce among the covered services. The web part of meeting it is unremarkable engineering. WCAG's Pause, Stop, Hide criterion is Level A and requires a mechanism to pause, stop or hide any moving content that starts automatically, runs longer than five seconds and sits alongside other content; honouring prefers-reduced-motion is how that usually gets built. The design consequence is what interests me here. If a concept's identity is the scroll-driven reveal, then the reduced-motion version is a different and worse design that nobody reviewed, and it is what a real share of visitors will see.
Ask the question during selection instead of after. What does this concept look like with motion disabled? If the answer is "flat and unfinished", the concept has a dependency rather than a feature. Tooling choice interacts with this too, which is part of why Webflow and Framer pull in different directions on motion and control.
What it costs when the popular direction wins anyway
Tropicana is the canonical case. In January 2009 the brand replaced its long-running straw-in-orange packaging with a cleaner, more contemporary design, and sales fell about 20 percent in two months — roughly 30 million dollars before the old packaging was reinstated in February.
The replacement was not ugly. It was better looking by most conventional measures, and it would very likely have won a poll against the packaging it replaced. Let me back up, because "better looking" is carrying too much weight in that sentence. It was cleaner. Cleanliness turned out to be a different property from findability, and findability was the one doing the work. What the new design lost was recognition: shoppers scanning a chiller could no longer find the brand, and some assumed they were looking at a house label.
Websites fail more quietly. There is no shelf, no sales figure moving inside eight weeks, and often no second measurement at all. The equivalent damage shows up as enquiries that do not arrive and a client who cannot say why the new site feels less like their company.
How do you present two directions without hiding the trade-offs?
Show both concepts against the same criteria, state your recommendation with its reasoning, and make the decision a choice between named trade-offs.
- Send the criteria before the concepts. Agree what the decision is about (continuity, trust, legibility, motion, cost, maintenance) while nothing is on screen to react to.
- Present the recommendation first, then the alternative. Presentation order biases preference, so being explicit about your own position is more honest than pretending the running order is neutral.
- Show each concept with real content. Real headlines in every language the site ships, a real service description, the actual legal text. Placeholder copy is how a concept passes review and then fails in build.
- Show the reduced-motion state of each. One screenshot with animation disabled, next to the full version.
- Give each concept a cost and a maintenance note. One sentence each, in plain numbers or plain hours.
- Record the decision and who made it. A short written rationale in the repository or the project doc, dated. It costs ten minutes and settles the question that arrives four months later about why the site looks like this.
- Include the poll as a labelled input. Put the result in the deck with its population named: twenty-one peers, not twenty-one buyers.
That last item is what turns an uncomfortable number into a useful one. You are not hiding the vote. You are telling the client exactly what it measured, which is the only way a client can weigh it properly.
Your turn to run the test: take the concept you are about to recommend, disable motion, replace every placeholder line with the client's real copy in their longest language, and look at what is left. Would it still win the poll? Should it have to?
