← All guides

Why rating scales fail on long lists

You put twenty features in front of 300 people, ask each to be rated 1 to 5 for importance, and get back a list where seventeen average between 4.1 and 4.6. Everything is important. Nothing is prioritised. The study cost real money and changed no decision.

This happens reliably, and it is worth understanding why, because the failure is in the instrument rather than in the people answering.

Nobody has to choose

A rating scale lets a respondent say every single item is important, and there is no cost to doing so. Ticking "very important" twenty times takes less thought than deciding which of two things matters more.

And in fairness to them, it is often true. Reliable sizing is important. Free returns are important. Asking somebody whether they want a good thing produces an unsurprising answer. The question you actually needed was which good thing they would give up first, and a rating scale never asks it.

People use scales differently

Two respondents can hold identical views and produce different numbers.

Some people never use the extremes: their scale runs 2 to 4 and a 4 means "critical". Others use only 1 and 5. This varies systematically by culture and personality, and it is not noise you can average away, because it biases whole subgroups in the same direction. If your enthusiastic segment is also your younger segment, you will conclude younger customers care more about everything.

The ceiling crushes the differences

With five points and twenty items, most items cluster in the top two. Real differences get compressed into the gap between 4.2 and 4.4, and then you are reading rounding as signal.

Going to a 1 to 10 scale helps less than people hope. It moves the pile-up from 4s to 8s and 9s.

What forcing a choice fixes

MaxDiff, also called Best-Worst scaling, shows a few items at a time and asks for the most and the least important. Both problems disappear at once:

  • Nobody can say everything matters. Every screen extracts a trade-off. That is the entire point.
  • Scale use stops mattering. There is no scale. A choice between four items means the same thing regardless of whether somebody is naturally generous or naturally stingy, which makes respondents genuinely comparable.

You also get a bonus: asking for the least important doubles the information from each screen, and negative choices discriminate especially well at the bottom of a list, which is exactly where a rating scale is most useless.

What it costs

Honesty about the downsides:

  • You lose absolute levels. MaxDiff tells you A matters more than B, and by how much relative to the rest. It does not tell you whether the whole list matters in absolute terms. If you genuinely need "is this important, yes or no", ask that separately.
  • It takes a few screens. A rating grid can be answered in one screen, badly. MaxDiff takes a handful, properly.
  • You need to build the list first. Garbage items produce a beautifully precise ranking of garbage.

When a rating scale is right

It is not a bad instrument, it is a misapplied one. Ratings work well for:

  • Satisfaction with something specific they have experienced. "How well does your current retailer handle returns?" has a real absolute answer.
  • A handful of items. With four or five, ceiling effects have less room to do damage.
  • Tracking over time, where you care about movement rather than rank order.

That first case is why SortedResearch's pairwise studies ask two rating questions at the end. Importance comes from the forced choices, where ratings fail. Satisfaction is a rating, where ratings work. Putting them together is what produces an opportunity matrix.

The practical test

Look at your last prioritisation study. If more than half the items scored above the midpoint, the instrument did not discriminate, and the ranking you took to the meeting was mostly measuring how generous your respondents were feeling.


Related: MaxDiff vs conjoint: what each one is actually for · Finding unmet needs with an opportunity matrix

Run one yourself

The free plan runs a full study with up to 10 participants, and needs no card.