Best–Worst projects
Best–Worst is the third kind of project you can run in SortedResearch, alongside card sorts and pairwise projects. Where a card sort shows how people group things and a pairwise project ranks a list two-at-a-time, a Best–Worst project ranks a longer list by what matters most. It shows each participant a small handful of items at a time and asks them to tap the one that matters most and the one that matters least. Those best-and-worst picks add up to a clear, reliable ranking of the whole list.
Best–Worst is sometimes called "best–worst scaling" or "MaxDiff". You don't need the jargon: it is simply the fastest way to sort a long list of options into order of importance.
When to use it
Reach for a Best–Worst project when you have a longer list, say 10 to 30 features, messages, benefits, or ideas, and you need one clean ranking of what matters most to your audience. It shines where a simple "rate each one 1 to 5" survey falls flat: people tend to call everything "important" on a rating scale, but forcing a most-and-least choice makes them trade options off against each other, so the ranking actually separates the list. Because each screen is quick, Best–Worst handles long lists without tiring people out.
Reach for pairwise instead when you also need to know how well each need is served today (the opportunity matrix), and for a card sort when you want to discover how people group a messy list rather than rank it.
Starting a Best–Worst project
Click New project on the dashboard and choose Best–Worst. You give it:
- Title (required).
- Description (optional).
- Client (optional): files the project into a folder on your dashboard, shared with your other projects for the same client.
You add the items to rank on the next screen (the Build tab), so the create step stays quick.
Adding your items
On the Build tab you can add items one at a time with + Add item, or click Bulk add to paste a whole list at once, one per line. Blank lines are ignored, so pasting straight out of a document is fine.
If you paste two columns from a spreadsheet, the first becomes the item and the second becomes its description, which participants see under the item. A project holds up to 200 items; if a paste would go over, we add what fits and tell you how many were left out.
Start from a card sort
Instead of typing your items, you can build a Best–Worst project straight from a card sort you have already run. In the New project window, choose Best–Worst and then From a card sort, pick the card sort, and give the project a title. Its final groups become the items to rank: each group's name becomes an item and Graham's one-line description of that group comes across with it. This turns "how do people group these ideas?" into "which of these groups matters most?" in one step. As with pairwise, the card sort needs a handful of completed responses first so its groups can be read and named.
Card sorts, pairwise projects, and Best–Worst projects live together on the same dashboard, each with a small type badge so you can tell them apart.
What are you ranking?
The items in a Best–Worst project don't have to be "needs". On the Build tab you choose what to call them, needs, cards, statements, features, ideas, or items, and that label is used everywhere in the project. This article says "items" as a stand-in for whatever you have chosen.
What you measure
A Best–Worst project produces one ranking on a single question. You choose what that question is:
- Importance (the default). Each item is ranked on how much it matters to people. The result is a 0 to 100 importance score for every item.
- Custom topic. Rank your items on a question you write yourself, which feature is easiest to understand, which message is most convincing, which name is most trustworthy. On the Build tab you write the words participants read: the most/least prompt shown on every screen ("Which of these is most convincing?" / "…least convincing?") and the labels for the two ends of the 1 to 10 rating used at the end. You can't publish until those are filled in, the same way you need enough items.
Unlike a "Both" pairwise project, Best–Worst does not produce an opportunity matrix, it measures one thing and ranks the list on it. To rank on a different question, start a new project. The custom wording can be tidied on the Build tab right up until the first response arrives; after that it is locked, so everyone in the same study reads exactly the same thing.
How many items per screen, and how many screens?
This is the one setting unique to Best–Worst, and the defaults are set for you, so you rarely need to touch them:
- Items per screen, how many options a participant sees at once. The default is 5 (the well-established sweet spot). For a short list we automatically step this down to 4 so no single screen shows too much of the list.
- Appearances per item, how many times each item is shown to one participant across their screens. The default is 3, which is the reliability floor: fewer, and individual items get too little data.
From these two numbers we work out how many screens each participant sees, and show it to you as you build. For example, a 20-item list at 5 per screen, each seen 3 times, is 12 screens, about 3 minutes of tapping. You set the two inputs; the screen count follows automatically. Every participant gets a freshly balanced set of screens, so every item is shown a fair, equal number of times and gets compared against a wide spread of the others.
The tabs
A Best–Worst project moves through the same six-step funnel as a card sort or a pairwise project: Build → Audience → Screener → Pre-launch → Launch → Results.
Build
Set the project title and description, choose what your items are called, add your list, and (for a custom topic) write the wording participants read. You need at least 5 items to publish. Item order does not matter, every participant gets their own balanced, randomised set of screens, so there are no reorder controls. Your items are locked once data collection has started, so the results stay consistent.
Each item has a short label and an optional description, a single line shown to participants under the item. Descriptions are handy when a bare label is too terse to judge; items ported from a card sort arrive with Graham's group description already filled in.
Under Participant experience you can set your own intro instructions and consent text (both fall back to sensible defaults), and turn on a short "please answer honestly" note. Consent is recorded per respondent. You can set a redirect URL on completion for recruitment panels, and tag a link with a recruitment source with ?src=yourcode, exactly as a pairwise project does.
Audience
Choose how you'll get participants, the same model as the other project types:
- Buy responses from Sorted Research (the default on accounts set up for it). Describe who you want to hear from and we recruit a matched audience. You'll see a live reach and a clean price per response, and you pay before launch (see Sharing & launching for the full details, it works the same for Best–Worst).
- Bring your own respondents. You get a participant link on the Launch step to share with people you recruit yourself.
A Target responses setting lives here too: how many completed responses you're aiming for.
Screener
Optional qualifying questions asked before the ranking, exactly as in a card sort or pairwise project. Add single-choice or multiple-choice questions, mark each answer as qualifies or disqualifies, and set an optional quota. You can also Import with Graham: upload a screener document (.docx or .pdf) and Graham drafts the questions for you to review. Anyone screened out sees a polite thank-you and does not go on.
Pre-launch
A review step before anything goes live: it shows your reach, the price, and a go/no-go readiness checklist. Nothing is spent here, it's the final check before you launch.
Launch
A Best–Worst project has a single card showing its status, a progress bar toward your target, and (for bring-your-own) its participant link. You can go live, pause, or close it. When you're buying responses, you place the order and pay here before Go live unlocks. A Preview as participant link runs the whole experience for you without recording anything, so you can check it before you share it.
Results
Once responses arrive you get:
- A filter bar (the same as the other types): Include checkboxes for each quality level so you can drop flagged responses, a Segment dropdown, and a running "Analyzing X of Y responses" count.
- The Recruitment funnel: total responses, completion rate, how many were screened out, and the median time to respond.
- A ranking of every item from most to least important, each with its 0 to 100 score and a ± confidence range (a bootstrap estimate: smaller means steadier).
- The must-have line: a marker on the ranking that separates the items that clear the bar from the ones that don't, so you can see at a glance where to focus and where it's safe to stop.
- A CSV download. Where an item has a description, it is shown under the item's name and included as its own column.
The Results tab has four views: Overview (everything above), Analysis, Participants, and Report.
The Analysis view goes past the ranking itself: Polarization shows how often each item was picked best against how often it was picked worst, so an item that splits the room stands out instead of averaging into the middle; Shortlist reach (TURF) answers "if we can only carry a few, which few reach the most people?", building the shortlist step by step and comparing it against simply taking the top few by score; and By segment shows how importance shifts across the audiences you screened for, for any segment with at least five participants.
The Report view lets Graham write a client-ready insight report from your results, the headline story, the must-haves, the also-rans, and recommended next steps, downloadable as Word, a branded PowerPoint, Print / save as PDF, or CSV.
The Participants view lists everybody who finished, newest first, with their quality flag and the reasons behind it, how many screens they answered, and how long they took. Anyone set aside for quality is marked, and is already left out of the scores. You can permanently remove a single participant from here, which deletes their choices, their end ratings and their screener answers with them, and the scores recompute without them the next time you open the tab. That covers a GDPR erasure request for one person.
For a GDPR access request, download that participant's full record as a CSV from the same view: their screener answers, every screen they saw with what they picked most and least important, their response times and their end ratings. An admin can also export the whole workspace as one JSON file from the account area.
Participants can leave an optional comment on the thank-you screen, which appears in your Participant feedback area alongside your other feedback.
How the 0 to 100 scores are worked out
Each item's score starts from all the best-and-worst picks: a model works out how strong each item is from how often it was chosen best, how often worst, and which items it was shown against, more reliable than a plain best-minus-worst count. That gives a ranking, which is then placed on a 0 to 100 scale using the short rating question at the end, where each person rates the item they picked most often and the one they picked least often from 1 to 10. The average "most" rating sets roughly where the top of the scale sits and the average "least" rating the bottom, so the numbers mean something in plain terms rather than only relative to each other. The must-have line sits on this same absolute scale. This is the same calibration a pairwise project uses, so scores read on a common scale across study types.
Only people who rated their most-picked item above their least-picked one are used to set that scale, so a careless or contradictory rating never squashes the spread. This only affects how far apart the scores sit, never which item ranks above which.
Response quality and anti-cheating
Best-Worst projects protect data quality in three layers, and only the middle one is ours to see on screen.
The panel, before they arrive. On a bought audience, PureSpectrum turns away fraud, duplicates and anyone who fails your screener before they ever reach your study. Those people never reach your results and never reach your bill.
Built-in integrity (always on). Every recorded pick is validated on the server (the chosen best and worst must both be items actually shown on that screen, and the two must differ), every participant endpoint is rate limited, sessions are token-authorised, and finished sessions cannot be resubmitted. A device that has already finished a study is turned away from it for up to 180 days. This stops bots and tampering.
Quality checks on completed responses. Every completed response is judged by five checks, and any one of them rejects it:
- Nothing answered: reached the end without answering a single screen.
- Speeder check that understands familiarity: a screen full of new items has to be read, but once items are familiar, fast honest answers are fine. Only genuinely impossible speeds are rejected.
- Impossible timings: claimed more time than the visit could actually have contained.
- Best position bias: put "best" in the same place on nearly every screen. Item order is randomised, so honest picks are spread out.
- Worst position bias: the same thing for "worst".
A rejected response is left out of the scores automatically, and if it came from a bought audience it is sent back to the panel as a quality reject, so it is replaced and you are not charged for it. This is the one place Best-Worst differs from a card sort or a pairwise study: those two give you a checkbox to put flagged responses back in, and Best-Worst does not, so the decision is made for you rather than by you.
Signals that flag without rejecting. Three more signals move somebody up the Respondents review list with a reason against them, without touching their data or their payment: contradicting themselves more than most people in the study did, rating their most-chosen item no higher than their least-chosen one, and going far faster than the study's own median. These are there for you to look at, not for us to act on.
How many responses do I need?
More responses give steadier rankings. A practical starting point is around 30 completed responses for a rough read, and 100 or more when you need to trust the ordering for a decision. A longer item list benefits from being on the higher end, since each item is seen by a smaller share of any one person's screens.
What participants see
Participants open the link on any device and, after the optional screener, see a small set of items, usually five, and tap the one that matters most, then the one that matters least. That repeats for a handful of quick screens. At the end they answer two short 1 to 10 rating questions. It is quick, works on a phone, and needs no account.