Pairwise projects
Pairwise is the second kind of project you can run in SortedResearch, alongside card sorts. Where a card sort shows how people group things, a pairwise project ranks a list of needs on what matters to you, most often how much each one matters to people and how well those needs are met today, or on a question you write yourself. It does this by asking participants lots of simple two-at-a-time questions ("which of these matters more to you?") and turning the answers into a ranking.
When to use it
Reach for a pairwise project when you have a list of customer needs, features, or jobs-to-be-done and you want to know which ones to prioritize. It is especially good for spotting opportunities: needs people care about a lot but are not well served on today.
Starting a pairwise project
Click New project on the dashboard and choose Pairwise. You give it:
- Title (required).
- Description (optional).
- Client (optional): files the project into a folder on your dashboard, shared with your card sorts for the same client.
You add the items to compare on the next screen (the Build tab), so the create step stays quick.
Adding your items
On the Build tab you can add items one at a time with + Add need, or click Bulk add to paste a whole list at once, one per line. Blank lines are ignored, so pasting straight out of a document is fine.
If you paste two columns from a spreadsheet, the first becomes the item and the second becomes its description, which participants see under the item. Bulk add drops your items into the list ready to review, and they are saved when you click Save needs. If you paste and click Save needs straight away, that works too: your paste is kept either way. A project holds up to 200 items; if a paste would go over, we add what fits and tell you how many were left out.
Start from a card sort
Instead of typing your items, you can build a pairwise project straight from a card sort you have already run. In the New project window, choose Pairwise and then From a card sort, pick the card sort, and give the project a title. Its final groups become the items to compare: each group's name becomes an item and Graham's one-line description of that group comes across with it (shown to participants under the item). This turns "how do people group these ideas?" into "which of these groups matters most, and how well are they served?" in one step.
Before importing, we make sure the card sort's groups have been named and described. If Graham has not done that yet, he names and describes them now, so the items are clear and each name is unique. The card sort needs a handful of completed responses first, so its groups can be read; very new card sorts will not appear in the list yet. Once imported, the items land on the Build tab where you can review, edit, or remove any of them before you publish.
Card sorts and pairwise projects live together on the same dashboard, each with a small type badge so you can tell them apart.
What are you comparing?
The items in a pairwise project are not always "needs". On the Build tab you choose what to call them, needs, cards, statements, features, ideas, or items, and that label is used everywhere in the project. This article says "needs" as a stand-in for whatever you have chosen.
What you measure
When you create a pairwise project you choose what it measures. You pick one of four options on the way in:
- Both (the default). The project measures importance (how much each need matters) and satisfaction (how well people are served on each need today). These are the two axes of the opportunity matrix, which is what makes pairwise so good at spotting opportunities. Each is answered by its own separate group of people, so if you want 100 responses you are really recruiting 200: 100 who answer the importance version and 100 the satisfaction version, each with its own participant link.
- Importance only or Satisfaction only. Measure just one of the two. You get a single ranking on that one thing, one participant link, and no opportunity matrix (the matrix needs both axes to plot).
- Custom topic. Rank your needs on a question you write yourself, for when importance and satisfaction are not what you are after. You might ask which packaging feels most premium, which feature is easiest to understand, or which name is most trustworthy. You pick "Custom topic" when you create the project and then, on the Build tab, write the words participants read: the question shown on every pair ("Which of these feels more premium?"), the two short rating questions asked at the end about their most- and least-chosen item, and the labels for the two ends of the 1 to 10 scale ("Not at all premium" through "Extremely premium"). You can't publish the project until those are filled in, the same way you need at least five items. The result is a single 0 to 100 ranking on your topic, with no opportunity matrix. Everything else, the two-at-a-time comparisons, the anti-cheating, and the scoring, works exactly as it does for importance or satisfaction.
You can tidy the custom wording on the Build tab right up until the first response arrives; after that it is locked, so everyone in the same study reads exactly the same thing. And what a project measures is fixed once responses start coming in: to measure something different, start a new project.
The four tabs
A pairwise project opens in the same four-tab layout as a card sort.
Build
Set the project title and description, choose what your items are called, add your list of items, and set two recruitment options:
Target responses: how many completed responses you are aiming for. When you are running both importance and satisfaction the number applies to each of them; for a single-dimension or custom project it is simply the target for your one link.
Comparisons per participant: how many two-at-a-time choices each person makes. Locked once responses start coming in.
Pick one number and use the same one for both importance and satisfaction. The two scores are meant to be read together on the opportunity matrix, so both sides should be collected the same way; if one dimension asks each person for many more comparisons than the other, that side is both better covered and more tiring to complete, and the two axes stop being a fair comparison. More comparisons per person means each pairing is seen more often, so the ranking is steadier, but a longer task also tires people and slightly raises the rate of rushed or careless answers. As a rough guide, aim for enough that every pairing is seen a few times across all your respondents: for a short list (around 10 items) a dozen or so comparisons each is plenty, and for a longer list (25 to 30 items) around 16 to 20 keeps coverage healthy without overloading anyone. Very low counts on a long list leave the ranking thin. Note that the number is locked once responses arrive, so set it before you launch and keep both dimensions in step.
Each item has a short label and an optional description, a single line shown to participants under the item on the compare screen. Descriptions are handy when a bare label is too terse to judge; items ported from a card sort arrive with Graham's group description already filled in. Leave it blank for items that speak for themselves.
You need at least 5 items to publish. Item order does not matter (every item is compared against every other, in a random order per participant), so there are no reorder controls. Your items are locked once data collection has started, so the results stay consistent.
Under Participant experience you can set your own intro instructions and consent text (both fall back to sensible defaults), and turn on a short "please answer honestly" note. Consent is recorded per respondent. You can set a redirect URL on completion for recruitment panels: paste their link including the word TRACKINGID and we swap in each respondent's tracking code, then send them back after they finish. You can also tag a link with a recruitment source by adding ?src=yourcode (or standard panel params like ?iid=) to it, which shows up against each response and is used as the tracking code; and a respondent who already finished a wave on a device sees a thank-you instead of taking it twice.
Screener
Optional qualifying questions asked before the comparisons, the same idea as the card sort screener. Add single-choice or multiple-choice questions, mark each answer as qualifies or disqualifies, and set an optional quota to stop accepting people who picked a given answer once that group is full. You can also Import with Graham: upload a screener document (.docx or .pdf) and Graham drafts the questions for you to review, keeping the single and multiple choice ones. Anyone who is screened out sees a polite thank-you and does not go on to the comparisons. With no screener questions, participants go straight in.
Share & publish
Each dimension you are running has its own card showing its status, a progress bar toward your target, and its own participant link. You can publish a dimension to take it live, pause it to stop accepting responses (you can resume later), or close it for good. If you chose both, publish both to run the full study; an importance-only, satisfaction-only, or custom project has a single card and link. Each card also has a Preview as participant link that runs the whole experience for you without recording anything, so you can check it before you share it.
Results
Once responses arrive you get:
- A filter bar (the same as card sort): Include checkboxes for each quality level (ok / suspect / rejected, with counts) so you can drop flagged responses from the scores and matrix, a Segment dropdown, and a running "Analyzing X of Y responses" count.
- The Recruitment funnel: total responses (split by dimension), completion rate, how many people were screened out, and the median time to respond.
- Per-need scores for importance and satisfaction.
- An opportunity matrix: every need plotted by importance (up) against satisfaction (across), split into four quadrants: Opportunity (top-left: important, not well served), Table stakes (top-right: important and already well served), Over-invested (bottom-right: not important but well served), and Low priority (bottom-left). Each dot is colored by its opportunity score, red for low, green for high, and hovering a dot highlights it in the list beside the chart.
- A needs ranking sorted by opportunity, and a CSV download. Where an item has a description, it is shown under the item's name in the ranking and included as its own column in the CSV, so a study built from a card sort still explains what each need means.
Each score carries a ± confidence range (a bootstrap estimate: smaller means steadier), shown next to the number and as whiskers on the matrix. If some needs are thinly compared, a coverage note suggests collecting more responses. The Segment dropdown in the filter bar recomputes all the scores and the matrix for just the respondents who gave a particular screener answer (e.g. new vs returning customers).
If your project measures a single dimension (importance only, satisfaction only, or a custom topic), the opportunity matrix does not apply: the Results show one 0 to 100 ranking on that topic, the ranking table and CSV carry a single score column labelled with what you measured, and Graham's report is written around that one ranking, its leaders and laggards and what to do about them, rather than the four quadrants.
The Results tab has three sub-tabs: Summary (everything above), Responses, and Report. The Report tab lets Graham write a client-ready insight report from your results, the headline story, the top opportunities, table stakes, over-invested areas, recommended next steps, and an honest note on data quality. Item descriptions come across into the report too: they sit under each need in the matrix key and Graham reads them when writing, and they carry into the Word and PowerPoint downloads. You can download it as a Word document or a branded PowerPoint deck (with the opportunity matrix as a slide), Print / save as PDF, or export the scores as CSV (which opens in Excel), the same report options as a card sort. The Responses view lists every respondent with their dimension, quality flag and reasons, number of choices and time taken. Expand a row to see each pairwise choice they made and their end ratings; download one respondent's full raw data (for a GDPR access request) or delete a response to drop it from the results.
Participants can leave an optional comment on the thank-you screen, which appears in your Participant feedback area alongside your card-sort feedback.
How the 0 to 100 scores are worked out
Each need's score starts from all the head-to-head comparisons: a model works out how strong each need is from how often it won and who it beat, which is more reliable than a plain win count. That gives a ranking, which is then placed on a 0 to 100 scale using the short rating question at the end, where each person rates their most and least favourite need from 1 to 10. The average "most" rating sets roughly where the top of the scale sits and the average "least" rating the bottom, so the numbers mean something in plain terms rather than only relative to each other.
Only people who rated their favourite need above their least favourite are used to set that scale. Some respondents answer the rating question carelessly and score their least favourite the same as, or higher than, their favourite, and including those contradictory answers used to squash all the scores into a narrow band (this showed up most on satisfaction, where people tend to rate everything fairly highly). Ignoring them lets the scores spread out and tell the needs apart properly. This only affects how far apart the scores sit, never which need ranks above which.
Response quality and anti-cheating
Pairwise projects protect data quality in two layers.
Built-in integrity (always on). Every recorded choice is validated on the server (the winner must be one of the two shown, and the pair must match the one actually assigned), every participant endpoint is rate limited, sessions are token-authorised, and finished sessions cannot be resubmitted. This stops bots and tampering.
Quality flags on completed responses. Each completed response is scored and marked suspect or rejected (or left clean) using:
- Speeder check that understands familiarity: the first time a card appears it has to be read, but once both cards are familiar, fast honest answers are fine. Only genuinely impossible speeds are flagged.
- Side-bias: someone who keeps tapping the same side of the screen (left/right is randomised, so honest picks are about 50/50).
- Consistency: a hidden repeated pair (a different answer the second time is a contradiction), plus a self-contradiction rate across their choices (A beats B, B beats C, but C beats A).
- Anchor sanity: rating their most-chosen card lower than their least-chosen one.
- Repeat participant: a browser that already finished this wave is flagged (a soft guard, not a hard block, so people sharing a network are never wrongly rejected).
On the Results tab, the Response quality panel shows how many responses were flagged and why, with a checkbox to exclude suspect and/or rejected responses from the scores and the opportunity matrix. Nothing is dropped unless you choose to exclude it, so you stay in control.
How many responses do I need?
More responses give steadier rankings. A practical starting point is around 30 completed responses per dimension for a rough read, and 100 or more per dimension when you need to trust the ordering and the opportunity matrix for a decision.
What participants see
Participants open the link on any device and, after the optional screener, see two needs at a time and pick one. After a set number of pairs they answer two short 1 to 10 rating questions. It is quick and needs no account.