← All help articles

Pairwise studies

Pairwise is the second kind of study you can run in SortedResearch, alongside card sorts. Where a card sort shows how people group things, a pairwise study ranks a list of needs on what matters to you, most often how much each one matters to people and how well those needs are met today, or on a question you write yourself. It does this by asking participants lots of simple two-at-a-time questions ("which of these matters more to you?") and turning the answers into a ranking.

When to use it

Reach for a pairwise study when you have a list of customer needs, features, or jobs-to-be-done and you want to know which ones to prioritize. It is especially good for spotting opportunities: needs people care about a lot but are not well served on today.

Starting a pairwise study

Click New study on the dashboard and choose Pairwise. You give it:

  • Title (required).
  • Description (optional).
  • Client (optional): files the study into a folder on your dashboard, shared with your card sorts for the same client.

You add the items to compare on the next screen (the Build tab), so the create step stays quick.

Adding your items

On the Build tab you can add items one at a time with + Add item, or click Bulk add to paste a whole list at once, one per line. Blank lines are ignored, so pasting straight out of a document is fine.

If you paste two columns from a spreadsheet, the first becomes the item and the second becomes its description, which participants see under the item. Bulk add drops your items into the list with Add to list, ready to review. There is no save button, the list saves automatically as you type, add, and remove, and a small "All changes saved" note confirms it. A study holds up to 200 items; if a paste would go over, we add what fits and tell you how many were left out.

Adding a picture to an item

Any item can carry a picture, shown above its text on the comparison card. On the Build tab, click Add picture on the item's row, pick a file, and it uploads and appears as a thumbnail straight away. Replace swaps it, Remove takes it away, and the small text box beside the thumbnail is the description read aloud by screen readers, which starts off as the item's own words and can be edited.

Pictures are available on the paid plans. On the Free plan the slot explains that and does nothing.

What you can upload. JPEG, PNG, WebP, GIF, and photos straight off an iPhone (HEIC), which we convert to JPEG for you because most browsers cannot display HEIC. Large photographs are automatically resized before they upload, so a phone photo does not become a slow download for your participants.

Animations. An animated GIF stays animated: we store it exactly as you uploaded it and never re-encode it, so it plays for participants. This is the way to compare things that move, a mechanism opening, a screen flow, an interaction. Animations can be up to about 4 MB; above roughly 1.5 MB we warn you that it is a lot for somebody answering on mobile data, but we do not stop you.

A study can be part illustrated. An item with no picture simply shows its text as usual, so you can add a picture only to the items that need one.

Pictures follow the item everywhere a participant meets it: on the comparison card, and again on the two rating questions at the end, where they are asked about their most- and least-chosen item. On an illustrated study the picture is what they were judging, so naming it in words alone at the rating step would ask them to recognise from a label something they only ever saw as an image.

Pictures are locked once responses arrive, exactly like the wording. What a respondent was shown cannot change underneath the data they gave you.

Start from a card sort

Instead of typing your items, you can build a pairwise study straight from a card sort you have already run. In the New study window, choose Pairwise and then From a card sort, pick the card sort, and give the study a title. Its final groups become the items to compare: each group's name becomes an item and Graham's one-line description of that group comes across with it (shown to participants under the item). This turns "how do people group these ideas?" into "which of these groups matters most, and how well are they served?" in one step.

Before importing, we make sure the card sort's groups have been named and described. If Graham has not done that yet, he names and describes them now, so the items are clear and each name is unique. The card sort needs a handful of completed responses first, so its groups can be read; very new card sorts will not appear in the list yet. Once imported, the items land on the Build tab where you can review, edit, or remove any of them before you publish.

Card sorts, pairwise studies, and Best–Worst studies live together on the same dashboard, each with a small type badge so you can tell them apart.

What are you comparing?

The things a pairwise study compares are called items throughout the app. They can be anything, customer needs, product concepts, features, statements, names. This article says "needs" in places as a stand-in for whatever your items are. (Older studies that picked a different label, like "needs" or "ideas", keep showing it.)

What you measure

When you create a pairwise study you choose what it measures. You pick one of two options on the way in:

  • Importance & satisfaction (the default). The study measures importance (how much each need matters) and satisfaction (how well people are served on each need today). These are the two axes of the opportunity matrix, which is what makes pairwise so good at spotting opportunities. For an importance & satisfaction study you choose, on the Build tab, who answers the two dimensions:
    • Separate groups (the default): different people answer importance and satisfaction, cleaner, independent samples. If you want 100 responses you're recruiting 200 (100 per dimension).
    • Same people, both (paired): each person answers both dimensions in one sitting, so 100 people give you 100 responses on each. This unlocks a per-person opportunity view in the results, you can see who personally finds something important and is unsatisfied with it, not just the averages.
  • Custom topic. Rank your needs on a question you write yourself, for when importance and satisfaction are not what you are after. You might ask which packaging feels most premium, which feature is easiest to understand, or which name is most trustworthy. You pick "Custom topic" when you create the study and then, on the Build tab, write the words participants read: the question shown on every pair ("Which of these feels more premium?"), the two short rating questions asked at the end about their most- and least-chosen item, and the labels for the two ends of the 1 to 10 scale ("Not at all premium" through "Extremely premium"). You can't publish the study until those are filled in, the same way you need at least five items. The result is a single 0 to 100 ranking on your topic, with no opportunity matrix. Everything else, the two-at-a-time comparisons, the anti-cheating, and the scoring, works exactly as it does for importance or satisfaction.

A description is required before a study can be published. Participants never see it: it is where you say what you are trying to find out, and it is the brief Graham writes the participant wording from. A study that was already live before this was introduced keeps working and can still be paused and resumed without one.

Graham can draft all six fields for you. On the Build tab, Draft all six with Graham writes the pair question, the two rating questions and the two scale labels for you, and you edit anything you like before saving. Graham writes from your description of the study, so a description is needed before the button will do anything: it is the only place you say what you are actually trying to find out. Your title and topic label come along as supporting context, but they are filing names, and a study called "Image Quality" could be about a hundred different questions. Participants never see the description.

You can tidy the custom wording on the Build tab right up until the first response arrives; after that it is locked, so everyone in the same study reads exactly the same thing. And what a study measures is fixed once responses start coming in: to measure something different, start a new study.

The tabs

A pairwise study moves through the same six-step funnel as a card sort: Build → Audience → Screener → Pre-launch → Launch → Results.

Build

Set the study title and description, add your list of items, and (for an importance & satisfaction study) choose whether the two dimensions are answered by separate groups or the same people (see What you measure).

Each item has a short label and an optional description, a single line shown to participants under the item on the compare screen. Descriptions are handy when a bare label is too terse to judge; items ported from a card sort arrive with Graham's group description already filled in. Leave it blank for items that speak for themselves.

You need at least 5 items to publish. Item order does not matter (every item is compared against every other, in a random order per participant), so there are no reorder controls. Your items are locked once data collection has started, so the results stay consistent.

Under Participant experience you can set your own intro instructions and consent text (both fall back to sensible defaults), and turn on a short "please answer honestly" note. Consent is recorded per respondent. You can set a redirect URL on completion for recruitment panels: paste their link including the word TRACKINGID and we swap in each respondent's tracking code, then send them back after they finish. You can also tag a link with a recruitment source by adding ?src=yourcode (or standard panel params like ?iid=) to it, which shows up against each response and is used as the tracking code; and a respondent who already finished a wave on a device sees a thank-you instead of taking it twice.

Audience

Choose how you'll get participants, the same model as a card sort:

  • Buy responses from Sorted Research (the default). Describe who you want to hear from and we find and recruit a matched audience for you. You'll see a live reach (how many people match) and a clean price per response, and you pay before launch (see Sharing & launching for the full buying, pricing, pay-before-launch and auto-refund details, it works the same for pairwise). A very large or specialised audience may show "request a quote" instead of an instant price.
  • Bring your own respondents. You get a participant link on the Launch step to share with people you recruit yourself.

Two recruitment settings live here too:

  • Target responses: how many completed responses you're aiming for. For an importance & satisfaction study run as separate groups it applies to each dimension; for a paired or custom study it's simply your one target.

  • Comparisons per participant: how many two-at-a-time choices each person makes. We work this out for you from the number of items in your study, and you can type over it whenever you disagree.

    The rule: enough comparisons for each person to see each item about three times, held to a minimum of 10 and a maximum of 40. Three sightings each is what makes one person's own ranking mean something, which is what the per-person view rests on. The floor of 10 is where the quality checks start working, the side-bias check needs 10 choices before it will run at all. The ceiling of 40 is fatigue: past roughly forty choices people stop deciding and start tapping, and that is worse than having less data, because it looks real. On a long list the ceiling wins, so each person sees each item fewer times and the coverage comes from having many respondents instead. The suggestion also tells you roughly how long the study takes each person, worked out with the same estimator that prices a panel order.

    The suggestion follows your item list as you add and remove, right up until you set a number yourself. After that the number is yours and nothing moves it. A study that already existed before this arrived keeps whatever number it had, and offers you the suggestion beside it.

    Use the same number for both importance and satisfaction. The two scores are meant to be read together on the opportunity matrix, so both sides should be collected the same way; if one dimension asks each person for many more comparisons than the other, that side is both better covered and more tiring to complete, and the two axes stop being a fair comparison. Note that the number is locked once responses arrive, so set it before you launch and keep both dimensions in step.

Screener

Optional qualifying questions asked before the comparisons, the same screener as every other study type (see The screener): single choice, multiple choice, short text, open-ended, number with a range, and matrix grids; "show if" branching on an earlier answer; qualifying, disqualifying and must-select answers; and quotas by count or by share of your target, which stop accepting people who picked a given answer once that group is full. You can also Import with Graham: upload a screener document (.docx or .pdf) and Graham drafts the questions, every kind included, for you to review before anything is added. Anyone who is screened out sees a polite thank-you and does not go on to the comparisons. With no screener questions, participants go straight in. If you're buying responses, a stricter screener means fewer people qualify, which can raise the price per response, the live reach and price update as you edit it, and you never enter any technical rates.

Pre-launch

A review step before anything goes live: it shows your reach, the price, and a go/no-go readiness checklist. Nothing is spent here, it's just the final check before you launch.

Launch

Each dimension you're running has its own card showing its status, a progress bar toward your target, and (for bring-your-own) its participant link. Take a part live with Use your audience, pause it to stop accepting responses (you can resume later), or close it, closing is final for a bought order; a bring-your-own part can be reopened. If you're running importance & satisfaction as separate groups, launch both to run the full study; a paired or custom study has a single card. When you're buying responses, you place the order and pay here before Go live unlocks. Each card also has a Preview as participant link that runs the whole experience for you without recording anything, so you can check it before you share it.

Results

Once responses arrive you get:

  • A filter bar (the same as card sort, shown above every view): Include checkboxes for each quality level (ok and rejected, with counts) so you can drop flagged responses from the scores and matrix, a Segment dropdown, and a running "Analyzing X of Y responses" count.
  • The Recruitment funnel: total responses (split by dimension), completion rate, how many people were screened out, and the median time to respond.
  • Per-need scores for importance and satisfaction.
  • An opportunity matrix: every need plotted by importance (up) against satisfaction (across), split into four quadrants: Opportunity (top-left: important, not well served), Table stakes (top-right: important and already well served), Over-invested (bottom-right: not important but well served), and Low priority (bottom-left). Each dot is colored by its opportunity score, red for low, green for high, and hovering a dot highlights it in the list beside the chart.
  • A needs ranking sorted by opportunity, and a CSV download. Where an item has a description, it is shown under the item's name in the ranking and included as its own column in the CSV, so a study built from a card sort still explains what each need means.

If you ran an importance & satisfaction study as Same people, both (paired), you also get a per-person opportunity view: because each respondent answered both dimensions, we can see who personally rates a need important while being unsatisfied with it, surfacing pockets of opportunity that the averages alone would hide.

Each score carries a ± confidence range (a bootstrap estimate: smaller means steadier), shown next to the number and as whiskers on the matrix. If some needs are thinly compared, a coverage note suggests collecting more responses. The Segment dropdown in the filter bar recomputes all the scores and the matrix for just the respondents who gave a particular screener answer (e.g. new vs returning customers).

If your study measures a single dimension (such as a custom topic), the opportunity matrix does not apply: the Results show one 0 to 100 ranking on that topic, the ranking table and CSV carry a single score column labelled with what you measured, and Graham's report is written around that one ranking, its leaders and laggards and what to do about them, rather than the four quadrants.

The Results tab has these views: Overview (the funnel and the ranking table), Analysis (the opportunity matrix, the items ranked by opportunity, and importance against satisfaction item by item, as a second row of buttons; a custom-topic study shows its one ranking), Per person where the study measures more than one thing, Participants, and Report. A Download row just under the tabs offers Excel (the results table, the per-person table where there is one, and the raw data as sheets) and Raw data (CSV) (one row per pair a person chose, plus one row per end rating, with their screener answers on every row). The Report tab lets Graham write a client-ready insight report from your results, the headline story, the top opportunities, table stakes, over-invested areas, recommended next steps, and an honest note on data quality. Item descriptions come across into the report too: they sit under each need in the matrix key and Graham reads them when writing, and they carry into the Word and PowerPoint downloads. You can download it as a PDF, a Word document or a branded PowerPoint deck (with the opportunity matrix as a slide), or print it, the same report options as a card sort. A custom-topic study gets the same PDF, Word and PowerPoint, built around its one ranking. The Participants view lists every participant with their dimension, quality flag and reasons, number of choices and time taken. Expand a row to see each pairwise choice they made and their end ratings; download one participant's full raw data, screener answers included (for a GDPR access request) or delete a response to drop it from the results.

Participants can leave an optional comment on the thank-you screen, which appears in your Participant feedback area alongside your card-sort feedback.

How the 0 to 100 scores are worked out

Each need's score starts from all the head-to-head comparisons: a model works out how strong each need is from how often it won and who it beat, which is more reliable than a plain win count. That gives a ranking, which is then placed on a 0 to 100 scale using the short rating question at the end, where each person rates their most and least favourite need from 1 to 10. The average "most" rating sets roughly where the top of the scale sits and the average "least" rating the bottom, so the numbers mean something in plain terms rather than only relative to each other.

Only people who rated their favourite need above their least favourite are used to set that scale. Some respondents answer the rating question carelessly and score their least favourite the same as, or higher than, their favourite, and including those contradictory answers used to squash all the scores into a narrow band (this showed up most on satisfaction, where people tend to rate everything fairly highly). Ignoring them lets the scores spread out and tell the needs apart properly. This only affects how far apart the scores sit, never which need ranks above which.

Response quality and anti-cheating

Pairwise studies protect data quality in three layers, and only the last one is yours to control.

The panel, before they arrive. On a bought audience, PureSpectrum turns away fraud, duplicates and anyone who fails your screener before they ever reach your study. Those people never reach your results and never reach your bill.

Built-in integrity (always on). Every recorded choice is validated on the server (the winner must be one of the two shown, and the pair must match the one actually assigned), every participant endpoint is rate limited, sessions are token-authorised, and finished sessions cannot be resubmitted. A device that has already finished a study is turned away from it for up to 180 days. This stops bots and tampering.

Quality checks on completed responses. Every completed response is judged by five checks, and any one of them rejects it:

  • Too many choices too fast: a large share of choices made faster than a person could read them. The check understands familiarity, so the first time a card appears it has to be read, but once both cards are familiar, fast honest answers are fine. It also understands pictures: a card with a picture is given time to be looked at, and a card with an animation is given time for the animation to play, so an illustrated study is not judged as though its cards were a line of text.
  • Typical response time below human reaction time: their usual pace is under what a person physically does.
  • Side bias: keeps tapping the same side of the screen. Left and right are randomised, so honest picks are about 50/50.
  • Self-contradiction: too many of their comparisons contradict each other (A beats B, B beats C, but C beats A).
  • Repeat participant: this browser had already completed this wave.

Signals that flag without rejecting. Three more signals move somebody up the Respondents review list with a reason against them, without touching their data or their payment: a couple of contradictions on the hidden repeated pair, a moderate share of choices too fast to have been read, and leaning heavily but not exclusively to one side. These are there for you to look at, not for us to act on.

On the Results tab, the Response quality panel shows how many responses were flagged and why, with an Include filter to keep rejected responses in or take them out of the scores and the opportunity matrix. That filter is for exploring your own data. It is separate from a panel reject: a respondent handed back to PureSpectrum has left the study altogether, and no filter brings them back, because they have been replaced and you were not charged for them.

Anchor sanity is handled separately, in the calibration above rather than here. Somebody who rates their favourite need no higher than their least favourite is left out of setting the 0 to 100 scale, but their comparisons still count towards the ranking.

How many responses do I need?

More responses give steadier rankings. A practical starting point is around 30 completed responses per dimension for a rough read, and 100 or more per dimension when you need to trust the ordering and the opportunity matrix for a decision.

Those numbers come from simulating this product's own scoring, not from a rule of thumb borrowed from another method. At 30 per dimension the ranking correlates about 0.90 with the true order and roughly four of the true top five needs come out in your top five, which is enough to tell you where to look. But about one need in four still lands in the wrong quadrant of the opportunity matrix at that size, and the quadrant is the part you make the decision from, so treat 30 as directional and not as a green light. By 100 per dimension the correlation is around 0.97 and quadrant placement is right about 86% of the time.

Pairwise needs roughly twice the sample a card sort does. It is the same participants doing a lighter task: a card sorter places every card and gives you a complete opinion in one sitting, while a pairwise respondent sees each item about three times and gives you a handful of quick either-or answers. That is what makes it easy to finish on a phone, and it is also why the coverage has to come from more people.

Two situations need more than the standard numbers.

  • A long list. Past about 30 items the comparisons per person hit the fatigue ceiling of 40, so each person sees each item about twice instead of three times. A 40-item study needs roughly 50 per dimension to reach the steadiness a 20-item study reaches at 30.
  • A shortlist where everything matters. If you have already cut the obvious losers and the remaining needs genuinely sit close together, telling them apart is much harder: that case wants 100 or more per dimension before the ranking settles at all. This is the common shape of a second-round study.

One thing pairwise will not give you, at any sample size. It reliably tells you which needs sit in the top tier. It does not reliably tell you which single need is number one, because when two needs are genuinely close, no amount of asking separates them cleanly. Read the top of your ranking as a shortlist of strong candidates rather than as a winner and a runner-up.

What participants see

Participants open the link on any device and, after the optional screener, see two needs at a time and pick one. If your items have pictures, each card shows its picture above the words, and an animation plays. The next pair's pictures are fetched while they are answering the current one, so the study does not stall between comparisons, and the timing we record starts when the pictures are actually on screen rather than when we began fetching them. If a picture fails to load, the card falls back to its text and the study carries on. After a set number of pairs they answer two short 1 to 10 rating questions. It is quick and needs no account.

Still stuck? Sign in and ask Graham, or email support@sortedresearch.com.