The method has held up for sixty years.

The tools kept changing. The reason card sorting works did not. This is the thinking SortedResearch is built on.

The origin

A method built from the field up

In the 1960s a Japanese ethnographer, Jiro Kawakita, built a way of working he called the KJ Method, set out across three books between 1967 and 1986. His premise was that real understanding of a subject comes from the field itself, and not from a theory laid over the top of it.

He called this field science. The tool was plain: one idea per card, spread on a table, grouped by hand until the patterns surfaced. From there the method spread through product development, planning and quality work across Japan, and then well beyond it.

A pile of cards Sorted into groups Ranked on a matrix

One idea per card. Card sorting groups them by affinity, the way Kawakita did by hand. Pairwise then ranks those groups by importance and satisfaction, so the biggest opportunities rise to the top.

What we preserve

What SortedResearch keeps from Kawakita

Three ideas carry straight over from the table of cards to the software.

01

Structure emerges from the data

Open sorts hand participants a deck with no groups set in advance. The structure comes out of the data. This is Kawakita's abductive principle: let the field speak before the researcher interprets.

02

Affinity as the unit of analysis

Kawakita's table encoded closeness as meaning. Cards that sat together belonged together. The similarity matrix and clustering dendrogram in SortedResearch are the computational form of that same idea.

03

Spatial structure reveals what lists hide

Kawakita watched how the groups sat in relation to each other, beyond which card went where. The 3D view in SortedResearch makes that spatial structure legible at scale, across hundreds of participants, in seconds.

What we extend

Where the tool goes further

Two places SortedResearch moves past what Kawakita could do at a table.

Scale

Kawakita worked with small teams in long, immersive sessions. SortedResearch aggregates across large consumer samples, trading depth of immersion for strength of pattern. Both produce valid knowledge. They answer different questions.

Language

Graham reads the pattern and drafts what it means. Kawakita held that the insight came through the act of grouping itself. Graham does not replace that judgment. He gives the researcher a faster place to start, and something to push against.

The second question

They agreed where it belongs. Can they find it?

Grouping and finding are not the same test. A card sort tells you where people think something belongs when they are looking straight at it. It does not tell you whether they will walk past it in a real menu, where the label has to compete with everything sitting beside it. Tree testing asks that second question, and asks it about the words alone.

The structure, with nothing to rescue it

People are given the navigation as plain text: no design, no colour, no search box, no page content. Anything that would let somebody succeed in spite of the structure is taken away, so what is being measured is the structure and the wording and nothing else. They are given a realistic task and asked where they would go looking for it.

The first tap says the most

The section somebody opens first is chosen on the labels alone, before any exploring, so it reads your wording more cleanly than the success rate does. It is also the strongest predictor of whether they end up in the right place. A label that pulls people in and cannot deliver costs more than one nobody notices, because they have already decided they are somewhere sensible.

The third question

From how they relate to what matters most

Card sorting answers one question: how do people group your ideas? Pairwise answers the next one. Which of those ideas matter most, and where are people underserved? It rests on two simple, well-established ideas.

A choice is steadier than a rating

Ask someone to score fifteen needs out of ten and the numbers drift. Ask them to pick between two needs at a time and the answer holds. Simple forced choices, repeated and aggregated across people, give a ranking you can trust far more than a long rating grid.

Importance is only half the picture

A need that matters is only an opportunity if it is poorly met. Measuring how much each need matters and how well it is served today, then reading the gap between the two, is a long-standing way to tell the real opportunities from the table stakes. That gap is what the opportunity matrix draws.

The fourth question

And across a long list, what rises to the top?

Pairwise weighs a focused set of needs, two at a time. When the list runs to twenty or thirty items, asking about every pair is more than anyone will sit through. Best–Worst keeps the same forced choice and scales it, by asking only for the extremes.

One screen, many comparisons

Show a handful of items and ask for the most and the least important. That single answer settles a whole cluster of comparisons at once: the best beats everything else shown, and everything else shown beats the worst. A few short screens rank a long list without ever putting every pair in front of anyone.

A 0 to 100 scale

Every item is shown the same number of times, then scored by how often it is picked best against how often it is picked worst. That gives each one a score from 0 to 100, so you can read the distance between items and see where the line falls between the must-haves and the rest.

One method, four questions

How the four fit together

The four study types share one line of enquiry. Each asks a question that follows from the last.

Card sorting opens the space. It takes a messy list and shows how people group it, so the real themes surface from the data itself. That is where a category first becomes legible.

Findability closes that loop. A structure people agreed with on cards still has to survive being read one level at a time, with nothing but the words to go on, so the same enquiry is put back to them as a task: where would you go looking for this? What they agreed with and what they can navigate are two different findings, and only the second one ships.

From there the question turns to priority. Pairwise takes a focused set of those needs and weighs each for importance against how well it is served today, so the underserved opportunities separate from the table stakes. Best–Worst takes a longer list and ranks the whole of it by importance alone, quickly, and on a scale you can measure distances from.

Run them in sequence, from structure to priority, or reach for the single one that fits the question in front of you. Each builds on the conviction Kawakita started from: let people show you the structure, and build on the choices they find easy to make.

The method is proven. The tools are new.

See the structure in your own category.

Run your first study free