Findability projects (tree testing)
Findability is the fourth kind of project you can run in SortedResearch, alongside card sorts, pairwise projects, and Best–Worst projects. It answers one question: can people find things in your menu?
You give us your navigation as plain text, no design, no page content, just the words. Participants get a realistic task ("you've got a broken kettle and want to send it back") and tap through the structure until they land where they think it lives. We record every tap, so you find out not only whether people got there, but where they went instead, how many backtracked, and which labels sent them the wrong way.
Because there is no visual design in the way, a Findability study tests the structure and the wording, and nothing else. If people get lost, it is the menu, not the colours.
When to use it
- You have a menu, sitemap, or settings structure and you want to know whether it works.
- You have just run a card sort and want to test the structure it suggested.
- Two people on your team disagree about how the navigation should be organized.
- You are about to rebuild the navigation and want a before-and-after number.
Reach for a card sort instead when you don't have a structure yet and want people to invent one. The two pair naturally: a card sort asks people to build a menu, a Findability study asks whether that menu works. You can start a Findability project straight from a finished card sort.
The three methods
When you create the project you pick one of three. They share the same runner, screener, audience, and results, so the difference is the question you are asking, not the machinery.
Findability
One menu, one round. The straightforward version: give people your structure and a handful of tasks, and get a clear answer on which labels are pulling their weight. Works from about 30 people; 40 to 60 is the sweet spot.
Findability A:B
Two versions, the same tasks, and the sample split evenly between them. This is the one that settles the argument. You need roughly 100 people, because each version needs about 50 of its own.
On the Build tab you choose what the two versions differ by:
- Two menus. Two structures, one question: which organizes better. This is the usual choice.
- Two ways of showing one menu. One structure, shown two ways: our one-level-at-a-time view against the conventional indented outline. Because both halves saw the same structure, any difference has to be the presentation.
Adaptive Tree
A study that improves while it runs. Four rounds of fifty by default. After each round closes, Graham reads where people actually got lost and proposes changes to the structure. You approve them before the next fifty see the revision. You end with navigation that has been rebuilt around real behaviour, and the evidence for every change.
An important piece of honesty about the numbers: four rounds means three refinements, so the last fifty people see a structure that nothing changes afterwards. That final round is a clean check on the structure you leave with. It is not right to say "200 people tested your menu", because each round saw a different menu. The two defensible numbers are the first round as a baseline and the last as the validated result.
Getting your menu in
You rarely have to type it. On the Build tab you can:
- Paste an indented list, a CSV, a sitemap, or a chunk of HTML. We work out which one you gave us.
- Type a website address and we read the navigation off the page.
- Drop screenshots on the page, including of a phone or desktop app, and Graham reads the menu out of them. This is the only route that reaches native apps.
- Drag a bookmark into your browser bar and click it on any page, including one behind a login, to capture that page's navigation.
- Start from a card sort you have already run: its groups become the menu.
Whatever the source, you can then edit it by hand, and drag rows to nest them.
Tasks
A task is what you ask people to find, written as a situation rather than as a label. "You want to send back a broken kettle" is a good task. "Find Returns and refunds" is not, because it hands over the answer, and we warn you when a task's wording contains the words of its own answer.
Each task needs a correct answer: the section (or sections, more than one is allowed) that counts as right. A task with no correct answer can never be answered right, so we refuse to publish until every task has one. Five to eight tasks is the usual range; you need at least two.
Your menu and your tasks are locked once the first response arrives, because the results reference them. Duplicate the project to test a different structure.
What you get back
- Found it: the share who landed in the right place, per task, with the range the true number plausibly sits in. A published benchmark: the median task in a usability study succeeds about 62% of the time, and 61 to 80% is considered good. We show which band each task falls in, so an ordinary menu is not reported as a disaster.
- Went straight there: the share who reached the answer without backtracking. Healthy is about 75%. A high success rate with low directness means people get there eventually, by a route that fights them.
- Where people went instead: the wrong destinations, worst first, with the full path. This is the number you act on.
- Misleading labels: the people who committed to a wrong answer and said they were certain. The most expensive kind of broken, because they never double back, so they never find it and they never complain.
- First-click magnets: a top-level section that is pulling people off the path. Its label is over-promising.
Everything is shown with its uncertainty. At fifty people a task that scored 40% might really be anywhere from about 26% to 55%, and we say so rather than reporting the 40% as a fact.
How Graham decides what to change (Adaptive Tree)
Worth knowing, because it is the part people ask about.
Graham never decides what is broken. That is worked out in code, before he sees anything, and the bar is deliberately high: a task only reaches him if we can be confident its true success rate is below the published median, or if a meaningful share committed to one wrong destination while saying they were certain, or if a section is confidently pulling first clicks away. Because a round judges every task at once, the confidence level is shared across them, so a study with eight tasks does not throw up a false alarm every round.
The consequence is deliberate: some rounds produce no changes at all, and that is an honest result rather than a failure. Changing a healthy label and then reporting the inevitable bounce-back as an improvement is the exact failure this is built to avoid.
What Graham is allowed to propose is up to you, and you set it on the Rounds tab before the study starts:
- Rename a section (on by default). The safest: the section stays where it is, so your tasks keep their answers.
- Move a section (on by default). Usually the highest-value fix, because the results already show where people expected to find things.
- Merge or split sections (off by default). Reshapes the menu, which makes round-to-round comparison harder to read.
- Add or remove sections (off by default). The most aggressive: Graham inventing navigation rather than repairing yours.
These are fixed once round one opens, along with the number of rounds and their size. They have to be: if the rules could change halfway through, a difference between two rounds might be your menu or might be the rules, and nobody could tell you which.
The Rounds tab (Adaptive Tree only)
One screen showing where the study is and the single next thing to do:
- Round N is collecting. It closes on its own the moment it is full. You can close it early if recruiting has stalled; Graham is told the real number either way, and a smaller round simply makes him more cautious about what counts as a problem.
- Round N is finished. Ask Graham to read it.
- Graham's proposal. His reasoning in his own words, then each change as a sentence about your own menu, with the evidence behind it. Every change starts ticked; untick anything you don't want. Approving opens the next round on the revision. You can also reject the lot and stop there, which leaves your menu exactly as the last round tested it and refunds the rounds you didn't use.
Between rounds the study is genuinely not taking people, and recruiting is paused while you decide. If you bought your audience from us, anyone who arrives in the gap is sent back to the panel rather than turned away, so you are not charged for them.
Two ways of showing a menu
Participants see your structure one of two ways, and you choose on the Build tab:
- One level at a time (the default). Full width, with a breadcrumb showing where they are, the way your phone's own settings app works.
- The conventional indented outline. Every level stays on screen under its parent. This is what tree testing has used since it was a desktop technique.
We default to the first because tree testing was invented for a desktop screen and most of your participants are on a phone, where every level of indent eats the width your labels need. We do not claim it is proven better, and the published evidence does not settle it. So both are offered on every study, every response records which one it used, and an A:B study can test the two against each other directly.
One number to read carefully if you do: how many went straight there is not comparable between the two presentations. Going back up a level means something different when every level is already on the screen. Success rate is the fair comparison.
Quality checks
From the Participants view you can download one participant's full record as a CSV, taps included: their screener answers, each task with where they ended up and which section they opened first, and then the trail itself, one row per tap with the full section path. You can also permanently remove that participant, which deletes every tap along with their answers, and the results recompute without them. Between them those two cover a GDPR access request and an erasure request for one person.
Every response is checked, and a careless one is set aside from the findings while still counting in your funnel. Findability carries more checks than any other study type, eleven in all, because a menu test leaves a much richer trail than a sort does. Any one of them sets a response aside:
- Nothing answered: reached the end without answering a single task.
- Speeder: answered faster than the tasks could be read.
- Too many tasks too fast: answered most individual tasks faster than they could be read. This catches somebody who does two tasks properly and then machine-guns the rest, which the median check alone would miss.
- Same answer every time: chose the same section for every task, mostly wrongly, without opening anything.
- Impossible timings: claimed more time than the visit could actually have contained.
- Never opened a menu: answered tasks whose answer was nested without opening, scrolling or pausing on anything.
- All skips: gave up on most tasks without opening, scrolling or pausing on anything.
- Skipped without trying: the same thing on several tasks rather than most.
- No exploration: chose from the first screen without opening a section, on most tasks.
- Branch fixation: opened the same section first on most tasks, and most of those answers were wrong.
- Flat confidence: gave the same confidence rating every time, and was mostly wrong.
Giving up is a real finding and stays payable, so the skip checks only fire when somebody gave up without ever looking. The shape-sensitive checks (no exploration, branch fixation, flat confidence) need enough tasks to be trustworthy, because on a shallow tree an honest respondent would otherwise trip them.
One further signal flags without setting anything aside: somebody far slower than the study's median moves up the Respondents list with that reason against them, and nothing happens to their data or their payment.
If you bought your audience from us, a rejected respondent is sent back to the panel as a quality reject, so you are refunded and replaced.
This matters more here than elsewhere. In the other project types a careless response drags an average slightly. In an Adaptive Tree study the round's data drives a change to your menu, so one person tapping the same box could otherwise invent a problem that Graham then proposes a fix for.
The insight report
Once anyone has finished, the Results tab has a Report view: Graham turns your results into a client-ready write-up. It leads with the headline, then where people get lost, then what is already working, then what to change first, and finishes with the task-by-task evidence.
Two things it does that a generic summary would not. Every task carries a verdict of Problem or Normal, and a task marked Normal performed within the range a typical task does whatever its percentage looks like, so nobody rewrites a healthy label because 64% sounded low. And on an Adaptive Tree study the report carries the round-by-round table, so the history of what you changed travels with the document instead of living only on a screen.
The report is written once and kept, so re-opening it is instant and costs nothing. It is rewritten automatically whenever your data moves: a new response, a quality verdict changing after the fact, or a new round's menu appearing.
On a paid plan you can export it to Print/PDF, Word, and PowerPoint. On Free it renders on screen.
Study settings
On the Build tab you can change the study title and description, choose how the menu is shown, and write what participants read. Everything is optional except the title; leave a piece of copy empty and we use our own wording, which is written to be clear on a phone and to tell people we are testing the menu rather than them.
- Intro instructions replace our explanation on the first screen.
- Consent text appears before they start, for organizations with their own wording. Consent is recorded per participant either way.
- Honesty note is a short please-answer-honestly line. Optional on purpose: priming is a real intervention and it is yours to decide on, not ours.
- Redirect URL is where a participant goes at the end, for anyone recruiting through their own panel. It has to start with http:// or https://. Respondents we recruit for you go back to the panel instead, so it is ignored for them.
- Ask how sure they were is on by default. Turning it off saves about two seconds a task and removes the misleading-labels findings entirely, because those depend on knowing who was certain.
- Shuffle the task order spreads order effects across your sample. Off by default so your first run reads in the order you wrote.
The title, description and redirect can be changed at any time. How the menu is shown and what is asked of people are frozen once responses exist, because changing them mid-study would put two different experiences in one column of results.
The journeys view
The Results tab has a Journeys view showing the actual routes people took, one task at a time. Every tap is recorded, so this is the route itself rather than a summary of it. Sections somebody opened and then left again are marked as backed out.
The number to read first is how many opened the right section and left again. Those people are invisible everywhere else in your results, because most of them come back and succeed, and a success rate cannot tell you they nearly did not. It is also the cheapest thing on the page to fix: the answer is already where you put it, and the wording is not reassuring enough to stop people second-guessing themselves.
The map
The Results tab has a Map view: your whole menu drawn as a tree, with every task's traffic on it at once. Colour is the outcome, not the volume, so green is people who found what they were after in that section and red is people who stopped there and were wrong. Branches thicken with the number of people who walked down them, so the shape tells you where everybody went before you read a word.
Sections people opened and then left again are marked, and the answer for each task is flagged. Use the task chips above the map to look at one task at a time, or leave it on all of them to see the menu as a whole.
First tap
The first section somebody opens is chosen on your labels alone, before any exploring, so it reads your wording more cleanly than the success rate does. It is also the strongest single predictor of whether they end up in the right place.
The First tap view shows, per task, where people started and how the people who started there went on to do. That second half is the point: a popular label and a right label are not the same finding. A section that takes a quarter or more of the opening taps and is not on the way to the answer is flagged as pulling people in and not being able to deliver, but only once there are enough people for that to mean anything.
Watching one person
The Participants view lists everyone who finished, one row each, numbered rather than identified. They are ranked so the ones worth watching come first: people who opened the right section and left again, then people who did not find it, then people who took far longer than the rest.
Open one and you can replay their whole session, tap by tap, across every task they did. This is the thing the rates cannot show you. Somebody who opens the right section, doubts the label and leaves is invisible in every other view, because most of them come back and succeed.
Downloading what you see
Every Results view has its own Download CSV on a paid plan, not just the raw data: the task table, first tap, the map's traffic, the journeys, the participant list, and round by round. Each one downloads what is on the screen in front of you, with the same numbers, so a file cannot disagree with the page it came from.
Your raw data
The Results tab shows your recruitment funnel: how many arrived, how many the screener turned away, how many finished, and how many the quality checks set aside. The share of screened people who qualified is your incidence rate, and on a bought audience it is the number that explains your bill.
On a paid plan you can download the raw data as a CSV: one row per person per task, with what they chose, the full path to it, whether they went straight there, which section they opened first, how long they took, and how sure they said they were. Participants are numbered rather than identified.
Round by round (Adaptive Tree)
An Adaptive Tree study gets an extra Results view answering the question the method exists for: did it work?
It compares the first round against the last, because those are the two that mean something. Round one is the structure you walked in with; the final round is the one nothing changed afterwards, so it is a clean check on what you leave with. The rounds in between each measured a menu that was then altered, so they are shown as progress and never quoted as the result.
The comparison is allowed to say too close to call, and often does. It is the same test the A:B comparison uses, so a study cannot call a difference real in one place and refuse to act on it two clicks away. It is also allowed to tell you the menu got worse, which matters: a tool that only ever finds improvements is a tool that invents them.
Underneath, each round shows its own numbers and exactly which changes were applied before it fielded, in plain English, with Graham's reason for each. So "we moved returns under My account and success went from 41% to 78%" is something you can read off the screen rather than reconstruct.
The view also states, in writing, the claim you may not make: four rounds of fifty is not "200 people tested your menu". Each round saw a different menu, so the rounds are not one sample.
Being told when a round finishes
An Adaptive Tree study stops between rounds on purpose, and it stops without warning you, because a round closes the moment it fills. So when a round closes we email whoever created the study: this round is done, your study is paused until you look, here is the link. When the last round closes you get a different email saying the study is finished.
Recruiting is paused while you decide, so a study waiting on you is not spending anything.
How many people do I need?
- Findability: from 30, with 40 to 60 the sweet spot.
- Findability A:B: about 100, because each version needs 50 of its own.
- Adaptive Tree: 50 per round, four rounds by default, paid once with unused rounds refunded.
A bigger menu needs more people per round, simply because there are more places to get lost.