Findability is tree testing, improved. Show people your navigation as words alone, give them a real task, and watch where they look. You get more than a score: you see what they nearly found, why they walked past it, and, if you want, a menu that improves itself round by round.
No design, no page content, no search box. Just the words, so what is being tested is your structure and your labels and nothing else.
Paste it, drop in a sitemap or spreadsheet, type a URL and we read the page, capture it from your own browser, upload screenshots, or turn a finished card sort into a starting structure.
“You want your money back on a broken kettle.” Mark where the answer lives. Graham checks your wording and tells you when a task gives the answer away.
Success rates with the uncertainty shown, the whole menu drawn with every task's traffic on it, the routes people took, and a replay of any one person hunting.
Tree testing was invented for a desktop screen, and most respondents today are on a phone. So we show the menu the way a phone app does: one level at a time, with a trail back up. Every tap is recorded, including the ones where they change their mind.
Every task gets a success rate with the range it could really be at your sample size, drawn against the bands the research literature actually publishes. The median task in a usability study succeeds about 62% of the time, so a menu in the sixties is ordinary rather than broken, and this says so instead of flattering you.
An ordinary tree test tells you your menu scored 48%. Then you guess at a fix, rebuild the study, and buy another sample. The Adaptive Tree closes that loop for you.
Field in rounds. When a round finishes, Graham reads it and proposes changes to the wording, in plain sentences about your own menu. You approve or reject each one. The next round fields the version you approved, and the answer key follows it across, so round two is scored against the menu round two actually saw.
You decide up front what Graham is even allowed to touch: renaming only, or moving things, or merging and splitting. He is checked against that twice, when he proposes and again when it is applied, so a model that ignores its instructions changes nothing.
Most tree testing tools draw one task at a time and shade the busy parts. That hides the failure that matters most: one task's answer quietly soaking up another task's traffic. We draw your whole menu once, with every task on it, and colour by outcome rather than volume. Green is people who found what they came for. Red is people who stopped there and were wrong.
A percentage says how many arrived. These say what happened on the way, which is the part you can act on.
Where somebody goes first is chosen on your labels alone, before any exploring, so it measures your wording more cleanly than anything else. Knowing 60% opened Products first is half a finding. We also show that of those, nine in ten went on to succeed while everyone who started elsewhere failed. That is the difference between a label that is popular and a label that is right.
People who opened the right section, doubted the label and left again. Most of them come back and succeed, so they never appear as a failure anywhere. They are the cheapest thing on the page to fix: the structure is already correct, the wording just is not reassuring enough to stop them second-guessing it.
Replay any respondent's whole session, tap by tap, at the speed they did it. The list puts the interesting ones first, so you are not scrolling fifty people to find the one worth watching. A column of numbers says a label is weak. Thirty seconds of somebody going in and out of the wrong section says why, in a way a stakeholder repeats back to you.
A tree test on 20 people can produce a chart that looks decisive and is not. This one refuses to. Every rate carries the range it could really be, the bar for acting on a task is set before anyone looks at the data, and a difference that could be sampling noise is reported as too close to call.
When the responses are in, Graham turns the findings into a client-ready insight report: what people could not find, which labels misled them, and what to change first.
A card sort tells you how customers naturally group a messy list, and that is where a good structure comes from. Findability is the other half: it takes the structure you landed on and finds out whether people can actually get to things in it. Turn a finished card sort into a starting menu in one click, then test it.
See card sortingUse it before a redesign to find out what is actually broken, after one to prove it improved, or on a competitor's navigation to see what you are up against. If you are still working out how customers group a messy list, start with a card sort and test the result here. If the question is which things matter most rather than where they live, that is pairwise or Best–Worst.
Tree testing gives people your navigation as words alone, with no visual design, no search box and no page content, and asks them to complete a real task. Because there is nothing else to go on, what you learn is whether the labels and the hierarchy actually work.
Card sorting designs a structure: it tells you how people would group things. Tree testing proves one: it tells you whether people can find things in the structure you have. They pair naturally, so most teams sort first and then test.
The same rule of thumb applies as for card sorting. Around 15 gives a directional read and 20 to 30 gives a firm one. The confidence indicator on your results tells you when a task's score has settled, so you are not guessing.
No. They see your menu as text, one level at a time. That is the whole point of the method: you are testing the words and the hierarchy, not the visual design and not the search box.
Findability can propose changes to your menu between rounds, based on where people actually went wrong, then retest with the revised menu once you approve it. Instead of a single score you get a menu that improves round by round.
Every task's traffic drawn on your whole menu at once, the first taps people made, the routes they took, and a replay of any individual session. A score tells you something is broken. This tells you where, and why.
Set up a Findability study today. Bring your own participants, or buy a matched audience in a few clicks.
Start your free study