Findability study · tree testing

Can people actually find it?

Findability is tree testing, improved. Show people your navigation as words alone, give them a real task, and watch where they look. You get more than a score: you see what they nearly found, why they walked past it, and, if you want, a menu that improves itself round by round.

Start your free study See how it works
How it works

Your menu, with everything else stripped away.

No design, no page content, no search box. Just the words, so what is being tested is your structure and your labels and nothing else.

1

Bring in your menu

Paste it, drop in a sitemap or spreadsheet, type a URL and we read the page, capture it from your own browser, upload screenshots, or turn a finished card sort into a starting structure.

2

Write real tasks

“You want your money back on a broken kettle.” Mark where the answer lives. Graham checks your wording and tells you when a task gives the answer away.

3

Read what happened

Success rates with the uncertainty shown, the whole menu drawn with every task's traffic on it, the routes people took, and a replay of any one person hunting.

What your participants see

One level at a time, on the phone in their hand.

Tree testing was invented for a desktop screen, and most respondents today are on a phone. So we show the menu the way a phone app does: one level at a time, with a trail back up. Every tap is recorded, including the ones where they change their mind.

  • Built for a thumb, so completion rates hold up
  • They can give up, and giving up is a finding, not a gap
  • We ask how sure they were, which is how we spot a label that misleads
Task 2 of 5 Where would you get your money back on a kettle? ‹ Top level › Help Returns & refunds Delivery help Product care Contact us It's in here Words only. No design, no page content, no search box. Every tap recorded, in order, including the backtracks.
The structure and the labels, with nothing else to lean on.
What you see

A score, and how much to trust it.

Every task gets a success rate with the range it could really be at your sample size, drawn against the bands the research literature actually publishes. The median task in a usability study succeeds about 62% of the time, so a menu in the sixties is ordinary rather than broken, and this says so instead of flattering you.

  • A confidence range on every task, so a small study cannot be over-read
  • A verdict in plain words, including “not enough people to say yet”
  • Slice by screener answer, with a test of whether the difference is real
How each task did 0% 100% TYPICAL TASK (62%) Money back on a kettle 75% When will my patio set arrive 33% Change the card on file 92% Keep the garden alive 50%
The bar is what we measured. The line through it is what it could really be.
The difference

A menu that improves while you test it.

An ordinary tree test tells you your menu scored 48%. Then you guess at a fix, rebuild the study, and buy another sample. The Adaptive Tree closes that loop for you.

Field in rounds. When a round finishes, Graham reads it and proposes changes to the wording, in plain sentences about your own menu. You approve or reject each one. The next round fields the version you approved, and the answer key follows it across, so round two is scored against the menu round two actually saw.

You decide up front what Graham is even allowed to touch: renaming only, or moving things, or merging and splitting. He is checked against that twice, when he proposes and again when it is applied, so a model that ignores its instructions changes nothing.

  • Nothing changes without you. Every refinement waits for your approval
  • A round that finds nothing says so. No invented changes to look busy
  • One audience, sized once. You buy the whole study, not a sample per round
Round 1 Deliveries My orders Help 48% found it Graham proposes Rename “My orders” to “Track an order”. It took first taps on a task it could not answer. Approve Reject Round 2 Deliveries Track an order Help 71% found it What you may claim Round 1 found things 48% of the time and round 2, 71%. That gap holds up at these sample sizes, so it is not noise. Each round saw a different menu.
Two rounds, one approved change, and a difference the numbers can defend.
The map

Every task's traffic, on one picture.

Most tree testing tools draw one task at a time and shade the busy parts. That hides the failure that matters most: one task's answer quietly soaking up another task's traffic. We draw your whole menu once, with every task on it, and colour by outcome rather than volume. Green is people who found what they came for. Red is people who stopped there and were wrong.

  • A busy section is either working or winning wrongly, and you can tell which
  • Branches thicken with the people who walked down them
  • Sections nobody touched stay on the drawing, because the shape is the context
Shop Garden Plants & seeds Tools & watering Deliveries Help My account ↩ 6 left again 4 6 wrong here 4 6 found it 2 4 found it ↩ 4 left again found it here stopped here, wrong ↩ opened it and left again
One task's answer taking another task's traffic, visible at a glance.
Beyond a success rate

What a score on its own will never tell you.

A percentage says how many arrived. These say what happened on the way, which is the part you can act on.

1

First tap, and what came of it

Where somebody goes first is chosen on your labels alone, before any exploring, so it measures your wording more cleanly than anything else. Knowing 60% opened Products first is half a finding. We also show that of those, nine in ten went on to succeed while everyone who started elsewhere failed. That is the difference between a label that is popular and a label that is right.

2

The ones who nearly found it

People who opened the right section, doubted the label and left again. Most of them come back and succeed, so they never appear as a failure anywhere. They are the cheapest thing on the page to fix: the structure is already correct, the wording just is not reassuring enough to stop them second-guessing it.

3

Watch one person hunt

Replay any respondent's whole session, tap by tap, at the speed they did it. The list puts the interesting ones first, so you are not scrolling fifty people to find the one worth watching. A column of numbers says a label is weak. Thirty seconds of somebody going in and out of the wrong section says why, in a way a stakeholder repeats back to you.

Numbers you can defend

It will tell you when it does not know.

A tree test on 20 people can produce a chart that looks decisive and is not. This one refuses to. Every rate carries the range it could really be, the bar for acting on a task is set before anyone looks at the data, and a difference that could be sampling noise is reported as too close to call.

  • Careless respondents are set aside, and still counted in your funnel, so the incidence rate that explains your bill stays honest
  • Rounds are never added together. Four rounds of fifty is not a study of two hundred, because each round saw a different menu
  • Two ways of showing a menu, one level at a time or the whole thing expanded. We do not claim ours is better, so you can run both and settle it on your own audience
! Act on this “When will my patio set arrive” is genuinely worse than a typical task. Even at the optimistic end of its range it is 61%, below the published median. This is one to act on. ? Not enough people to be sure “Keep the garden alive” looks low, but with 12 people it could really be anywhere from 25% to 75%, which is either side of typical. More people would settle it. Changing a label on this evidence risks rewriting something that was working.
The second card is the one most tools will not show you.
Graham

Graham writes the report

When the responses are in, Graham turns the findings into a client-ready insight report: what people could not find, which labels misled them, and what to change first.

  • The headline story, the labels to fix, and clear next steps
  • An honest note on data quality, so you know what to trust
  • Download as Word, PowerPoint, PDF or CSV, and every view exports
Better together

Card sorting designs it. Findability proves it.

A card sort tells you how customers naturally group a messy list, and that is where a good structure comes from. Findability is the other half: it takes the structure you landed on and finds out whether people can actually get to things in it. Turn a finished card sort into a starting menu in one click, then test it.

See card sorting
Card sort groups Getting help Orders & delivery Payment Products A menu to test Getting help Returns & refunds Contact us Orders & delivery
The groups become the top level, ready to nest and test.
When to use Findability

Reach for it when the structure already exists.

Use it before a redesign to find out what is actually broken, after one to prove it improved, or on a competitor's navigation to see what you are up against. If you are still working out how customers group a messy list, start with a card sort and test the result here. If the question is which things matter most rather than where they live, that is pairwise or Best–Worst.

Questions

Tree testing, answered.

What is tree testing?

Tree testing gives people your navigation as words alone, with no visual design, no search box and no page content, and asks them to complete a real task. Because there is nothing else to go on, what you learn is whether the labels and the hierarchy actually work.

What is the difference between tree testing and card sorting?

Card sorting designs a structure: it tells you how people would group things. Tree testing proves one: it tells you whether people can find things in the structure you have. They pair naturally, so most teams sort first and then test.

How many participants do I need for a tree test?

The same rule of thumb applies as for card sorting. Around 15 gives a directional read and 20 to 30 gives a firm one. The confidence indicator on your results tells you when a task's score has settled, so you are not guessing.

Do participants see my actual website?

No. They see your menu as text, one level at a time. That is the whole point of the method: you are testing the words and the hierarchy, not the visual design and not the search box.

What is an adaptive tree test?

Findability can propose changes to your menu between rounds, based on where people actually went wrong, then retest with the revised menu once you approve it. Instead of a single score you get a menu that improves round by round.

What do I get beyond a success score?

Every task's traffic drawn on your whole menu at once, the first taps people made, the routes they took, and a replay of any individual session. A score tells you something is broken. This tells you where, and why.

Test your menu, free

Set up a Findability study today. Bring your own participants, or buy a matched audience in a few clicks.

Start your free study