← All guides

What a tree test tells you that analytics cannot

"We have analytics, we know where people go." You know where people went. That is not the same thing, and the difference is most of what you need to fix navigation.

Analytics only sees the site you built

Every number in your analytics is conditional on the current structure. You can see that 3 percent of visitors reach the returns page. You cannot see:

  • How many were looking for it. Three percent of visitors reaching returns is excellent if 3 percent wanted it, and a catastrophe if 12 percent did and most gave up.
  • Where the other nine percent went instead. You see their sessions, but not that they were hunting for returns while they were in Help, then Orders, then out.
  • What they would have found in a different structure. Analytics cannot evaluate a menu that does not exist, which makes it useless for deciding a redesign and only useful for grading one afterwards.

Analytics answers "what happened". Navigation research answers "what were they trying to do, and did the words help them".

What a tree test measures instead

You strip the navigation back to words. No design, no images, no search. You give a real task: "your trainers do not fit, find where to send them back." They tap through, one level at a time, until they commit.

Because you set the task, you know the intent. That single fact is what analytics can never give you, and it unlocks everything else:

Success rate per task, against a known denominator. Not 3 percent of everyone, but 62 percent of the people who were definitely looking.

The first section they opened. The strongest single predictor of whether somebody eventually succeeds, and the number that tells you a top-level label is wrong. If half your participants open Help when the answer is under Orders, the problem is the word Orders.

Where the failures went. Not that they failed, but that eleven of them committed to Customer Service, which means Customer Service is making a promise the content does not keep.

Whether they went straight there. Somebody who lands correctly after wandering through four sections had a worse experience than the success rate suggests.

The finding analytics structurally cannot produce

The single most fixable problem in navigation is somebody who opens the right section, doubts the label, and leaves again.

They usually come back. They usually succeed in the end. So they appear in your success rate as a win and in your analytics as a normal session with a bit of back-and-forth. Every aggregate view you have says these people are fine.

They are not fine. They found your answer, read your label, and did not believe it. That is a wording problem with a precise location, and it is cheap to fix once you can see it. A tree test surfaces it because it records every tap, including the ones that went nowhere.

What a score alone will not tell you

A success rate tells you something is broken. It does not tell you where or why. The useful outputs are the ones that keep the individual detail:

  • Every task's traffic drawn on the whole menu at once, so you can see where everybody went before reading a word.
  • The routes people took, one task at a time, with the sections somebody opened and backed out of marked as such.
  • A replay of one person's session, which is the only view that does not aggregate, and the only one where the hunting behind a choice can be watched.

Where analytics wins

Analytics beats tree testing at scale, at reality, and at everything post-launch. It sees every real visitor with real intent, not 25 people with an assigned task. It catches seasonal patterns, device differences and the effect of a campaign.

The right relationship is: analytics tells you where to look, tree testing tells you what is wrong. A page with high traffic and high exits is a question. A tree test is how you answer it before spending a quarter on a redesign.

And afterwards

The other use is proof. Run the test before the redesign to get a baseline, then again after. A success rate going from 48 to 81 percent on the tasks that mattered is the kind of evidence that survives contact with a stakeholder who preferred the old menu.


Related: Card sorting vs tree testing: which one, and when

Run one yourself

The free plan runs a full study with up to 10 participants, and needs no card.