UXR InstituteUX ResearchCard Sorting
Method Guide
Information Architecture

Card Sorting: How to Design Navigation Users Actually Understand

A large part of usability is finding something where you expect to find it. Card sorting is the research method that helps organize information to align with users' mental models.

Card sorting is a UX research method where people organize labeled cards into groups that make sense to them. Product teams use it to help design menus, page groupings, and content structures around users' actual mental models. Organizing information close to the way users expect improves usability and user satisfaction. Doing card sorting early in the development process also prevents structural flaws in a design that are more difficult to fix later.

Key takeaways

  • Card sorting surfaces how users group content, which supports the creation of information architecture (IA) that aligns with how users think.
  • The data shows which cards landed together, not why. Two identical piles can come from opposite reasoning.
  • Capture the reasoning as you go, with think-aloud or a short post-sort interview.
  • Open, closed, and hybrid sorts answer different questions: discover a structure, test one, or find its gaps.
  • A sort produces a hypothesis. Only a tree test shows whether people can find anything in it.

A familiar card-sort failure looks like this. A product designer at a payroll company asks customers to group cards about payroll, taxes, employees, and benefits. Many customers put "tax filings" near the payroll cards, so the team decides tax content belongs under Payroll. After launch, support starts hearing the same question again: "Where do I find my tax filings?" The customers had grouped tax with payroll because paychecks and taxes are related in their work. That did not mean they would look for tax filings inside a Payroll menu. The grouping was real, but the team didn't elicit the "why" behind the grouping that would help them understand participants' thinking. That is often the key to actionable insight from card sort activities.

What a card sort actually shows you about how people think

You are not asking participants to design your menu. You are watching which piles feel obvious to them, which cards make them pause, and which labels they reach for without thinking. A sorting task is useful because people often cannot explain their categories until they have to put something in one.

The method works because sorting makes hidden knowledge visible. When Chi, Feltovich and Glaser gave physics problems to experts and novices and asked them to sort, the novices grouped by literal features like inclined planes, while the experts grouped by the principle needed to solve the problem, such as conservation of energy (Chi, Feltovich & Glaser, 1981)Chi, Feltovich & Glaser, Categorization and Representation of Physics Problems by Experts and Novices, Cognitive Science, 1981. Same cards, same instructions, different piles. The difference was what each group noticed first.

That is what a sort gives you: visible piles that point back to the rules people are using. Rugg and McGeorge describe sorting techniques as a way of finding out how people categorise their world, and treat them as knowledge elicitation rather than menu design (Rugg & McGeorge, 2005)Rugg & McGeorge, The sorting techniques: a tutorial paper on card sorts, picture sorts and item sorts, Expert Systems, 2005 (orig. 1997). The lineage runs back further, to George Kelly's use of sorting to elicit the personal constructs an individual uses to organize a domain (Kelly, 1955)Kelly, The Psychology of Personal Constructs, Norton, 1955.

Nielsen Norman Group's own guidance points the same way. Their card sorting article is titled "Uncover Users' Mental Models," and it defines the method as participants placing cards into groups "according to criteria that make the most sense to them" (Nielsen Norman Group, 2024)Nielsen Norman Group, Card Sorting: Uncover Users' Mental Models, 2024.

That distinction changes how you use the results. If customers and administrators sort the same cards differently, you do not average the two structures and call it consensus. You have learned that the two groups walk into the product with different starting points. That can change the navigation, but it can also change search synonyms, onboarding order, help-center taxonomy, and the glossary your own team uses.

A card sort tells you what people grouped. It cannot tell you why.

Two participants can make the same pile for different reasons. One puts "tax filings" with payroll because taxes are part of paying employees. Another puts it there because payroll is the only place they have ever seen a tax form in your product. Your similarity matrix treats those as the same answer. It only knows that two cards landed together.

That is why a high agreement score is not proof that participants meant the same thing. It proves they made the same move. Those two facts come apart more often than the tidy chart suggests.

Labels make the risk worse because words feel like explanations. A participant can accept one of your category names, sort confidently into it, and still mean something by that name that your team would never assume.

Field Note I was running a card sort with employee benefit brokers and noticed they kept using a term in a way that did not fit what I expected. They talked about "implementation" as though it covered an enormous amount of ground. I probed on it, and found that for brokers, implementation named the entire process of getting a client set up, from the sold case all the way through to employees being enrolled and able to use their benefits. The company used the same word for one step inside that process.

Nobody was confused in the moment. That is what made it dangerous. Both sides had a word they used fluently and neither had any reason to suspect the other meant something different by it. In the sorting data alone, it was invisible. The cards went where they went. It only surfaced because I asked someone to keep talking while they sorted, and then followed up on a phrasing that sounded slightly off.Leo Hoar

The sorting data would have looked clean. The cards landed together under a label everyone recognized. The problem was the scope of the word. For brokers, "implementation" meant a long client setup process. For the company, it meant one step inside that process. The only way to catch that difference was to ask someone what belonged inside the category.

So what? Agreement tells you where people made the same move. It does not tell you whether they made it for the same reason. Treat every strong cluster as a prompt: ask why those cards belong together.

The Thick Card Sort

The pile is the starting point. The explanation is what makes it useful.

Most card-sorting tools tell you which cards were grouped together, how often that happened, and where the strongest clusters sit. That is the similarity matrix, the dendrogram, and the agreement percentage. Those numbers are useful. They tell you what happened. They do not tell you what the participant thought the pile meant.

A Thick Card Sort adds the participant's own explanation of the piles, captured through think-aloud during the sort or a short post-sort interview with a subset of participants. The piles suggest a taxonomy. The explanation tells you what the taxonomy would need to mean in the product. Without that layer, the team is left naming clusters from the outside.

The distinction has a serious name. In anthropology, Clifford Geertz drew the line between recording a behavior and interpreting its meaning as thin versus thick description, adapting a distinction from the philosopher Gilbert Ryle (Geertz, 1973)Geertz, Thick Description: Toward an Interpretive Theory of Culture, in The Interpretation of Cultures, Basic Books, 1973, pp. 3-30. Ryle's example is simple: two boys contract an eyelid. One has a twitch. The other is winking. A camera sees the same movement. A person who understands the situation sees two different acts (Ryle, 1968)Ryle, The Thinking of Thoughts: What Is 'Le Penseur' Doing?, University Lectures no. 18, University of Saskatchewan, 1968.

A thin sort records the twitch. A thick sort recovers the wink.

The benefit-broker example works the same way. The thin record says the card went under "implementation." The thick record says brokers and the company used that word for different amounts of work.

[diagram pending: two-panel thin sort vs thick sort figure]

Open, closed, and hybrid card sorting answer different questions

The three types look like formats. They are really three study questions. Pick the wrong one and the data can still look clean, but it will answer the wrong problem.

Open card sorting

Participants group the cards and name the categories themselves. Use an open sort when you do not trust the current structure, when the domain is new to you, or when the product still reflects the org chart. It is the version that can show you a category your team would not have created.

Do not copy participant category names straight into the navigation. People name piles quickly, with no style guide and no need to make the label work in a real interface. Read their labels as clues. If several people call a pile "money stuff," the final label probably should not be "money stuff," but you have learned which cards they experience as financial tasks.

Closed card sorting

You supply the categories and participants place cards into them. Use a closed sort when you already have a proposed structure and need to know where it breaks. It is good for checking fit, especially when a stakeholder says, "This menu already makes sense."

Closed sorts force a choice unless you give people an escape hatch. A participant who has no good home for a card will still drop it somewhere, and that placement can look just like a confident one. Add a "does not fit" pile, a confidence rating, or a prompt to say when a placement feels forced.

Hybrid card sorting

Participants start with your categories and can create their own. Use a hybrid sort when you mostly believe the structure, but you suspect it has gaps. Every new category a participant creates is a useful complaint: "your menu did not give me a place for this."

Open, closed, and hybrid card sorting compared
What it asks What it produces Best when
Open How would you group this if we gave you no menu? New groupings, rough labels, and vocabulary clues The structure is undecided or the current one feels internal
Closed Can our categories hold these cards? Fit, misfit, and forced placements A structure exists and needs a stress test
Hybrid Where does our structure run out of room? Fit, plus the missing categories people create You partly believe your scheme and want to find its gaps

How to run a card sort that captures the reasoning

The mechanics are simple. The study succeeds or fails in the setup.

  1. Pilot the cards before launch, then keep the live deck stable. If the same three cards confuse everyone in the first real sessions, do not rewrite them and mix those results with the old deck. Pause, fix the cards, and treat the next run as a new phase. The ONRR case study on Digital.gov makes the lesson concrete: unclear card labels cost participants time, overwhelmed some users, and led the team to commit to better labels and consider fewer cards for the next sort (Digital.gov, ONRR case study)Office of Natural Resources Revenue, Open-source Information Architecture Design, Digital.gov, 2022. Early review is there to catch a broken deck, not to mutate the study while pretending the data stayed comparable.
  2. Choose the right level of granularity for the cards. Each card should be one thing a participant can sort: a task, topic, page, product, or content item. Keep the deck at roughly the same level. Do not mix a broad bucket like "Resources" with a specific task like "download last year's tax form" unless that mismatch is what you are intentionally testing. Digital.gov describes both open and closed sorts as giving people content cards, then asking them to group those cards or place them into predefined categories (Digital.gov, Card sorts)Digital.gov, Card sorts, Research & Collaboration guide. In an open sort, broad bucket cards push people toward your proposed navigation. In a closed sort, those broad terms usually belong as the categories people sort cards into.
  3. Keep the deck small enough for attention. An 80-card deck is not automatically more rigorous. It is often just more tiring. Digital NSW puts a number on it: limit the deck to 30 to 40 cards at the most, especially for an open sort, and stay mindful of participant fatigue (Digital NSW)Digital NSW, Card sorting, Digital Service Toolkit. UXPA's Usability Body of Knowledge sets the outer bound wider, noting that recommendations range from about 20 to 200 cards for a sort that usually takes one to two hours, and that past 200 you need some form of sampling rather than a longer session (UXPA, 2009)Killam, Preston, McHarg & Wilson, Card Sorting, Usability Body of Knowledge, User Experience Professionals' Association, 2009. It also names the variable most guidance leaves out: participants who already know the content can handle a bigger deck than people meeting it for the first time. So the ceiling is not a fixed number. If the deck is getting large, first remove duplicates, internal-only items, and cards that sit at a different level from the rest. Then size what is left for the people you are actually recruiting.
  4. Use sample-size guidance as an anchor, not a law. Tullis and Wood found that the similarity matrices from 20 to 30 participants correlated about .93 to .95 with the one from their full 168-participant sample, with little further gain beyond that (Tullis & Wood, 2004)Tullis & Wood, How Many Users Are Enough for a Card-Sorting Study?, Proceedings of UPA 2004. A closed-sort case study with 191 participants, 45 cards, and 7 predefined categories found that 8 to 12 participants were enough for reliable results in that e-commerce context (Dougalis & Katsanos, 2025)Dougalis & Katsanos, How Many Participants Do You Need for Closed Card Sorting? A Case Study of an E-commerce Website, CHIGreece 2025, ACM. Do not spend the whole budget chasing a cleaner dendrogram. Spend some of it hearing why people made the piles.
  5. Use a think-aloud protocol, the same way you would in a usability test. Ask participants to say what they are noticing, what feels obvious, and where they are unsure while they sort. Do not explain the cards or defend the categories; prompt lightly with questions like "What made that card fit there?" or "What are you deciding between?" Digital.gov includes asking why in both open and closed sorts, and Digital NSW tells facilitators to have participants talk out loud so the team can hear their rationale and frustrations. If most sessions have to be unmoderated, moderate a subset with think-aloud or add short explanation prompts after each group. A post-sort interview can help, and the probing discipline you would bring to user interviews transfers directly, but it is second best because some reasoning gets reconstructed after the decision has already been made.
  6. Probe the vocabulary specifically. When a participant uses one of your terms, ask what belongs inside it and what does not. Righi and colleagues make the same point from the analysis side: final labels require understanding what categories participants created and why they put those items there (Righi et al., 2013)Righi, James, Beasley, Day, Fox, Gieber, Howe & Ruby, Card Sort Analysis Best Practices, Journal of Usability Studies, 8(3), 2013, pp. 69-89. This is the move that surfaced the brokers' "implementation." It takes about fifteen seconds and can prevent months of shared-word, different-meaning confusion.

Tool choice rarely rescues a weak protocol. Mainstream card-sorting platforms can randomize cards, collect sorts, and generate the standard matrices and dendrograms. Physical cards still work beautifully in moderated sessions because a hand hovering between two piles is a thing you can ask about. The important tool is the one your protocol often forgets: a place for the participant to explain the pile.

How to analyze card sort results step by step

Card sorting analysis is not "look at the dendrogram and pick the neatest branches." The job is to move from raw participant groupings to a defensible IA hypothesis: these items belong together, this is the evidence, this is what we think the category means, and these are the cards that still worry us.

The workflow below works whether you used a card-sorting platform, sticky notes, a whiteboard, or a basic spreadsheet. Righi and colleagues describe card-sort analysis as a sequence of cleaning the data, reviewing item-by-item relationships, using dendrograms and matrices, standardizing labels, and then making judgment calls when multiple interpretations are possible (Righi et al., 2013)Righi, James, Beasley, Day, Fox, Gieber, Howe & Ruby, Card Sort Analysis Best Practices, Journal of Usability Studies, 8(3), 2013, pp. 69-89.

  1. Clean the data before you combine it. Review each participant's sort on its own. Look for incomplete sessions, a deck dumped into one giant category, many meaningless labels, or a completion time that suggests the participant did not really do the task. Do not delete a weird response just because it is inconvenient. Decide whether it is low-quality data or a real difference in mental model, and keep a short exclusion note for anything you remove.
  2. Preserve the raw sort. For each participant, keep the original group names, the card numbers or card IDs inside each group, and any comments they made while sorting. Digital NSW says to photograph each participant's results if you used physical cards, and if you did not, to record the names each participant gave each grouping and the numbers of the cards under it (Digital NSW)Digital NSW, Card sorting, Digital Service Toolkit. Either route gets you a complete session record. A simple table is enough: participant, card ID, card label, raw group label, notes.
  3. Group similar participant labels before you analyze them. In an open sort, participants may use different names for the same kind of pile. One person may write "Support," another "Help," and another "Customer service." Put those raw labels into a working analysis bucket, such as "Help and support," so you can compare how cards moved across participants. Joe Lamantia's spreadsheet method makes the same move: open-sort categories have to be standardized before you can compare card placement across participants (Lamantia, 2003)Lamantia, Analyzing Card Sort Results with a Spreadsheet Template, Boxes and Arrows, 2003. Do not treat that bucket name as final navigation copy yet. [diagram pending: raw labels merging into one working analysis bucket]
  4. Look for cards that repeatedly travel together. This is the plain-language version of a similarity matrix, the chart most card-sorting tools generate for you. In a spreadsheet, list the cards down the left side and again across the top. Each cell answers one question: how many participants put these two cards in the same group? If 18 out of 24 participants grouped "download tax form" with "view pay history," that relationship has a 75% similarity score. Then build a second view: for each card, count how often it landed in each standardized category. The first view tells you which cards seem connected. The second tells you whether each card has a clear home or is scattered across categories.
  5. Read the strong, weak, and split signals. A strong connection means most participants kept two cards together. A weak connection means your team expected two cards to go together, but users kept them apart. Split cards are the most interesting: "tax filings" might land with payroll half the time and compliance half the time. That split is not noise. It may mean the item needs a cross-link, a clearer label, a different card description, or a category that does not yet exist.
  6. Check segments before averaging people together. If administrators and customers sort the same deck differently, the average can hide both models. The ONRR case study on Digital.gov analyzed results by user type after normalizing categories, separating industry participants from internal users to understand audience differences (Digital.gov, ONRR case study)Office of Natural Resources Revenue, Open-source Information Architecture Design, Digital.gov, 2022. Segment only on variables that matter to the product decision, such as role, experience level, buyer type, or domain expertise.
  7. Use the dendrogram as a sketch, not a judge. A dendrogram is a tree-shaped chart that suggests possible groups from the similarity matrix. It starts with the cards that participants most often put together, then adds weaker relationships as the branches get larger. Read it as a draft grouping, not as the answer. If one branch looks too broad, split that branch and leave the others alone. If a card joins a group late and awkwardly, go back to the spreadsheet and see whether participants were actually divided about it. Righi and colleagues recommend moving back and forth between the dendrogram and the item-by-item matrix instead of trusting a single cut point. Report where you accepted the tree, where you overrode it, and why those choices served the product decision. [diagram pending: similarity matrix turning into a dendrogram]
  8. Name categories after you understand the group. Participant labels are evidence, not final copy. Look for repeated words, variants, synonyms, and phrases that point to the same idea. Then check the label against the cards inside the group, participant comments, site-search language, editorial rules, and any legal or brand constraints. A label can be popular and still too vague. A label can be internally approved and still fail users.
  9. Turn analysis into recommendations. For each proposed category, write the evidence in plain language: the core cards, the strongest card relationships, the common participant labels, the cards with weak agreement, and what you recommend doing next. Some cards will belong in the category. Some need a cross-link. Some need a better name. Some should be removed from scope. The deliverable is not the chart; it is the decision record behind the draft structure.

You do not need a paid tool to do this work. Use the tool you already have, as long as it lets you preserve the participant-level data.

Analyzing a card sort without a paid tool
No-paid-tool option Best for What to make
Photos plus notes Small moderated or physical sorts One session record per participant, then a spreadsheet of card IDs and group labels
Spreadsheet Most small to medium studies Raw-data table, standardized-label map, item-by-category table, and a simple similarity matrix
Pivot tables Closed sorts and open sorts with standardized labels Counts and percentages showing where each card landed
Open-source script Larger open sorts or teams comfortable with CSV files Distance matrix, dendrogram, and cluster-label extracts; the cardsort Python package, for example, can create distance matrices, dendrograms, and label extracts from a five-column CSV (cardsort docs)cardsort Python package documentation, cardsort.readthedocs.io

Reliability deserves a precise mention here. Open card sorting can be more reproducible than skeptics assume: Katsanos and colleagues found high similarity across repeated open-sort studies with the same content (Katsanos et al., 2019)Katsanos, Tselios, Avouris, Demetriadis, Stamelos & Angelis, Cross-study Reliability of the Open Card Sorting Method, CHI 2019 Extended Abstracts. That helps the method. It does not make the analysis automatic. The groupings can repeat while your data cleaning, label standardization, segment choices, clustering view, cut point, and final label interpretation still need to be defended.

So what? Analyze card sorts like evidence, not output. Preserve the raw sort, standardize carefully, count the patterns, read the exceptions, and write down the reasoning that turns the piles into a structure.

A card sort produces a hypothesis. Only a tree test tests it.

This is the limitation most card-sorting guides soften. A card sort puts all the items in front of someone and asks them to organize. A live product does the opposite. The person has a task, sees one level of navigation at a time, and is trying to choose the next plausible click before patience runs out.

Grouping content and finding content are different jobs. A structure can make sense when all the cards are visible and still fail when someone has to navigate it one label at a time.

NN/g reaches the same conclusion from the other direction. Their guidance lists "lacking context" as a limitation, notes that card sorting does not account for the site's images, structure, and links, and says it reveals only limited navigation pathways. They recommend tree testing [link: /tree-testing], which they describe as reverse card sorting, over closed card sorting for validation work (Nielsen Norman Group, 2024)Nielsen Norman Group, Card Sorting: Uncover Users' Mental Models, 2024.

Treat the workflow as two steps. First, use the card sort to draft a structure from the way users group content. Then use a tree test to see whether people can find real things in that structure. Until the tree test works, you have a promising draft, not a proven IA.

Skip the card sort when the method does not match the decision. If the content set is small and obvious, the sort will make the obvious expensive. If the structure already exists and you are allowed to change it, tree testing will show where people get lost. If the problem is what people do inside the interface around that structure, usability testing is the better instrument. If the question is a single label, use a label test or first-click study. If most items legitimately belong in several places, a card sort will force a false neatness.

So what? Budget for the tree test when you plan the card sort. A structure you cannot test will still be treated like a decision once it enters a roadmap.

Tree testing vs card sorting

Teams often compare the two methods as alternatives. For IA work, they usually belong in sequence.

Tree testing vs card sorting
Card sorting Tree testing
Question How do people group and label this content? Can people find things in this structure?
Stage Generative, before the structure exists Evaluative, once a structure is proposed
What participants see All items at once, no goal One level at a time, given a task
Primary output Similarity data, candidate categories, vocabulary Task success rate, directness, where people go wrong
Answers findability? No Yes
Typical sequence First Second, on the structure the sort produced

Run the sort to learn which content feels related. Draft a structure from it. Tree test [link: /tree-testing] the draft to see whether people can use it for real tasks. If you only have budget for one study and a structure already exists, tree testing is usually the better spend because it measures the outcome you care about.

The two methods are strongest as a loop. The sort gives you the draft. The tree test shows where you misread it.

Frequently asked questions

What is card sorting in UI/UX? Card sorting is a UX research method where participants group labeled content items into categories that make sense to them. Researchers use it when a menu, taxonomy, or help center needs to match how users already group the work. It helps you see which items belong together, which labels users reach for, and where your current structure feels internal.

What are the three types of card sorting? Open, closed, and hybrid. In an open sort, participants create and name their own categories. In a closed sort, they place cards into categories you provide. In a hybrid sort, they start with your categories and can add their own. Use open when the structure is undecided, closed when you need to check an existing structure, and hybrid when you mostly believe the structure but want to find its gaps.

What is the difference between card sorting and tree testing? Card sorting asks people to group content. Tree testing asks people to find content in a proposed structure. That means card sorting helps you draft the structure, while tree testing checks whether the structure works for real tasks. Card sorting cannot measure findability by itself.

How many participants do you need for a card sort? Around 20 to 30 is a good target for the quantitative structure. Tullis and Wood found that the similarity matrices from 20 to 30 participants correlated about .93 to .95 with a 168-participant baseline, and returns diminished after that. Our stance is to stop around 30 and spend the remaining budget on a moderated subset, because the explanation behind the piles is what makes the result usable.

Is card sorting qualitative or quantitative? Both. The grouping data is quantitative: it gives you similarity matrices, agreement scores, and dendrograms. The explanation is qualitative: it tells you why the cards felt related. If you collect only the numbers, you get a proposed structure with no explanation attached.

What are the limitations of card sorting? Card sorting cannot tell you whether anyone will find something in the product. Participants see all the cards at once, with no task and no interface. Real navigation happens one level at a time, under a specific goal. Card sorting can also force items into one category when they belong in several, and its dendrogram depends on the clustering view and cut point you choose.

Leo Hoar, PhD
About the author

Leo Hoar, PhD is the founder of the UXR Institute, where experienced researchers sharpen the methodological and strategic skills that turn findings into decisions. He led UX research at Beam Benefits and Samsung Research America, and writes and teaches on qualitative analysis and the interpretive work that turns data into a finding. Read more at his bio page.

Created with