Quantitative UX Research: A Practitioner's Guide to Methods, Concepts, Tests, and Tools
A practitioner's guide to the methods, statistics, and tests of quantitative UX research, written for researchers who have to turn a result into a product decision.


Quantitative UX research is the systematic study of user behavior and attitudes, expressed as numbers rather than narratives. Researchers measure things like behavioral metrics from analytics, benchmark scores, and standardized survey scales, then use statistics to determine whether the findings from a sample are likely true of the wider user population. Quantitative UX research methods include surveys, A/B tests, usability benchmarking, and analytics reviews, used to describe how big a problem is, compare design options, or test whether a change actually caused an effect. Teams rely on it to size issues, track the user experience over time, and back decisions with evidence.
Key takeaways
- Quantitative UX research measures behavior and attitudes at scale and tests whether patterns are real, so you can generalize from a sample to your user base.
- It answers how many, how much, how often, and, with an experiment, whether a change caused an effect. It does not explain why that effect happened. Its natural partner is qualitative research.
- The methods sort along two axes at once: research purpose, meaning generative, evaluative, or summative, and data source, meaning behavioral versus attitudinal.
- Statistical significance tells you a result would be surprising if nothing were going on. Practical significance tells you whether it is big enough to act on. You need both.
- A result matters only when it changes a decision. The discipline that separates strong quant from busywork is deciding the rule before you see the data.
Most quantitative training stops at the p-value. You learn to run the test, read the output, and report that a result is significant. The harder work starts after that, when a product team has to decide what to do about the number, and it is the work most guides leave out. This one is organized around it. It covers the methods, the statistics, and the tests you will actually use, and treats all of that as the setup for the question a stakeholder is actually asking, which is what we should do now. That is where quantitative research earns its keep, and everything below is pointed at it.
The term is less settled than it looks. There is no field-endorsed definition of quantitative user research, the line between quantitative and qualitative work is drawn by convention rather than by any fixed property of a method, and a measure can look rigorous while barely capturing the thing it claims to. None of this is a reason to distrust the numbers. It is a reason to know where they bend.
This guide covers one half of the discipline. For the full map of methods, including the qualitative work quant depends on, start with our guide to UX research.
What quantitative UX research actually means
"Quantitative" is used in three different senses in UX research: that the data are numbers, that the study was mediated by an instrument such as a survey or an analytics tool, and that the study was designed so that statistical inference from a sample to a population is warranted. The third is the one worth adopting, and it is the standard this guide runs on.
The term "quantitative UX research" looks clear and stable. There is even a job title to match it, the quantitative UX researcher. Look closely, though, and the word "quantitative" turns out to be working in several different registers at once, and two researchers can both claim to do it while meaning genuinely different things. The label itself is part of the problem. What defines the tradition is a package of assumptions borrowed from the natural sciences, not the mere presence of numbers (Bryman, 1984)Bryman, The Debate about Quantitative and Qualitative Research: A Question of Method or Epistemology?, British Journal of Sociology, 1984. Once you see that, the single word splits into at least three distinct senses, and they are not the same claim.
- The data-type sense. Quantitative means the data are numbers. This is the folk definition and the emptiest of the three, for two reasons. First, not every number is a quantity. Stevens' levels of measurement separate numbers that only label or rank from numbers that carry magnitude, so a satisfaction rating of 4 is an ordinal position rather than a measured amount, and averaging it already assumes the distance from 3 to 4 equals the distance from 4 to 5 (Stevens, 1946)Stevens, On the Theory of Scales of Measurement, Science, 1946. Second, if this were true, you could manufacture numbers from almost anything. Counting how often a theme recurs across twelve interview transcripts, what mixed-methods researchers call quantitizing, produces numbers without producing quantitative research (Sandelowski, Voils & Knafl, 2009)Sandelowski, Voils & Knafl, On Quantitizing, Journal of Mixed Methods Research, 2009. Beneath both sits the problem Joel Michell named. A genuine quantitative science has two separable jobs, the scientific one of showing that the attribute really has additive structure and the instrumental one of building procedures to estimate its magnitudes, and the tradition from Fechner onward simply skipped the first (Michell, 1997)Michell, Quantitative Science and the Definition of Measurement in Psychology, British Journal of Psychology, 1997. UX research inherited that omission wholesale. A SUS score of 72 is read as standing in a real magnitude relation to a 68, yet that additive structure is an empirical hypothesis, almost never tested, quietly converted into a premise the moment the number is written down. The data-type definition does not merely tolerate that move. It is that move, promoted to a definition.
- The instrumentation sense. This is the sense that treats quantitative work as instrument-mediated work: surveys, analytics tools, logging systems, dashboards, and other apparatuses that stand between the researcher and the user. It is seductive because instruments look rigorous, but mediation is the wrong property. An open-ended survey is an instrument, yet it produces qualitative text. Watching users and tallying task successes with a stopwatch is direct and unmediated, yet the output is quantitative. The real question is not whether a tool sat between you and the user. It is whether the instrument produced valid, reliable measures of the construct you claim to be studying, which is exactly the problem Hornbæk raises in his review of 180 usability studies (Hornbæk, 2006)Hornbæk, Current Practice in Measuring Usability: Challenges to Usability Studies and Research, 2006.
- The inferential sense. Quantitative means the study is designed so that statistical inference from a sample to a population is warranted. This is the sense operative in Sauro and Lewis, whose whole project is applying inferential statistics to research problems, confidence intervals, sample-size estimation, standardized questionnaires, and the long-running controversies in measurement (Sauro & Lewis, 2016)Sauro & Lewis, Quantifying the User Experience, 2016. On this reading, a study with numbers but no defensible sampling logic is not quantitative research at all. It is arithmetic decorating a qualitative study.
The inferential sense is the one worth adopting. It ties the word "quantitative" to what quantitative work is for, a warranted move from a sample to the population you care about, and it is the standard the rest of this guide runs on.
Adopting it changes what you scrutinize. The question stops being "do I have numbers?" and becomes "does my sample license a claim about the population I actually care about?" That has consequences a working researcher feels on a Tuesday.
- Your sampling is what earns the word "quantitative," and recruiting is where you win or lose it. A dashboard of metrics from a self-selected pop-up survey is numbers without warranted inference, so it is not "quantitative" in the rigorous sense. If sampling hasn't taken place, any claims from that data cannot be inferred of the wider population.
- You have to name the population you are inferring to. "Our users" is not a target until you say which ones, new or returning, one market or all of them, and confirm your sample stands for them. The inference is only ever as good as that match.
- It tells you when not to quantify. If you cannot get a sample that warrants the inference, adding statistics is theater. Honest qualitative work beats arithmetic dressed as generalization, and knowing which one a situation calls for is part of the craft.
One more caution matters in practice. A score is not automatically a real measurement just because it is a number. A SUS score, a satisfaction average, or a rating-scale mean can be useful, but it is still an estimate built from a particular instrument and sample. If one design scores two points higher than another, do not treat that difference as a shipping argument until you know what the score represents, how precise the estimate is, and whether the difference is large enough to matter.
So what? Simply having numbers for data does not, on its own, constitute quantitative research. A study is quantitative only when its sample lets you infer a claim about the population you care about. This assumes that your research has defined the population of interest, and is sampling from it.
What counts as quantitative user research, and where the definition frays
Quantitative user research is defined by three moves: operationalizing a construct into something you can record, measuring it behaviorally or attitudinally, and holding a sampling logic that warrants a claim beyond the users you observed. A study is quantitative to the degree it makes all three. The full working definition:
Working definition. Quantitative user research is the systematic study of users, their behaviors, and their evaluative responses to designed systems, in which the phenomena of interest are operationalized as measurable variables, observed or elicited from samples or populations of users, and analyzed numerically in order to describe magnitudes, estimate parameters, test hypotheses, or support causal inference about design decisions.
The components that make a study quantitative
A study is quantitative user research to the degree it makes these moves. Reading them as a checklist is more useful than arguing over a label.
- Operationalization. A construct such as findability, perceived ease, trust, or effort is mapped to an indicator you can record. This is where most validity is won or lost, and it is the least examined step.
- Measurement. Either behavioral, such as task completion, time on task, error counts, and log-derived events, or attitudinal through a psychometric instrument such as SUS, UMUX-Lite, SUPR-Q, the UEQ, or PSSUQ. A survey of nearly 100 summative usability tests found practitioners typically collect some mix of completion rates, errors, task times, task-level and test-level satisfaction, help access, and problem lists (Sauro & Lewis, 2009)Sauro & Lewis, 2009, reported in Quantifying the User Experience, 2016.
- Sampling and inference. A defensible warrant for generalizing beyond the units you observed, so the numbers estimate something about the population rather than describing only the sample.
Four places the definition strains
Calling research "quantitative" sounds precise and solid. It usually isn't. Four things you would expect to be nailed down are shakier than the word implies. Where quantitative work even begins is a judgment call. The numbers often go unchecked for whether they measure anything real. What you are measuring may not be one measurable thing, and the biggest sources of quant data, analytics and A/B tests, do not clearly fit the picture at all. Treat "it's quantitative" as a claim worth checking, not a badge of rigor.
Where "quantitative" begins is a judgment call. The line between qualitative and quantitative testing is usually drawn at a participant count, and that count is a judgment about precision rather than a property of the method. The familiar five-user rule governs qualitative, formative testing. Quantitative sample size follows from the estimate you need, the variability in the measure, and how wide a confidence interval you are willing to live with, which is a decision about the cost of being wrong (Sauro & Lewis, 2016)Sauro & Lewis, Quantifying the User Experience, 2016. There is no fixed count where research turns quantitative, and the line moves with the stakes.
Most numbers are never checked for validity. A score is only as good as the link between the thing you care about and the thing you actually recorded, and that link is rarely examined. A systematic review of measurement practice across four years of CHI papers found that studies gave a complete rationale for scale selection only 20% of the time, provided every scale item about 30% of the time, adapted more than a third of the scales they used, and reported any check on scale quality only about a third of the time (Perrig et al., 2024)Perrig et al., Measurement Practices in User Experience Research, 2024. Flake and Fried call these questionable measurement practices, decisions that raise doubts about whether a measure is valid at all and let a number look rigorous while measuring very little (Flake & Fried, 2020)Flake & Fried, Measurement Schmeasurement, 2020.
What you are measuring may not be one thing. Validate every measure perfectly and you can still report a meaningless number, because the label may cover several things that do not move together. Usability is the stock example. In a study of 87 participants across 20 information-retrieval tasks, Frøkjær, Hertzum, and Hornbæk found the correlation between efficiency, measured as task time, and effectiveness, measured as solution quality, was negligible (Frøkjær, Hertzum & Hornbæk, 2000)Frøkjær, Hertzum & Hornbæk, Measuring Usability: Are Effectiveness, Efficiency, and Satisfaction Really Correlated?, 2000. Hornbæk's later review of usability measurement across 180 studies raised the same worry as an open problem, alongside the need to study long-term use and to validate the growing pile of subjective questionnaires (Hornbæk, 2006)Hornbæk, Current Practice in Measuring Usability, 2006. When effectiveness, efficiency, and satisfaction do not move together, a single "usability score" averages away the disagreement and stands for nothing in particular.
At scale, the sampling logic falls away. Quantitative research is supposed to generalize from a sample to a population. Log data has no sample in that sense. It is found data, with no sampling frame, no recruitment, no informed participation, and no control over what was instrumented. Rodden, Hutchinson, and Fu introduced the HEART framework and a goals-signals-metrics process at CHI 2010 precisely because large-scale behavioral measurement had no established method the way small-scale observation did (Rodden, Hutchinson & Fu, 2010)Rodden, Hutchinson & Fu, Measuring the User Experience on a Large Scale, 2010. The experimental branch goes further, importing causal-inference machinery from a lineage largely outside HCI (Kohavi, Tang & Xu, 2020)Kohavi, Tang & Xu, 2020. Whether A/B testing is quantitative user research or a neighboring craft that researchers sometimes practice is genuinely unsettled.
What quantitative UX research can answer, and what it can't
Quantitative research is the right tool for three kinds of questions:
Estimation with known precision. Quantitative research can estimate how many users experience a problem and how uncertain that estimate is. A completion rate from eight participants and one from eighty may look identical as percentages, but they do not support the same decision. The confidence interval tells you whether the estimate is tight enough to use, which is why the interval, not the point estimate, is often the finding (Sauro & Lewis, 2016)Sauro & Lewis, Quantifying the User Experience, 2016.
Comparison on evidence. Quantitative research can compare a design against a benchmark, competitor, previous release, or decision threshold. The useful answer is not just "Design B was better." It is "Design B was better by this much, with this much uncertainty." That is the form of evidence you need when a team has to decide whether a difference is real enough, large enough, and important enough to act on (Sauro & Lewis, 2016)Sauro & Lewis, Quantifying the User Experience, 2016.
Causal adjudication, of a specific kind. A controlled experiment can tell you whether a change caused a measurable effect. It answers the question "did this change cause that outcome?" It does not explain the mechanism, what users experienced, or what else might have worked better (Kohavi, Tang & Xu, 2020)Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments, 2020. Quantitative research tests hypotheses you already have. Qualitative research is usually better at finding the hypothesis you did not know to test.
Quantitative research also has clear limits.
It is weak when the problem is still vague. Quantitative research needs a defined population, a measurable question, and enough observations to estimate uncertainty. Early in discovery, those conditions often do not exist. You can still compute a percentage from a handful of users, but the interval around it will usually be too wide to guide a decision. Be especially skeptical of surprising numbers. Twyman's law captures the habit: the more surprising the result, the more likely it is an error (Kohavi, Tang & Xu, 2020)Kohavi, Tang & Xu, 2020.
Self-report measures what people say, not what they will do. Survey data is real evidence about beliefs, preferences, satisfaction, and intent. It is not a direct forecast of behavior. A meta-analysis of experimental studies found that even medium-to-large changes in intention produce only small-to-medium changes in behavior (Webb & Sheeran, 2006)Webb & Sheeran, Does Changing Behavioral Intentions Engender Behavior Change?, Psychological Bulletin, 2006. Use self-report to understand attitudes. Use behavioral data or experiments to evaluate behavior.
A survey number is only as good as the people behind it. A survey mean describes the people who answered. It represents the wider user population only if the sample can stand in for that population. That means naming the population, checking who responded, weighting the sample when known groups are over- or underrepresented, and screening out fake or low-quality responses. These are not statistical niceties. They decide whether the number means anything. The mechanics live in our guide to UX surveys [link: /ux-surveys].
Tracked numbers become targets. Teams often start treating a metric as the goal itself, a failure mode called surrogation (Harris & Tayler, 2019)Harris & Tayler, Don't Let Metrics Undermine Your Business, Harvard Business Review, 2019. Once that happens, people optimize the metric instead of the user or business outcome it was supposed to represent. Every metric is a proxy. Part of the quant researcher's job is to notice when the proxy has expired.
Access can decide the project before analysis begins. Behavioral data often lives in product analytics or a warehouse owned by data or engineering. Survey data lives in the tool that fielded it. Support data lives in a CRM owned by the support team. Before you choose a test, map who owns the data, what was instrumented, and who can pull it. That step often decides whether the project is feasible.
From the Classroom A researcher in our Statistical Methods course put the precondition bluntly early on: garbage in, garbage out. Most of quantitative work is not the test. It is the cleaning, labeling, and structuring of the data before the test, and a beautiful analysis on messy inputs is still a wrong answer.
Quantitative and qualitative: two kinds of knowledge Quantitative and qualitative research are not rival grades of rigor. They generalize differently, so they answer different questions: quantitative tells you what is true across your users, qualitative what is true about their experience. The strongest programs sequence the two rather than choosing a side. Full treatment: Quantitative vs Qualitative Research
[link: /qualitative-vs-quantitative].
Quantitative user research methods, by purpose and data source
The main types of quantitative user research methods:
- A/B testing and online experiments test whether a change caused an effect.
- Quantitative usability studies estimate task success, time on task, errors, and task ease with enough precision to compare designs or benchmarks.
- Survey studies and standardized instruments capture attitudes, preferences, satisfaction, and perceived usability.
- Tree testing and first-click testing check findability and where people start a task.
- Card sorting shows how users expect information to be grouped.
- Product analytics studies describe instrumented live behavior through funnels, cohorts, retention, and paths.
- MaxDiff and conjoint studies surface priorities and trade-offs among options.
Each belongs somewhere on two axes, purpose and data source. The point of sorting them this way is to classify each method by its practical function in product development, so the grid tells you which method fits the decision in front of you and which claim that method can defend.
Use the matrix by asking two questions. First, what job does the study need to do? If you are trying to find problems or needs, you are doing generative work. If you are testing a design in progress, you are doing evaluative work. If you are judging a finished or live product against a target, a competitor, or an earlier release, you are doing summative work. Those labels come from the older distinction between formative and summative evaluation (Scriven, 1967)Scriven, The Methodology of Evaluation, 1967.
Second, where will the evidence come from? Behavioral methods measure what users do, in logs, tasks, clicks, and completion rates. Attitudinal methods measure what users say, in ratings, rankings, and self-report. That difference matters because doing and saying do not reliably collapse into the same signal. Usability research has found the same problem inside its own core measures: effectiveness and efficiency are performance-oriented, while satisfaction is self-reported, and those measures do not always move together (ISO, 2018; Frøkjær, Hertzum & Hornbæk, 2000; Hornbæk, 2006)ISO, ISO 9241-11:2018; Frøkjær, Hertzum & Hornbæk, Measuring Usability: Are Effectiveness, Efficiency, and Satisfaction Really Correlated?, 2000; Hornbæk, 2006.
The same caution applies outside usability testing. Changes in what people intend to do produce smaller and less certain changes in what they actually do, so self-report is evidence, not a behavior forecast (Webb & Sheeran, 2006)Webb & Sheeran, Does Changing Behavioral Intentions Engender Behavior Change?, 2006.
Organizing methods this way prevents the usual category mistake. A method name alone does not tell you what claim the evidence can support. A/B testing belongs where live behavioral data can support a causal claim. MaxDiff and conjoint belong where attitudinal choice data can support priority or trade-off claims. Product analytics belongs where instrumented behavior can describe patterns across a live product. The point is not to label methods neatly. The point is to match the decision to evidence that can actually defend it.
| Behavioral: what users do | Attitudinal: what users say | |
|---|---|---|
| Generative: discover problems and needs | Exploratory analytics, log mining, open card sorting | Discovery surveys, scaled open-ended surveys, problem-ranking studies |
| Evaluative: test a specific design | A/B testing, quantitative usability testing, tree testing, first-click testing, closed card sorting | Concept surveys, preference surveys, MaxDiff studies, conjoint studies |
| Summative: benchmark overall performance | Usability benchmarking, product analytics studies using funnel, path, cohort, retention, or drop-off analysis | Standardized survey instruments such as SUS, UMUX-Lite, SUPR-Q, PSSUQ, SEQ, NPS, CSAT, and CES |
The grid is a map, not a filing cabinet. Purpose is intent, not a fixed property of a technique. A survey is generative when you are exploring, evaluative when you are testing a preference, and summative when you are benchmarking satisfaction, which is why it appears in more than one place. Analytics can be generative when you are looking for unexplained behavioral patterns and summative when you are tracking a product's ongoing performance. Place a method by the job you are giving it, not by its usual label.
Four terms are easy to blur in quantitative UX work. A method is the study design that generates evidence. A metric or instrument is what you measure, such as task success, time on task, SUS, SEQ, CSAT, or NPS. An analysis is how the data are summarized or tested, such as a funnel, cohort, correlation, regression, ANOVA, or chi-square. A tool is where the work happens. The table below keeps the study design separate from the evidence it produces and the claim that evidence can support.
| Method | Typical evidence or metric | Claim it can support | Common mistakes |
|---|---|---|---|
| Survey study | Rating scales, rankings, self-reported behaviors, coded open-ended responses | Among the population sampled, this many users report, prefer, believe, or experience something | Treating a self-selected sample as representative, or reading stated intent as future behavior |
| Standardized usability survey | SUS, UMUX-Lite, SUPR-Q, PSSUQ | This experience scores higher, lower, or about the same as a baseline on a validated usability instrument | Treating the score as a complete measure of usability, or over-reading tiny differences |
| Relationship or satisfaction survey | NPS, CSAT, CES | Users report this level of satisfaction, advocacy, or effort at a defined moment in the journey | Presenting relationship or satisfaction scores as usability measures |
| MaxDiff study | Best-worst choices across balanced item sets | This list has a relative priority order under forced trade-off conditions | Using it to estimate feature demand or price trade-offs it was not designed to model |
| Conjoint study | Choices among profiles made from attributes and levels | Users appear to value these attributes this much relative to one another within the modeled choice scenario | Treating modeled preference as guaranteed market behavior |
| Quantitative usability study | Task success, time on task, errors, SEQ | Under these tasks and conditions, users succeed, fail, struggle, or rate task ease at measurable rates | Reporting percentages from tiny samples as if they generalize cleanly, or generalizing artificial tasks to all product behavior |
| A/B test, or online controlled experiment | Randomized exposure, outcome metrics, guardrail metrics | This change caused, or failed to cause, movement in the chosen metric for eligible live traffic | Shipping on a statistically significant but practically meaningless lift, or trusting a polluted experiment |
| Tree testing | Task success, directness, path, time | Users can or cannot locate target content through this information architecture at this rate | Mistaking findability in a tree for the whole navigation experience |
| Card sorting, open or closed | Item-category assignments, similarity, category agreement | Users tend to group these items together, or they can fit items into these categories | Treating an open sort as a finished navigation design |
| First-click testing | First-click location, first-click success, time to first click | Users start in the right or wrong place at this rate | Treating the first click as proof that the whole task will succeed |
| Eye tracking | Fixations, dwell time, scan order, areas of interest | Users looked at these areas, in this order, for this long | Treating gaze as understanding, preference, or intent |
| Product analytics study | Event logs, sessions, users, paths, conversions, retention | Instrumented users behaved this way in the observed system | Assuming uninstrumented behavior does not exist, or treating found data as a clean sample |
| Benchmarking study | Matched task metrics, survey scores, competitor or prior-release measures | This product performs better, worse, or about the same as the chosen baseline | Comparing numbers that were collected under different conditions and calling the difference a benchmark |
Notice what is not in the table: a fixed sample-size column. Sample size is not a property of the method alone. It depends on the precision you need, the smallest effect that would change the decision, the variability of the measure, the confidence level, and, for live experiments, the traffic that can actually be exposed. Numbers like 40 participants for a usability benchmark or a few hundred respondents for MaxDiff and conjoint are planning anchors, not permissions to skip the power or margin-of-error question.
Go deeper on the individual methods:
- Card sorting
- Tree testing
[link: /tree-testing] - Quantitative usability testing
[link: /usability-testing] - Surveys and standardized metrics
[link: /ux-surveys] - A/B testing
[link: /ab-testing] - UX analytics, funnels and cohorts
[link: /ux-analytics]
Statistical thinking
Statistical thinking is the habit of treating every result as a glimpse of a larger population rather than a fact about it. It asks two questions of any number: where did this sample come from, and would a different sample have shown the same thing? Everything else in statistics is machinery for answering the second.
You're standing at a window looking out at a street. You want to know what the neighborhood is like, who lives there, how they move around, what they're up to. But you only ever get to look for two minutes at a time, and only through this one window.
Whatever you see in those two minutes is real. It happened. But it isn't the neighborhood. It's two minutes of the neighborhood, framed by one window. The discipline is in never confusing those two things. You are always describing a glimpse and reaching for a place.
Once you hold that, two questions become automatic.
Where's the window? If your window faces the school, you'll see children. Not because the neighborhood is mostly children, but because that's where you're pointed. In our work the window is recruiting, screening, who shows up, who drops out, and what you say to people before the session starts. A perfectly executed study aimed at the wrong window gives you a crisp, confident, useless answer.
Would two other minutes have looked different? This is the one that separates statistical thinking from ordinary observation. You saw three people walk by carrying umbrellas. Is it raining, or did you happen to catch three people who own umbrellas? The test is imagining the same window at a different two minutes. If you'd almost certainly see umbrellas again, it's raining. If it easily could have gone the other way, you don't have a finding yet. You have a glimpse you got attached to.
That's the whole discipline. Before you commit to a pattern, ask whether a different two minutes would have shown you the same thing.
Two consequences of this way of thinking:
Short looks aren't worthless. They just answer a different question. Two minutes is enough to establish that something is possible. If you see one person trip on the loose paving stone, the paving stone is a hazard, full stop, and a second look can't take that away from you. What two minutes can't tell you is how often people trip on it. The mistake isn't the short look. It's letting "this can happen" get written up as "this happens 62% of the time."
Longer looks help. Watch for two hours and the flukes even out. The three umbrella-carriers stop dominating the picture. That's all a bigger sample buys you: fewer accidents of timing masquerading as facts about the street.
Essential statistical concepts for quantitative UX research
A handful of concepts carry most of the statistical work in UX research: the two data types, the null hypothesis, statistical significance, effect size, confidence intervals, Type I and Type II error, the parametric and non-parametric split, sampling distributions, and degrees of freedom. Each is defined below alongside the mistake it invites, and the p-value gets its own treatment underneath the table. The stats guide takes each further, with worked examples and the tests that use it, in our guide to statistics for user research [internal link: anchor "statistics for user research" → /statistics-for-ux-research].
| Term | Working definition | UX watch-out |
|---|---|---|
| Categorical data | Values that are labels. Nominal categories have no natural order, like chat, email, or phone. Ordinal categories are ranked but unevenly spaced, like low, medium, and high severity. | Count categories and compare proportions. Be cautious about averaging ordinal ratings, because the arithmetic assumes spacing the scale never promised. |
| Continuous data | Values on a scale where the distances are real and consistent. Twelve minutes is six more than six minutes and half of twenty-four. | Means and standard deviations are meaningful only when the scale behaves like a real quantity. Task times are continuous, but usually skewed. |
| Null hypothesis | The default starting position that the difference in front of you came from sampling variation rather than a true difference in the population. | Evidence can push you away from the null, but a non-significant result does not prove "there is no difference." Equivalence needs its own threshold or test. |
| Statistical significance | The label applied when the p-value falls below a threshold agreed on in advance, conventionally .05. | A cutoff chosen after seeing the data is not a test. It is a description dressed up as one. |
| Effect size | The size of the observed difference, often standardized as Cohen's d so it can be compared across scales. | Significance tells you how incompatible the data are with the null hypothesis. Effect size tells you how big the observed difference is. You need both. |
| Confidence interval | The range of values your data is consistent with, such as "about 37%, plausibly 29% to 45%." | The width is part of the finding. A wide interval says the study cannot yet distinguish between decisions. |
| Type I and Type II error | A Type I error is a false positive. A Type II error is a false negative, often caused by a sample too small to detect the effect. | Type II error is common and quiet. Nothing in the output announces that you missed a real effect. |
| Parametric and non-parametric | Parametric tests assume a particular form for the data or model. Non-parametric tests assume less, often by working from ranks. | Non-parametric tests help with ordinal scales, skew, and outliers, but they do not rescue an underpowered study. |
| Sampling distribution | The spread of results you would get if you repeated the same study again and again with different samples. | It is the bridge from one sample to uncertainty about the population. Small, skewed samples may not behave like the textbook version. |
| Degrees of freedom | The number of independent pieces of information available for an estimate. | Tests use it to calibrate how much natural wobble to expect. Fewer degrees of freedom usually means wider tails and a higher bar. |
From the Classroom Our Statistical Methods course teaches degrees of freedom with a Sudoku analogy. When the values in a set have to add to a known total, you can fill in numbers freely until the last one, which is then fixed by everything before it. Degrees of freedom is just a count of how many values were free to vary. The intimidating term turns out to describe something ordinary.
p-value
The one term above worth slowing down on, because it is quoted most and understood least.
How often pure luck would hand you a result this extreme.
Say you and I are flipping a coin for money and you're losing. Fifteen flips, I've won twelve. You start to wonder whether the coin is rigged. You can't cut it open. All you can do is ask: if this coin were perfectly fair, how weird is 12 out of 15?
Turns out a fair coin gives you 12 or more heads about 1.8% of the time. Rare. Not impossible. But rare enough that you push the coin across the table and ask to use a different one.
That 1.8% is the p-value. It answers a conditional question: if the coin were fair, how often would you see a result at least this lopsided? It does not tell you the probability that the coin is fair, and it does not tell you the probability that the coin is rigged. It tells you how well the fair-coin explanation fits the result you saw. A small p-value means a fair coin would rarely produce a result this extreme, so the fair-coin explanation is getting hard to defend. The American Statistical Association (ASA) is blunt about the two misreadings to avoid here: a p-value is not the probability that the hypothesis is true, and not the probability that chance alone produced the result (Wasserstein & Lazar, 2016)Wasserstein & Lazar, The ASA Statement on p-Values: Context, Process, and Purpose, 2016.
Every p-value uses the same logic. Say twelve of fifteen participants preferred design B. The null hypothesis is that users have no real preference, so each participant is equally likely to prefer A or B. Under that assumption, a 12-to-3 split for either design would happen about 3.5% of the time, or p = .035. That does not prove users prefer B. It says a split this lopsided would be rare if there were no real preference.
The "either design" part matters. In the coin example, you cared only whether the coin was tilted against you, so the test looked in one direction. In a design preference study, you usually care about a lopsided result in either direction: B beating A, or A beating B. That is a two-sided test. You have to choose one-sided or two-sided before the data arrives. If you wait, see that B won, and then choose the one-sided test because it gives you a smaller p-value, you are no longer testing a prediction. You are making the result look stronger after the fact.
The core statistical tests, and how to choose between them
Software runs the test. What it cannot do is choose the right test for you, or warn you when the answer should not be trusted. That choice is not a matter of taste or habit. Two things settle it: the shape of your question, and the type of data on each side of it. Get those straight and the test is nearly determined. Reach for a familiar test first and bend the data to fit, and you get a clean number that answers a question you never asked.
Questions come in a few shapes. Are you comparing groups, checking whether two measures move together, or asking whether two labels are associated? Cross that with your data type and you land on one of a handful of workhorses.
| Your question | Data on each side | Default test |
|---|---|---|
| Do two independent groups differ on a continuous outcome? | one continuous outcome, two independent groups | Welch's t-test |
| Do the same people differ before and after? | one continuous outcome, measured twice on the same users | paired t-test |
| Do three or more independent groups differ? | one continuous outcome, three or more independent groups | one-way ANOVA, then adjusted follow-up comparisons |
| Do two measures move together? | two continuous or ordinal measures | Pearson correlation for a linear continuous relationship |
| Are two categorical variables associated? | categorical counts from independent observations | chi-square |
| Did a paired yes-or-no outcome change? | categorical outcome measured twice on the same users | McNemar test |
| What predicts an outcome once other variables are included? | continuous or binary outcome plus one or more predictors | linear regression for continuous outcomes; logistic regression for binary outcomes |
Every row has a fallback when its assumptions do not hold, and knowing the fallback is most of the skill. Those alternatives, what each test is actually doing, and the assumption checks in order are the substance of our guide to statistics for user research [internal link: anchor "statistics for user research" → /statistics-for-ux-research].
Two cautions carry more weight than the choice of test itself. The first is that a correlation is a lead, never a cause, because a third variable can drive both sides at once. That is not a footnote; it is the single most common way a quant readout misleads a room. The second is that the arithmetic runs whether or not the test's assumptions hold. Check the unit of analysis first, since repeated observations from the same user are not independent, then the shape of the outcome, the spread across groups, outliers that dominate the estimate, and whether the sample is large enough to detect the effect you came to find. A badly violated assumption returns a number that is precise and wrong, which is more dangerous than no number at all.
From the Classroom When a class example showed that time spent in an app correlated with money spent, a researcher in our course flagged the trap before anyone wrote it up as a finding: a third variable, like overall engagement, could be driving both. Correlation is a lead worth chasing, never a verdict to ship.
Statistical significance vs practical significance
Statistical significance asks whether a result would be surprising if nothing were going on, which is what the p-value reports. Practical significance asks whether the result is big enough to matter, which is what effect size, a confidence interval, and business judgment tell you. Treating the two as one is the most common way a quant finding misleads a team, and holding them apart is what separates a quantitative researcher from a person who runs tests.
The two are independent. With a large enough sample, a difference far too small to care about will cross the significance line, because the same ASA statement is explicit on a second point: a p-value does not measure the size or importance of an effect (Wasserstein & Lazar, 2016)Wasserstein & Lazar, 2016. Cohen made the same point earlier and more bluntly (Cohen, 1994)Cohen, The earth is round (p < .05), 1994.
The failure mode is familiar. A test comes back "significant, 23% lift," it goes on a slide, and the team ships, without anyone asking whether a 23% lift on that metric is worth the cost of the change. The fix is a matter of habit. Report effect size and a confidence interval next to every p-value, so a result is judged by the magnitude of the effect and not the p-value alone (Kirk, 1996)Kirk, Practical Significance: A Concept Whose Time Has Come, 1996. Better still, decide the smallest effect that would actually change your decision before you run the study rather than after you see the result. Jeff Sauro's framing is the one to keep in mind: a statistically significant result can end up being inconsequential, and the call is informed by math but driven mostly by professional judgment (Sauro)Sauro, From Statistical to Practical Significance, MeasuringU.
The opposite failure is just as common. Be most skeptical of the most exciting numbers. By Twyman's law, a result that looks surprisingly large is more often a logging or instrumentation error than a genuine effect. Before a striking figure goes on a slide, check the plumbing that produced it.
From the Classroom In a live ANOVA exercise, two designs came back statistically different on a usability score, but a researcher in our course noticed the effect size, Cohen's d, was small and said out loud that it might not be practically significant. That instinct, to check the magnitude the moment significance appears, is the whole habit in one sentence.
Hold this line all the way to the decision. A significant result that changes no decision is noise with a p-value. That is exactly the job of the decision layer.
Put these methods to work on your own product data.
Our 6-week live course, Quantitative UX Research: Statistical Methods for Product Development, walks you through t-tests, ANOVA, correlation, and chi-square applied to real product questions — no code, and built for researchers with no statistics background.
Explore the course New to statistics? Start with Statistics Prep for UX Researchers.Programming languages and tools for quantitative user research
Four kinds of tool cover almost all quantitative UX work: a spreadsheet, a menu-driven statistics application such as JASP or jamovi, a programming language such as R or Python, and SQL for getting hold of the data in the first place. Which one you need follows from two things, how often you will repeat the analysis and where the data lives.
The spreadsheet, first, honestly
Excel and Google Sheets run t-tests, correlation, chi-square, and descriptive statistics, and for an analysis you will run exactly once on a few hundred rows, they're the fastest path from data to answer. The spreadsheet's real limits are not statistical. They're that every step you take is hard to audit afterward: no record of what you clicked, no way to re-run last quarter's analysis on this quarter's data without redoing it by hand, and no easy answer when someone asks what exactly you did to get that number. A spreadsheet is a fine calculator. It is a weak record of analytical reasoning.
R: built by statisticians, for exactly this
R is a programming language designed for statistical analysis, and it shows in the best way: the things a researcher does daily are one line each. t.test(), wilcox.test(), chisq.test(), done. Three parts of the ecosystem matter most for UX work. The tidyverse handles the unglamorous 80% of every project, getting messy exported data into analyzable shape. ggplot2 produces the charts you'll actually put in a readout. And the community has spent decades writing free packages for nearly everything a researcher might touch, from sample size planning to choice modeling. R's syntax is odd if you've programmed before and merely unfamiliar if you haven't, which is why researchers with no programming background often find it easier than engineers expect.
Python: strongest when quant work touches the rest of the data stack
Python does statistics through pandas for data handling, SciPy and statsmodels for the tests, and matplotlib or seaborn for the charts. For pure analysis, R is more direct: what R does in one built-in line often takes Python an import and some assembly. Python's case is everything around the analysis. If your quant work touches log files, APIs, text processing at scale, or anything a data engineering team built, Python speaks that world natively. The practical rule: researchers who mostly analyze choose R, researchers who mostly wrangle choose Python, and choosing "wrong" costs little because the concepts transfer entirely.
SQL: the language of getting your hands on the data
Nobody lists SQL as a statistics tool, and it decides more quant projects than any statistics tool. Behavioral data lives in a warehouse, and the gate between you and it is either a query you can write or a ticket in someone else's backlog. A researcher who can write a competent SELECT with a JOIN and a GROUP BY stops asking permission to see their own product's data. It's also the least intimidating language here: a useful working subset is a weekend of practice, not a semester.
JASP and jamovi: menus with the receipts
Two free, open-source applications cover the gap between spreadsheet and code. Both give you point-and-click statistics with output built for reporting. JASP is the stronger of the two for Bayesian analysis, which it makes as easy as the classical versions. jamovi was started by JASP developers and has one feature we'd call pedagogically load-bearing: it shows the R syntax behind every analysis as you click, so using it quietly teaches you the code you'd need to graduate. If SPSS is the software you learned in grad school and stopped having a license for, either of these is the free replacement, and jamovi is the one that leaves a door open.
Notebooks, and AI in them
Wherever you land on language, do the work in a notebook, Jupyter if you are in Python, Quarto or R Markdown if you are in R. It keeps code, results, and your reasoning in one document that re-runs top to bottom. The notebook is the answer to the spreadsheet's lab-notebook problem, and it's also where AI assistance is genuinely useful, because modern AI writes competent R and Python from a plain-language description of what you want. That dissolves the syntax barrier, which was always the worst argument for avoiding code. It does not dissolve the judgment: AI will cheerfully choose a test, run it, and narrate a confident result without noticing the assumptions are violated. Let it type. Don't let it decide. We cover this in our guide to using AI for UX research.
How to choose without agonizing
The decision looks bigger than it is, because the tools sort themselves once you're honest about two things: how often you'll repeat the analysis, and where the data lives.
Start with repetition. If you'll run this analysis exactly once, on a dataset that fits comfortably in front of you, use the spreadsheet and move on with your day. Nothing here improves on it for a one-time answer. The moment you know you'll run the same analysis again next month or next quarter, that changes: write it as a script, because in code the second run is free, and in a spreadsheet you pay full price in clicks every single time.
Then look at where your data lives. If the answer is a company warehouse, learn enough SQL before you learn anything else, because no statistics tool can analyze data you can't retrieve. If your data arrives as exports and downloads, you can skip straight to the analysis layer.
Within that layer, let temperament and job description pick the tool. If you want real statistics but aren't ready for code, jamovi gives you menus today and shows you the R behind them, so you're learning the next tool while using this one. If analysis is the core of your job, learn R, because it was built for exactly this work. If your projects lean toward wrangling logs, text, and pipelines, learn Python, because it lives closer to that world. And if you're worried about picking wrong, don't be. The concepts are the hard part, they transfer completely, and switching languages later is a far smaller project than learning the first one.
Visualizing quantitative results for decisions
A chart is not decoration on a finding. For most stakeholders, the chart is the finding, which means how you visualize a result decides whether it lands. Two jobs sit here, and researchers often do the first and skip the second.
The first job is to show the numbers honestly. Match the chart to the question, show the distribution and not just the average, put uncertainty on the page rather than hiding it, and strip anything that is not carrying information (Nussbaumer Knaflic, 2015; Cairo, 2019)Nussbaumer Knaflic, Storytelling with Data, 2015; Cairo, How Charts Lie, 2019.
The second job is to make the number land for a decision. Lead with the takeaway in the title, so a chart says "Design B saves users eight to thirteen seconds" rather than "Task time by design." Put one message on one chart. Annotate the specific point that matters. And translate the statistic into a frame a non-statistician holds naturally, because "roughly one in four users" moves a room that "a 23% difference, p under .05" does not. Presenting an outcome as a natural frequency rather than a bare probability is one of the most reliable ways to be understood (Gigerenzer, 2002)Gigerenzer, Calculated Risks, 2002, and showing uncertainty without either overclaiming or hedging into uselessness is a craft worth studying (Spiegelhalter, 2019)Spiegelhalter, The Art of Statistics, 2019.
Choose the chart by the job it has to do:
| What the reader needs to see | Better chart choices | Avoid |
|---|---|---|
| Compare designs or groups | dot plot, bar chart with confidence intervals, slope chart for paired comparisons | 3D bars, unlabeled averages, ranking without uncertainty |
| Show a distribution | histogram, box plot, violin plot, strip plot | only the mean, especially for skewed task times |
| Show a proportion | stacked bar, icon array, natural-frequency annotation | pie charts with many slices |
| Show change over time | line chart with event annotations and uncertainty where it matters | dual-axis charts unless the relationship is truly the point |
| Show a funnel | step funnel with denominators at every step | percentages with no raw counts |
| Show uncertainty | confidence intervals, plausible ranges, shaded bands | hiding uncertainty in an appendix |
From the Classroom After the final presentations in our Statistical Methods course, the instructor's guidance on communication was pointed. Executives do not ask about t-statistics or confidence intervals. They want the high-level takeaway. Keep the R-values and t-values in a back-pocket appendix slide, lead with the message, and handle the one question that does reliably come up, whether the result is significant, with an asterisk and a key rather than a table.
The decision layer: deciding the rule before you see the data
A significant result is only the beginning of the work that matters. It counts once it changes what a team builds or funds, and reaching that point depends on a decision that has to be made before the study runs rather than after.
There is a humbling reason this discipline pays off. When ideas are put to a controlled test instead of shipped on conviction, most do not move the metric they were built to move. Kohavi, Tang, and Xu report the broad Microsoft pattern as roughly a third positive, a third flat, and a third negative across well-designed changes (Kohavi, Tang & Xu, 2020)Kohavi, Tang & Xu, Trustworthy Online Controlled Experiments, 2020. In heavily optimized products such as Bing and Google, Kohavi and Thomke put the success rate lower, around 10 to 20% of experiments (Kohavi & Thomke, 2017)Kohavi & Thomke, The Surprising Power of Online Experiments, Harvard Business Review, 2017. The lesson is not the exact percentage. It is that conviction is a bad forecasting tool.
Decide the rule before you see the data. Concretely, that means three habits that all point the same way. Ask what decision a measurement would actually change before you run it, and do not measure what will not move a decision, because across a huge range of real analyses most variables carry almost no decision value and only a handful are worth the effort (Hubbard, 2014)Hubbard, How to Measure Anything, 2014. Commit to a single success metric in advance rather than fishing for one that looks good afterward, the discipline experimenters call an overall evaluation criterion, a composite measure agreed on before the experiment so the whole team is aligned on what winning means (Kohavi, Tang & Xu, 2020)Kohavi, Tang & Xu, 2020. And set your default decision and decision criteria before the numbers arrive, so the data informs the call instead of being reverse-engineered to justify it (Kozyrkov, 2019)Kozyrkov, The First Thing Great Decision Makers Do, Harvard Business Review, 2019.
Pair every effect with its size and its business impact, so a significant result is not automatically treated as an important one. And deliver the finding as a recommendation the team can act on rather than a report it has to interpret. Carl Pearson, a staff quantitative UX researcher, models research impact as three layers, the execution of the study, the influence on a stakeholder's decision, and the outcome the organization measures, and argues that combining the right statistical measures is what turns a report from a loose claim into a story a stakeholder can act on (Pearson)Pearson, Research impact: a researcher-centric framework.
A null result is not automatically a green light. It becomes decision-grade only when the study had enough precision to rule out differences that would have mattered, or when the team declared in advance how small a difference it was willing to ignore.
From the Classroom In one exercise from our Statistical Methods course, two features tested no different from each other on ease of use. That was useful only because the team had a practical decision to make: if ease of use did not separate the options within the study's precision, cost and resourcing should decide. In another, a beta discount showed no meaningful relationship to either satisfaction or likelihood to subscribe, which supported a recommendation to stop discounting, because it was cannibalizing revenue without buying loyalty or growth.
Decide the Rule Before You See the Data The value of a quantitative study is set before it runs, by naming the decision it will change and the threshold that will change it. Ask what a measurement is worth, commit to a success metric in advance, and fix your decision criteria before the data arrives.
The business value of quantitative UX research
The business value of quantitative UX research concentrates in three moves: settling decisions that are too expensive to get wrong, giving experience quality a number leadership can watch, and correcting for confident intuition. You already know how to find the truth in a conversation. Quantitative work is what lets you put a number on that truth, defend it in a roadmap meeting, and track whether shipping it actually changed anything.
Settling decisions that are too expensive to get wrong. Some calls cannot be made responsibly from qualitative evidence alone, not because the evidence is weak but because the decision needs a winner, at scale, across segments. Carl Pearson, who has worked as a quant UX researcher at Meta, describes a familiar version: a team has to commit engineering time to one navigation design, and no interview study can confidently declare which variant wins on behavior. A quantitative usability test measuring completion rates across prototypes can. So can a tree test built from nothing but the menu labels (Pearson, 2023)Pearson, What Does a Quantitative UX Researcher Do?, 2023. The same logic covers the sizing problem every qualitative study creates: interviews might surface twenty real customer needs, and the roadmap holds three. MaxDiff and conjoint studies are how you rank those twenty by what users will actually trade for them. The value here is concrete and easy to state to a stakeholder: engineering time not spent building the losing version.
Making experience quality a number leadership can watch. Revenue has a dashboard. Until you build one, user experience does not, and things without dashboards lose prioritization fights. This is the problem Kerry Rodden's team was solving when they created the HEART framework at Google, treating large-scale behavioral data as one more research method and mapping product goals to user-centered metrics like adoption, retention, and task success (Rodden, Hutchinson & Fu, 2010)Rodden, Hutchinson & Fu, Measuring the User Experience on a Large Scale, ACM CHI 2010. Adoption is the most concrete of them: something as ordinary as the percentage of users who apply Gmail labels turns a fuzzy question about whether people actually use a feature into a defined number you can track over time, instead of one that gets argued out in a meeting (Rodden)Rodden, How to Choose the Right UX Metrics for Your Product, GV Library. Give an experience a metric like that and it stops losing prioritization fights to the things that already have one.
Correcting for confident intuition, including everyone's. The most honest argument for measurement is that experienced people, at every level, are unreliable at predicting what will work. Harvard Business School's Stefan Thomke, who has spent years studying experimentation at Booking.com, notes that around nine in ten new ideas fail to improve the metrics they target, and that team confidence is no predictor of which ones (Thomke, 2019)Thomke, At Booking.com, Innovation Means Constant Failure, HBR Cold Call, 2019. His favorite illustration involves the company's own leadership. In late 2017, Booking.com's design director proposed stripping the meticulously optimized home page down to a bare search box. The CEO was skeptical it would confuse loyal customers, and the head of the experimentation team bet a bottle of champagne the test would tank conversion. The test ran anyway, because the company's operating principle is that data outranks opinion and any employee can launch a test without management's sign-off, a principle it exercises roughly 25,000 times a year (Thomke, 2020)Thomke, Building a Culture of Experimentation, Harvard Business Review, 2020. For the record, the skeptics were right: the dense, battle-tested page beat the minimalist one. Which is the point. Sometimes the confident person is correct, and sometimes they are the nine in ten. The test is the only way to know which conversation you are in, before the redesign ships to everyone.
If you want to build this skill on your own product data, our course on quantitative methods for product teams is designed for researchers with no statistics background [internal link: /course/fundamentals-of-quantitative-ux-research], and you can browse the full catalog of UXR Institute courses [internal link: /ux-research-courses].
Frequently asked questions
What is the difference between qualitative and quantitative UX research? Quantitative UX research estimates what is happening, how much, how often, and whether a tested change caused an effect. Qualitative UX research is better suited to explaining why something is happening, what users experience, and what you did not know to ask. They produce different kinds of knowledge rather than different grades of rigor, and the strongest studies sequence the two rather than choosing a side.
What are the different types of quantitative user research methods? They fall on two axes at once. By purpose they are generative, evaluative, or summative, and by data source they are behavioral or attitudinal. Common study designs and evidence families include A/B testing, quantitative usability studies, survey studies, standardized usability surveys using instruments like SUS, tree testing, card sorting, product analytics studies, and choice-modeling studies like MaxDiff and conjoint. Funnel and cohort analysis are common analyses inside product analytics, not separate methods. The same technique can serve different purposes depending on how you deploy it.
What is the difference between statistical and practical significance? Statistical significance means a result would be surprising if nothing were going on, which the p-value reports. Practical significance means the result is large enough to matter, which effect size, a confidence interval, and business judgment tell you. They are independent: a large sample can make a trivial difference statistically significant. Always report an effect size alongside the p-value, and decide the smallest effect that would change your decision before you run the study.
How much sample size do you need for quantitative user research? It depends on the method and the size of the effect you want to detect. You will often hear that significance tests like t-tests and ANOVA start around thirty participants per group, but that figure is a widely repeated rule of thumb with no real power basis, and choice-based methods like MaxDiff and conjoint usually need a few hundred. Treat those numbers as rough anchors only. The real answer comes from a power calculation based on the smallest effect that would change your decision, the variability of your measure, and the confidence you need, not from a fixed figure.
Do you need to know statistics or how to code to do quantitative user research? No. The core skill is matching a clear question to the right test and interpreting the result honestly, not deriving the math. Researchers new to quantitative work can get a solid start in a spreadsheet or a free, menu-driven app like JASP or jamovi, and the concepts are learnable in weeks. Code becomes useful when you outgrow those tools, not before.
Statistics Prep for UX Researchers
The sampling, distribution, and significance concepts the quant courses assume you already have.
Conjoint Analysis & MaxDiff
Choice modeling for prioritization, bundling, and roadmap trade-offs.
How to Design Survey Questions Like a Pro
Write questions that hold up under scrutiny: ambiguity, double-barrels, and hidden assumptions.
Survey Methodology for Product Impact
End-to-end, rigorous survey design, from construct development to analysis and delivery.

Leo Hoar, PhD is the founder of the UXR Institute, where experienced researchers sharpen the strategic and quantitative skills that turn findings into decisions. He writes and teaches on research methods and the craft of making research matter. Read more at his bio page.

