← Back to list

UX Research Borrowed Half of Statistics, There’s More to Go

UX Research borrowed a lot of methodologies that produce a number but completely forgot the methodologies that are meant to make sure that…

Alina Bezchotnikova in Bootcamp · 2026-05-06 11:15 · 40 claps · 5.2 min read paywalled
#ux-design #ux-research #researchops #product-development #product-management
Open on Medium ↗
Wiki topics: UX · UI/UX Design BIZ · Business Strategy 📋 · Product Management 📐 · Mathematics

UX Research Borrowed Half of Statistics, There’s More to Go

UX Research borrowed a lot of methodologies that produce a number but completely forgot the methodologies that are meant to make sure that that numbers make sense.

Photo by Bozhin Karaivanov on Unsplash

Photo by Bozhin Karaivanov on Unsplash

User research has many parents with the biggest being anthropology, HCI, human factors, design research, and statistics. As statistics is my specialty, in this article I’m going to focus primarily but not exclusively on that.

What we borrowed (and mostly broke)

Sampling. We took the word but dropped all the theory around it. We say “we talked to 8 users” like it’s supposed to represent a sample but a sample supposed to have a frame, a selection mechanism, and a defined population it generalises to. To avoid that problem we call people participants, which is fine, as long as we are aware that this might not scale the way sampling does.

Significance testing. A/B testing brought p-values into prodcut development. A p-value is the probability of observing data at least as extreme as what you saw, assuming the null hypothesis is true — which is not “the probability your variant is winning” but is treated that way constantly. We borrowed the ritual without the reasoning, which is how we ended up with the peeking problem or the multiple comparisons problem.

Survey design. Likert scales, the entire NPS apparatus, satisfaction batteries — psychometrics, all of it, lifted wholesale. The scales mostly survived the trip but the thinking around scale validation, reliability, and measurement invariance didn’t.

Descriptive statistics. Means, medians, percentages. Used constantly, usually correctly — the lowest bar a quantitative discipline can clear.

Correlation. Borrowed and routinely confused with causation in research readouts. Nothing new.

Thematic analysis has roots in content analysis, which has roots in early 20th-century quantitative communications research. Worth mentioning that not all qualitative researchers think coding should have inter-rater reliability. Constructivist approaches reject the idea that there’s a single “true” coding scheme to converge on but the field rarely makes the position explicit, so coding decisions get treated as obvious when they’re contested.

What we should have borrowed but mostly didn’t

This is the longer list, which tells you something.

Power analysis. Before-the-fact. The single most useful planning tool in applied research. The famous “five users find 85% of usability issues” comes from a 1993 model by Nielsen and Landauer that assumed a problem-detection probability of about 0.31. Subsequent work has found that probability varies enormously by task complexity, user heterogeneity, and study type, sometimes higher but often lower. The rule is not particularly wrong but it was always context-dependent and got flattened into a universal. Power analysis would tell you what your specific study can and cannot detect before you run it. We don’t do it because, most likely, we are not going to like the answer.

Bayesian updating. Frequentist A/B testing asks: “assuming nothing is happening, how unusual is this data?” Bayesian asks: “given what we already knew and what we just saw, what should we now believe?” The second framing maps better onto how research actually works — you have priors, you run a study, you update. The catch is that priors have to be specified and defended, which is its own argument. But for the inferential question most product teams actually want answered, the Bayesian framing is closer to fit.

Sequential analysis. A whole branch of statistics built around “can I stop early?” Wald developed it for WWII munitions testing. Product teams reinvent it badly every time someone watches a dashboard and asks “do we have enough data yet.”

Survival analysis. Time-to-event modeling with censoring made for retention, churn, time-to-first-value, time-to-task-completion. We have all of this data but we almost never use the right tools on it. Kaplan-Meier curves should be in every PM’s vocabulary.

Mixed-effects models. Repeated measurements on the same user, or users nested inside teams nested inside accounts, are hierarchical data. Treating it as flat inflates your effective sample size and produces overconfident conclusions. UX research has hierarchical data constantly and ignores the hierarchy.

Causal inference. Difference-in-differences, regression discontinuity, instrumental variables, Pearl’s do-calculus, it’s all different tools for different identification problems, all under serious assumption burdens — no unmeasured confounders, parallel trends, valid instruments. This is not something easy at all but product teams ship interventions and then guess at causation, when at least some of these methods could give a defensible answer with the data we already have.

Measurement theory. UX surveys should include construct validity, reliability, convergent and discriminant validity, measurement invariance. The bedrock of psychometrics.

Experimental design beyond A/B. Factorial designs, fractional factorials, conjoint. When you have four variables to test, running four sequential A/B tests is the worst possible answer and the most common one.

Missing data theory. Rubin’s MCAR/MAR/MNAR framework tells you when dropouts are dangerous and when they’re benign. UX research has dropout in every longitudinal study and survey, and the standard treatment is to ignore it.

Effect sizes with uncertainty intervals. Significance tells you something is happening but effect sizes tell you whether it matters. Confidence intervals tell you how confidently you know that. Reporting “p < 0.05” without an effect size and an interval is half a finding — and the modern statistical consensus, post-ASA’s 2016 statement on p-values, has moved firmly in this direction.

Pre-registration. The cleanest defense against peeking, p-hacking, and the garden of forking paths. Common in psychology now and almost unheard of in product research, despite product research having all the same problems.

Why this gap exists

Two stories get told about why UX research isn’t more statistically rigorous. Both are partially true and they pull in opposite directions.

The first is training: the methods above require real coursework, and the field has chosen accessibility over depth at most forks in the road. You can run a usability study after watching a YouTube video but it’s treacky to run a power analysis or fit a survival model that way. That’s real.

The second is a structural: most UX research isn’t underpowered because researchers don’t know about power analysis. It’s underpowered because the populations are small, recruiting is expensive, the timelines are quarterly, and the audience for findings is a PM with two weeks on the roadmap. The behavioral data we’d need for survival analysis or causal inference often lives in product analytics and not in research. Pre-registration assumes you have the seat at the table to commit to a design before stakeholders see preliminary numbers. Many researchers don’t have that luxury.

Neither story is the whole picture. A researcher who has tried to introduce power analysis in a product org felt that structural constraints. A researcher who has only invoked structural constraints knows the training gap is real too. Both can be true, and the path forward involves both — researchers building deeper quant fluency where they can, and product leaders building the org conditions that make rigor possible.

What’s actually at stake

The discipline that gets serious about its quantitative methods will absorb work that the discipline that doesn’t currently does. That’s already happening — it’s why product analytics teams are increasingly making decisions that used to belong to research, and why “research ops” keeps drifting toward measurement infrastructure. Some of that drift is healthy specialization. Some of it is research ceding ground it didn’t have to cede.

The shelf is right there. Statistics is one parent discipline among several, not the parent discipline. But it’s the one we’ve been most selective about, and the parts we left behind are the parts that would push back on our conclusions. That’s worth fixing — not because every researcher needs to become a statistician, but because the methods that interrogate a number are how a discipline grows up.


메타데이터
post_id
d3cf419eb7d5
slug
ux-research-borrowed-half-of-statistics-theres-more-to-go-d3cf419eb7d5
url
https://medium.com/design-bootcamp/ux-research-borrowed-half-of-statistics-theres-more-to-go-d3cf419eb7d5
canonical_url
https://medium.com/design-bootcamp/ux-research-borrowed-half-of-statistics-theres-more-to-go-d3cf419eb7d5
author_url
https://medium.com/@alina.bezchotnikova
status
ok
fetched_at
2026-06-09 15:37:30