Going Beyond “Lady Tasting Tea”
From one woman’s claim to many people’s answers
Going Beyond “Lady Tasting Tea”
From one woman’s claim to many people’s answers
At *8 CUPS AND A LADY*, a small café hidden inside an old apartment building in Asakusa, Tokyo, we offer a café-au-lait version of the famous “Lady Tasting Tea” experiment.

This is the third post in the series. In the first post, I explained how our Lady Tasting Flight works: a blind tasting of six small cups of café au lait. Three are milk-first, and three are coffee-first. In the second post, I wrote about why recreating this experiment is harder than it sounds.
In this post, I want to move from preparation to analysis.
Once customers start submitting their answers, how should we read the data? What does a perfect answer mean? And what can many people’s answers tell us that one person’s answer cannot?
How rare is a perfect answer?
Suppose a participant identifies all six cups correctly.
In our version, the participant knows that exactly three cups are milk-first and three are coffee-first. So the task is to choose three cups out of six as “milk-first.”
There are twenty possible ways to choose three cups out of six. Only one of them is exactly correct. Therefore, even if everyone is guessing completely at random, about 5% of participants are expected to give a perfect answer.
In statistics, 5% is often treated as rare enough to be worth attention, although this is more of a convention than a law.
A perfect answer is certainly worth noticing. But by itself, it does not prove that the participant truly has the ability to distinguish milk-first from coffee-first café au lait. A perfect answer can happen by chance.
This is one of the lessons of the original “Lady Tasting Tea” experiment. The question is not only whether someone got the answer right. The question is how surprising that answer would be if the person were only guessing.
From one woman’s claim to human tasting ability
The original “Lady Tasting Tea” experiment was designed to test a claim about the tasting ability of a specific woman. Fisher’s question was narrow and precise: was her performance in the tasting too good to be explained by chance?
Our café version changes the scale of the question.
If many people try the same tasting, we can ask something broader. Do the accumulated answers suggest that human taste can distinguish milk-first from coffee-first café au lait?
Of course, our data come from customers who choose to try the menu at one café in Tokyo. That is not a random sample of humanity. So we should be modest about what we can conclude.
Still, the direction of the question is different from the original experiment. We are no longer only asking whether one remarkable woman could taste the difference. We are trying to collect enough answers to say something, however small, about human tasting ability.
This does not mean proving that every individual participant has that ability. Some people may be guessing. Some may be more sensitive than others. Some may simply get lucky. But if the overall pattern of answers looks different from random guessing, then the data may contain a signal about what human taste can distinguish.
The task is stranger than it sounds
Having six cups of café au lait may sound like a lot. But statistically, the scoring is more constrained than it first appears.
Because the participant knows that three cups are milk-first and three are coffee-first, the score cannot take every value from zero to six. If you correctly identify one milk-first cup, you must also correctly identify one coffee-first cup. Correct answers come in pairs.
As a result, the possible scores are only zero, two, four, and six.
This already makes the experiment a little unusual. But there is another twist.
If the task is to identify which cups are milk-first, then a score of six is the only perfect answer. A score of zero means the participant labeled every cup incorrectly.
But if our interest is simply whether the participant can tell that there are two different types of café au lait, then a score of zero is also interesting. It means the participant perfectly separated the cups into two groups, but gave the two groups the wrong names.
In that sense, a score of zero may indicate perfect discrimination with reversed labels.
So there are two different notions of “perfect” here.
A label-specific perfect answer has fraction 1 out of 20, or 5%, under random guessing. But a perfect separation of the two groups, allowing the labels to be reversed, has fraction 2 out of 20, or 10%.
This distinction matters because our experiment is not only about whether people know what “milk-first” tastes like. It is also about whether people can detect any consistent difference between milk-first and coffee-first café au lait.
A symmetry that hides one kind of bias
There is another interesting constraint in this experiment.
You might think that we can analyze whether people tend to mistake milk-first café au lait for coffee-first café au lait more often than the other way around.
But in this setup, that bias cannot be evaluated.
Suppose a participant incorrectly labels two milk-first cups as coffee-first. Since the participant must choose exactly three cups as milk-first, those two mistakes automatically imply two mistakes in the opposite direction: two coffee-first cups must have been labeled as milk-first.
In other words, false milk-first and false coffee-first errors are always paired. The structure of the task forces the two types of mistakes to occur in equal numbers.
This is not a limitation of the data size. It is a limitation of the experimental design. With this six-cup forced-choice format, we can analyze accuracy, perfect answers, score distribution, and position bias. But we cannot analyze whether participants are more likely to confuse milk-first as coffee-first than coffee-first as milk-first.
This is a useful reminder that the design of an experiment determines what we can and cannot learn from it.
What we can hope to learn
This café version of “Lady Tasting Tea” cannot answer every question about human taste. The data come from customers at one café in Tokyo, not from a random sample of humanity. The setting is practical, not laboratory-perfect. And the six-cup format itself limits what we can analyze.
Still, the data can be meaningful.
We can ask whether perfect answers occur more often than chance would predict. We can compare the observed score distribution with random guessing. We can look for position bias. And if we collect enough records, we may be able to say something modest about whether human taste can distinguish milk-first from coffee-first café au lait.
The original experiment tested a claim about one woman’s tasting ability. Our version collects many small answers and asks what they suggest together.
Once we have collected enough records, I hope to return to these questions with actual data.
If you have a chance to visit Tokyo and are interested in our blind tasting, please stop by *8 CUPS AND A LADY*.
https://eight-cups-and-a-lady.com/en
[embed]
메타데이터
- post_id
- ebebd6c70a5d
- slug
- going-beyond-lady-tasting-tea-ebebd6c70a5d
- url
- https://medium.com/@tatsurokawamoto/going-beyond-lady-tasting-tea-ebebd6c70a5d
- canonical_url
- https://medium.com/@tatsurokawamoto/going-beyond-lady-tasting-tea-ebebd6c70a5d
- author_url
- https://medium.com/@tatsurokawamoto
- status
- ok
- fetched_at
- 2026-07-08 19:15:55