← Back to list

How visualizations of uncertainty can impact people’s perceptions

Link to original paper this article references…

Thomas · 2025-03-08 04:29 · 0 claps · 8.4 min read
#visualization #reserach #confidence-interval #prediction-interval #statistics
Open on Medium ↗
Wiki topics: 📐 · Mathematics

How visualizations of uncertainty can impact people’s perceptions

Link to original paper this article references: https://www.dangoldstein.com/papers/Hofman_Goldstein_Hullman_Visualizing_Uncertainty_Mislead_Scientific.pdf

The Problem at Hand

Imagine you’re a researcher and you’ve created what you believe might be a slightly improved prototype of a product being sold to consumers. A slightly more effective medication, a toothbrush that brushes just a bit better, a shoe that can run more miles before breaking down, or anything else you can think of! When it comes time to publish your findings, you want to make sure your visualizations accurately display the results of your findings. You don’t want customers to be misled that your new iteration is substantially better or worse than it is. If someone underestimates the effects of some treatment, for example, they may be less inclined to use it and thus have a worse outcome than they otherwise could have. However, if they overestimate the effects of some treatment, they may be inclined to spend a lot of money for something that doesn’t really help them.

This raises a question: how does the visualization you choose impact how people view your experiment? Which visualizations are better under different circumstances?

How we currently visualize

Scientists usually display their results using one of two ways: inferential uncertainty or outcome uncertainty. Inferential uncertainty involves the uncertainty of the average value of a population, whereas outcome uncertainty involves the uncertainty of the range of values an individual can have. A fun example of this may be understood with height. Given we measured a sample of men, we might say the inferential uncertainty of the average height of men is between 5’9” and 5’11”, but the outcome uncertainty for the range of heights we expect most men to be is between 5’4” and 6’4”.

Inferential and outcome uncertainty are commonly shown using confidence intervals (CI) and prediction intervals (PI), respectively, with a confidence level of 95%. We use these intervals because we don’t expect our samples to be perfect representations of the population. For example, let’s say we know the “true” average height of men is 5’10”. If we sample a group of men and the average height is 5’9”, we might make a CI that shows the true average height of men can be between 5’8” and 5’10”. In this case, although our average is slightly off, the “true” average is contained within the interval we created. The confidence level also means that if we sampled men’s heights 100 times, 95 of those samples will have the “true” population average within the interval created. The other five times we might have been unlucky and happened to sample a lot of basketball players or gymnasts, skewing the average. Here’s what this, along with alternative visualizations, looked like in the study which looked at a sliding distance rather than height:

Do either of these do a better job at communicating the expected average value and range of values? Are they better together? Are there other ways to communicate that outperform these two options? Just because scientists commonly use these two methods doesn’t mean they’re the best options available to use.

Running an Experiment

The researchers solved both of these questions: how people judged results from inferential and outcome uncertainty, then how people judged results from alternative visualizations. They did this by creating an online game and recruiting participants via Amazon’s Mechanical Turk. Participants were in a head-to-head athletic competition with equal probability for winning or losing/not receiving a prize of $250, but they were given an option to upgrade their item, at a cost, to get an advantage in winning.

  1. In the first experiment, participants were shown either a visual CI or PI for the results of the new item, and some were additionally shown text that either matched their visualization or contained extra information.

  2. In the second experiment, participants were shown either a CI, PI, a rescaled CI axis, or an alternative visualization (all are shown in image above). Participants additionally either had a small or large effect size, meaning the new item gave a small or large advantage.

Participants had to record what they thought their new odds of winning would be with their item advantage and how much they’d be willing to pay for the upgraded item. This was compared against the mathematical expected value (the authors also used the terms risk-neutral and normative value) of the upgraded item.

Results: the good, the bad, and the ugly

Here are the results of the second experiment below. The first experiment is nearly identical in outcome, so I will omit showing it.

As benchmarks, we can see dotted, vertical lines on the right graphs which represent the mathematical probability of winning with the upgraded item, and the expected value of upgrading the items would be $17.50 for the small effect size and $65 for the large effect size.

One theme is made abundantly clear: people who view CIs tend greatly overestimate effect sizes while those who view PIs, although not perfect, do a better job at accurately understanding the effect size. This was the same conclusion reached in the first experiment. We can see that the rescaling of the axis might help with small effect sizes, but there is still substantial over-valuation.

Given how messy HOPs are (it would have taken 6.3 minutes to see the entire HOP visualization and it’s an animation rather than a static, unchanging, printable image) and their nearly identical performance to PIs, I’m doubtful this will take off for generalized purposes. Also, given how far off the benchmarks CIs were, I think these results will lead to CIs falling out of fashion when communicating statistics to general audiences. Perhaps they still have legitimate purposes when dealing with professionals who are used to dealing with statistics, but soon using a CI to demonstrate results to the masses may be seen as unethical because we see how it tends to cause people to overvalue the effect of the experiment.

But before we start contemplating how good actors can benefit from using a more perceptually accurate visualization and how bad actors can do the opposite, we should look at where this study falls short. As interesting as it is, there are some potentially very important shortcomings in this study.

Limitations

There are always some obvious limitations like people tending to round to the nearest multiple of five or an even number, and the researchers did a tremendous job by adding an “attention check” which removed people who didn’t say an equal chance to win 1-on-1 was equivalent to 50%. But some serious problems remained.

From the bottom right chart of the conclusions, we can see that there was a fair number of people who stated they would have a 100% probability of winning when shown a CI. In fact, the mode of these groups was nearly 100%! Surely people would know that in a game of chance, increasing your odds doesn’t guarantee your chances of success, right? To make matters worse, the stated probability of winnings doesn’t come close to what you’d expect someone’s willingness to pay should be. If you thought you’d have a 100% chance of winning $250 with the new item, why would no one have been willing to pay almost $250 for it? Why would anyone be willing to pay over $125 as that is the maximum expected value for an item that would guarantee a win?

Clearly, people’s choices might not be as interoperable as we think. To throw another wrench, this paper doesn’t even address risk aversion. These limitations could hinder some of the results one may like to draw from this experiment.

Questions I Still Have

  • Should we use alternative visualizations?

The visualizations in this experiment primarily relied on lines, notably for PIs. This single dimension could cause perceptual problems such as giving a sense that outcomes were equally distributed instead of normally distributed. I wonder if using something like an overlapping histogram or boxplots would be better at showcasing differences. Here’s an example of what this could look like:

If the researchers just used the lines seen in PIs, this would just be two long lines with a lot of overlap. Instead, it might be clearer to see how often you’d expect to come out ahead when looking at a histogram or boxplots where you can see quartiles. Really, any graph that can show a distribution could be applied using this logic.

  • If people are bad at figuring these things out, should we be more explicit when communicating results?

It might be common for people to use very technical visualizations in papers which might require a lot of interoperability from readers. But especially for results that are designed to be used for laypeople, I think it’s reasonable to expect there might be some irreducible error in the ability for them to interpret technical graphs into everyday questions such as “how much better off will I be with this slightly improved product?” It might need to become popularized to provide sections in papers which are designed for less-technical readers to be able to get a more accurate understanding of what certain results mean in different contexts.

For example, here is a potential chart that could have been used in this paper’s first experiment to explicitly show their chances of winning using the upgraded item.

As simple as it is, it visually shows that the chances of having an additional win are low. Since this is one of the core questions an end user will have, should these sorts of explicit visualizations be more readily available in papers?

  • Will we need to create new standardizations for visualizations?

I won’t recap my conclusions from this paper again here, but there may be legitimate needs to create standardizations in which visualizations should be used. A researcher knowingly using a CI instead of a PI to persuade an athlete to buy a new item as seen in the experiment may be seen to some as sneaky or immoral, but this would be small beans compared to something like cancer research. Imagine the researcher who knowingly uses CIs instead of PIs to have users overpay for medication that isn’t much better!

I would like to see what dissenting views are for the case against standardizing visualizations in certain settings. I am unaware of any explicit benefits end users may receive when looking at CIs instead of PIs in similar contexts, although my knowledge starts and ends here with this paper; it would surprise me if nothing else was out that there that contradicts this paper or shows important use cases for CIs. Additionally, this sort of regulation always has the potential to backfire. However, we do have rules against deceptive marketing practices which presumably help consumers, so this might be argued as just another extension of that.

  • Does interpreting visualizations under different circumstances change the results?

There’s a lot that can be hashed into some benchmarks the researchers used, notably in the use of risk-neutral or expected values. This benchmark can be unreliable depending on how participants judge risk. Additionally, this can potentially change based upon the situation.

We’ve spent a lot of time learning how graphs appeal visually to users, but there’s another medium how people interpret information: emotions. In the example the researchers used, it was a pretty low risk game. If participants did nothing, they either broke even or won $250 imaginary dollars! The worst case is that they spent money upgrading an item yet didn’t win, but all they do is lose imaginary dollars! This is extremely low stakes. Additionally, participants likely had little to no prior knowledge of the game was simulated. These settings may not be how people experience most data.

Let’s say the circumstances have changed: there’s a new treatment for a chronic illness you have, but it comes with risks. You’re shown CIs, PIs, rescaled axis, and alternative visualizations. Will you still overestimate the improvement in the outcome of the new treatment? I don’t think you would, at least not as much. I think having additional background knowledge and a higher vested interest will cause you to be more skeptical of the study. You might do more vetting and think through the outcomes more.


메타데이터
post_id
5ae9b2fdfba5
slug
how-visualizations-of-uncertainty-can-impact-peoples-perceptions-5ae9b2fdfba5
url
https://medium.com/@reedyt22/how-visualizations-of-uncertainty-can-impact-peoples-perceptions-5ae9b2fdfba5
canonical_url
https://medium.com/@reedyt22/how-visualizations-of-uncertainty-can-impact-peoples-perceptions-5ae9b2fdfba5
author_url
https://medium.com/@reedyt22
status
ok
fetched_at
2026-08-26 14:33:53