← Back to list

How M+ Tier Lists Are Calculated: The Math Behind Rankings

Open three different World of Warcraft ranking sites on the same Tuesday. The first will tell you that one specialization is the best…

Sam Morris · 2026-05-26 12:54 · 0 claps · 4.8 min read
#gaming #worldofwarcraft #wow #midnight
Open on Medium ↗
Wiki topics: GEN · Genomics & Sequencing 📐 · Mathematics 🎮 · Gaming 🛠️ · Crafts & DIY

How M+ Tier Lists Are Calculated: The Math Behind Rankings

Open three different World of Warcraft ranking sites on the same Tuesday. The first will tell you that one specialization is the best damage dealer in the game. The second will rank it second. The third will put it in the middle of the pack and add a caveat about how it depends on the affix week.

All three are right. None of them is lying. They’re measuring different things, and which thing matters depends on what you’re trying to do.

The trick to using ranking sites isn’t picking which one to trust. It’s understanding what each one is actually measuring, so you can read three of them and triangulate the truth in the middle.

The three places the numbers come from

Almost every ranking site pulls from one of three sources. The first is a major log-aggregation platform that ingests damage and healing data from completed runs and lets writers query the parses. The second is a rating service that scores every completed run on level, timer, and dungeon difficulty, then ranks players by aggregate score. The third is custom data feeds — smaller sites that subscribe to the larger aggregators’ APIs and build their own dashboards on top.

The wow.gg m+ tier list uses live aggregated data that refreshes every few hours. It ranks specializations by their average dungeon score, weighted by how often they appear in the top key bracket. The weighting is the important part: a specialization that does many high-level runs at a strong average score outranks one that does many low-level runs at the same score, because the harder bracket carries more weight.

That’s one methodology choice. Other sites make different ones. Some skip the weighting and rank by raw damage parses. Some use the rating service’s numbers directly. Some build composite scores from multiple sources. The same specialization can land in different tiers on different sites without any of them being incorrect.

The personalities behind the rankings

The data-only rankings have one personality. They’re fast, they’re objective in the narrow sense of being formula-driven, and they’re missing the context that makes a ranking actually useful.

The editorial rankings have a different personality. A guide writer with multiple seasons of high-level experience weights specializations based on what the data says plus what the game actually feels like to play at the top. The result captures things data alone misses — “this specialization is technically near the top but feels exhausting to play.” That’s a real piece of information. It just doesn’t show up in any spreadsheet.

The hybrid rankings, which the wow.gg approach is closer to, calculate the rankings from live data but group the tiers based on actual gaps in that data rather than arbitrary thresholds. If three specializations are within touching distance of each other, they go in the same tier. If there’s a real gap to the next-best specialization, the tier breaks there.

The honest answer about which approach is best is that they’re all useful for different things. Data-only rankings are best for tracking patch-to-patch changes. Editorial rankings are best for picking a class to learn. Hybrid rankings are best for getting a quick read on what the meta actually looks like.

The biases the data can’t fix

Every methodology has blind spots, and the data ones are interesting because they’re invisible to the people relying on them.

Popular specializations have larger data sets than unpopular ones, which makes their numbers more statistically reliable. It also makes them easier to inflate, because the same specialization being played by many different players will produce more outlier good runs than a specialization being played by fewer players. A demonology specialization with a thousand parses in the top bracket will have its average lifted by the ten best parses in the sample. A less-popular specialization with two hundred parses won’t have the same statistical luck working in its favor.

Editorial rankings have a different bias. A guide writer who plays one specialization will subconsciously rank it higher than the data supports. This isn’t dishonest — it’s how human judgment works under uncertainty. The writer has more lived experience with the specialization they play, which means they know its strengths in detail and forgive its weaknesses, while they only know the other specializations from theorycraft and the occasional run.

The bias that matters most is one neither methodology handles well: specializations that excel in narrow situations. A Warlock specialization is great on dungeons with specific affix combinations and mediocre on others. A season-aggregate ranking buries that variance, even though week-to-week the specialization moves through several tiers depending on the week.

What the math actually tells you

A specialization at the top of the rankings outperforms one in the middle by roughly one effective key level — the difference between timing a run with a comfortable margin and timing it at the buzzer. That’s significant, but it’s not “you can’t play the middle-tier specialization.” It’s “if you play it, your margin is thinner.”

Within a tier, the differences are mostly noise. Two specializations in the same bracket are performing identically in any practical sense, which means picking between them based on the rankings is missing what the rankings are actually telling you.

This is why ranking sites shouldn’t be read as commands. A middle-tier specialization you’ve played for three expansions will outperform a top-tier one you picked up last week. Hours of practice beats hours of theory, every time. The rankings capture what a specialization can do at the top of its ceiling, not what it does in your hands.

How to actually use these sites

When a ranking surprises you, check the methodology. Is it editorial or data-driven? What time window? Top hundred runs or all completed keys? The same specialization can look top-tier in the top hundred and middle-tier across the top ten thousand, and both numbers are true.

Live data sites update fastest after tuning patches. Editorial sites update with more context but lag by days. If you’re trying to keep up with a fast-changing meta, you need the live data. If you’re trying to understand why the meta is what it is, you need the editorial.

When multiple sites agree on a ranking, trust the consensus. When they disagree, the specialization is probably situational — strong in some contexts, weaker in others. The disagreement is more useful information than any single ranking would be.

What a ranking actually is

It helps to remember that a ranking is a compromise between speed and context. The faster the ranking updates, the less context it can include. The more context the ranking includes, the slower it updates. You can have either, but not both at the same site, which is why looking at multiple sources is more useful than picking a favorite.

The number on the screen is a snapshot. It captures a moment in time, against a specific set of conditions, with specific assumptions baked in. Reading it as anything more than that — as a permanent truth about which specialization is best — is reading something into it the methodology never claimed.


메타데이터
post_id
48bc2f03ba6d
slug
how-m-tier-lists-are-calculated-the-math-behind-rankings-48bc2f03ba6d
url
https://medium.com/@sam.gamingguides/how-m-tier-lists-are-calculated-the-math-behind-rankings-48bc2f03ba6d
canonical_url
https://medium.com/@sam.gamingguides/how-m-tier-lists-are-calculated-the-math-behind-rankings-48bc2f03ba6d
author_url
https://medium.com/@sam.gamingguides
status
ok
fetched_at
2026-06-25 16:53:31