Moderation evaluates messages: players experience conversations
What competitive gaming reveals about the limits of moderation systems

AI generated illustration
Moderation evaluates messages: players experience conversations
What competitive gaming reveals about the limits of moderation systems
Valorant is a 5-v-5 online shooter where strangers rely on text and voice communication to coordinate under pressure and prevent the opponent team from planting or defusing a “spike” (bomb) by eliminating them across 24 total rounds. Like many competitive online games, it’s also a place where harassment has become an ordinary part of the experience for many players — and the examples here, though mostly from Valorant, are common across many such games like Counter-Strike, Call of Duty, League of Legends, or PUBG.
My hands were shaking.
Not because we were losing — we weren’t; we’d won most rounds and the match was basically decided. But somewhere in the previous ninety seconds, a stranger had decided I deserved to be spoken to differently, and my body responded the way bodies respond to threats: elevated heart rate, tight chest, unsteady hands.
My brain disagreed. I knew he was nobody — someone I’d never meet again, saying things that had nothing to do with me. But my nervous system wasn’t interested in the argument.
That gap — between what I knew and what my body did — is what stayed with me a long time after that game ended. And it turns out moderation systems have exactly the same blind spot.
I don’t think this story is unusual.
What surprised me afterwards wasn’t the harassment itself — it was realizing how neatly it exposed a design assumption built into almost every moderation system I’ve encountered.
The moment everything changed

AI generated illustration
The first few messages were nothing. Genuinely nothing.
“You’re throwing” (gaming slang for costing your team the match). “Stop peeking” (looking around a corner). “You’re useless.”
If you’ve played competitive games for long enough, these aren’t remarkable. They’re like London weather — you know it’s going to rain so you always keep that umbrella with you. If you’re the guy with the perfect aim who’s never had strangers yell at you over a video game, congratulations — this article probably isn’t about you.
Then I answered a question over voice. Until that moment, I was just another anonymous teammate. The moment I spoke, another player could infer something about the person behind the screen.
There was a pause.
And suddenly we weren’t talking about the game anymore.
Now I should be in the kitchen. Then I should be doing dishes. Then women shouldn’t play this game. And then I was a useless bitch — and, for the record, I was sitting just one kill behind the man saying this. Eventually, he moved on to telling me what should happen to women like me.
Nothing about the match had changed. Not the score. Not my performance. Not the player shouting. Exactly one new piece of information had entered the match: He realized I was a woman.
That’s the thing I want you to hold onto.
Because the abuse wasn’t triggered by anything I had done. It was triggered by something I was, the instant another player could infer it. And for anyone interested in AI systems, that’s a surprisingly important observation.
The single most important variable in that interaction wasn’t in the chat log.
It was in my voice.
It was the variable that no text classifier could ever see.
But this article is not about being a woman in competitive gaming. This is much broader than that.
“Just mute him”

Photo by Josh Eckstein on Unsplash
Whenever conversations about moderation happen online, someone inevitably says the same thing. You might be thinking that too.
“Just mute him.”
It’s common advice, but it’s also more complicated than it sounds. Competitive games rely on communication, so muting everyone from the start often means giving up information that helps your team.
Eventually, I did.
And muting works. It stops the next message. What it can’t do is un-happen the previous ten. It doesn’t lower your heart rate, it doesn’t steady your hands. It doesn’t magically restore the concentration you’ve spent the last three rounds trying not to lose.
It doesn’t give you back the enjoyment that disappeared somewhere between “you’re throwing” and “go back to the kitchen.”
The intervention arrives after the injury.
That’s not a criticism of the mute button. But it reveals something deeper.
Even when moderation works exactly as designed, most moderation tools are designed to stop harassment after it has happened — it is structurally reactive. It doesn’t prevent it. Those aren’t necessarily the same thing. And for a lot of targets, by the time the tool kicks in, the thing you needed protection from has already gone through you.
What these systems actually do
It’s worth saying something else before going any further.
This isn’t an article about Riot Games failing.
In many ways, Riot has invested real effort in trust and safety. Since a 2022 patch, Valorant runs real-time text evaluation that automatically mutes messages flagged as offensive — it acts as the message is sent, not days later. Their text moderation is designed to account for some nuance rather than blunt keyword-matching. And in 2022 they began evaluating voice comms using language models. Call of Duty has a comparable voice system that has flagged millions of accounts. This is a serious, well-resourced field, and the people working in it are not naïve. Millions of messages and voice interactions happen every day, making human review impossible except after the fact.
If we look closely at how the voice system works, because it matters: Riot only records and evaluates voice **after a player submits a report.* The comms aren’t monitored live. Which means, for the abuse that actually shakes you — the voice abuse, the escalating kind — the system was never going to step in during* the match. By design, it looks only afterwards, at a fragment, once I’ve already been through it.
So the real picture isn’t “no moderation.” It’s this: text is moderated in real time but message by message, and voice is moderated reactively, after the fact, on report. And that’s the whole problem.
The systems improved, the experience didn’t
I find one statistic really interesting.

Source: VALORANT Systems Health Series — Voice and Chat Toxicity
In 2022, Riot reported that despite increasing punishments and improving detection systems, the proportion of players who felt they were experiencing harassment hadn’t meaningfully changed.
Read that again.
The systems improved, the experience didn’t. Riot themselves described the work as foundational rather than complete. And it isn’t only Riot saying so. Independent reporting has noted that recording and punishing toxic chat doesn’t necessarily translate into a safer environment for players. I don’t think that’s because the classifiers weren’t good enough. Rather, I think it’s because they were solving a subtly different problem. The system was counting messages while you were living an interaction.
This isn’t unique to Riot.
Every large multiplayer game faces the same trade-off. Moderation systems have to operate at enormous scale, across languages, cultures, and rapidly changing communities. The challenge isn’t that developers don’t care. It’s that they’ve inherited a unit of analysis — the individual message — that is easy to evaluate, easy to explain, and relatively inexpensive to moderate. The question is whether it’s the same unit players actually experience.
It’s not one message

Photo by Volodymyr Hryshchenko on Unsplash
Here’s the thing a message-level view can’t see: harassment has a shape, and it has mass.
Take the actual arc of what gets said to someone over a match:
“You’re bad.” (later) “Go back to the kitchen.” (later) “Worthless bitch.” (later) “Kill yourself.”
Look at those the way a message-level classifier does — one at a time, even with context and nuance. “You’re bad” is fine; that’s just competitive chat. “Go back to the kitchen” contains no slur at all — I could be telling my friend to go grab the ice cream. “Kill yourself” might be a genuine attack, or the throwaway hyperbole players use about themselves fifty times a night. In isolation, each one is defensible. As a sequence, it’s unmistakable: generic frustration, then identity discovered, then gendered abuse, then dehumanisation, then a threat. The harm isn’t in any single message. It’s in the direction they travel.
And it isn’t only the shape. It’s the volume.
When I sat down to write this, my first instinct was frustration that I hadn’t recorded any of it. No screenshots. No clips. No proof. Then I realised that instinct was the point. Because harassment rarely arrives as one spectacular message you could screenshot. Most of the time it’s death by a thousand cuts:
Round two: “Nice one.” (sarcastic) Round four: “Still throwing.” Round six: “Kitchen.” Round nine: “Worthless.” Round twelve: “Kill yourself.” Round fifteen: “Report her.”
No single one of those captures what the match felt like. Together, they do. By the end, I wasn’t reacting to the latest comment — I was reacting to everything that had come before it. Psychologists have known for decades that people don’t remember experiences as isolated moments. We remember narratives, patterns, and emotional trajectories. Harassment is no different. We don’t experience harassment message by message. We experience it as a whole: building in one direction, accumulating the entire time.
A system that evaluates messages sees fragments, and quite reasonably decides most of the fragments are fine.
Meaning doesn’t live in the sentence. It lives around it.

Photo by Glen Carrie on Unsplash
There’s another reason message-level moderation struggles: Meaning doesn’t live inside a sentence.
This is the part that makes it genuinely hard, and genuinely interesting.
“You should be doing dishes” contains no bannable word. It is also, in context, unmistakable misogyny.
Take another phrase like “Go back to the kitchen.” There are no slurs. No profanity. Nothing that immediately stands out to a keyword filter. Yet almost everyone understands it as gendered harassment because of everything surrounding the sentence rather than the sentence itself.
The same problem appears across different kinds of identity-based harassment. My friend is Indian, and the moment he speaks, his accent becomes part of the match. Sometimes it’s a genuine question. But most times it’s the beginning of mockery. The words themselves aren’t always remarkable. The context is. Sometimes it is mockery of the Indian character, Harbor, he plays with. “Nice shot” can be encouragement or sarcasm. “Kill yourself” might be a genuine threat, or the kind of hyperbole players casually direct at themselves after a terrible round. Similarly, words like “gay” or “faggot” aren’t always used as direct attacks — they’re often thrown around as generic insults, sometimes reclaimed between friends, and sometimes used with unmistakably homophobic intent. The word is the same. The context is what determines the harm.
The words alone don’t tell you which conversation you’re in. Because meaning isn’t contained inside a sentence. Meaning comes from who is speaking, who they’re speaking to, what happened immediately beforehand, how the interaction has unfolded, and what everyone involved already understands.
Humans reconstruct that context almost instantly. Moderation systems have to infer it.
That’s what makes contextual harassment such a difficult technical problem. The messages that hurt people most are often the ones with the cleanest words.
And sometimes the harassment isn’t in the words at all.
In another match, a man decided the way to respond to me was a running commentary of sexual innuendo — such as jokes about how big his “helicopter” was. When I ignored him and kept playing, he escalated. While I was defusing the spike, he repeatedly moved his character against mine, using the game’s animations to mimic an unconsented sexual act.
From a moderation perspective, that’s a fascinating failure mode. Nothing was typed, nothing was said over voice. There was no message to classify, no slur to detect, no audio to review. The harassment was communicated entirely through player behaviour.
Every moderation system I’ve talked about so far — text classifiers, voice models, keyword detection — is effectively blind to it. The most invasive part of that interaction left no trace in the data those systems were designed to analyse.
It’s a reminder that meaning doesn’t just live around a sentence. Sometimes it never enters language at all.
The players who quietly disappear

Photo by Chuck Fortner on Unsplash
For many people, online games aren’t just games. They’re where friends meet after work, where relationships are built, and where communities form. Walking away isn’t always as simple as closing the client.
But here’s what all of this actually costs the community — and it isn’t hurt feelings. It’s selection.
This isn’t really about me. I’ve played games long enough that I’ll queue again tomorrow.
A lot of people won’t.
Women, players of colour, LGBTQ+ players, maybe teenagers trying competitive games for the first time, or just someone already having a bad day who decides tonight simply isn’t worth it. This isn’t only anecdotal; the ADL’s 2022 surveys of gamers found the majority of Valorant players report harassment, disproportionately targeting women and players of colour
The people who leave rarely announce it. They don’t post on Reddit explaining that one more match finally tipped the balance. They just stop joining voice chat. Then they stop solo queuing. Eventually, they stop playing altogether.
That’s the cost I think moderation struggles to see. Not the messages but the absences.
A moderation system that can’t see this kind of harm isn’t neutral. It quietly shapes the community around what it can detect. More messages are flagged. More players are punished. The dashboards suggest progress. If the people most likely to be targeted quietly leave, there are fewer opportunities for that harm to happen, fewer reports to submit, and a community that’s gradually selected for the people least likely to experience it. Meanwhile, the players who decided the experience wasn’t worth coming back to were never a metric in the first place.
That’s true far beyond games. Every system decides what it can see. And over time, what it can’t see starts to look like it doesn’t exist.
The unit is wrong

Photo by Siora Photography on Unsplash
This isn’t just a moderation problem. It’s a measurement problem.
As product managers, we spend a lot of time talking about proxy metrics: A food delivery app measures delivery time because it can’t directly measure satisfaction; a fitness app measures streaks because it can’t directly measure long-term health; Duolingo measures daily retention because fluency takes years to observe.
The proxy isn’t wrong. It’s simply easier to measure. Moderation has the same problem. Messages are measurable. Experiences aren’t. So we optimise what we can count. The trouble is that players don’t remember individual messages. They remember the match. They remember the conversation. They remember the moment something changed.
That’s why I find Riot’s own survey so interesting. Punishments increased, yet players still reported experiencing harassment at roughly the same rate. That doesn’t necessarily mean the work failed. It may simply mean the system was getting better at measuring one thing while players were experiencing another.
Because harassment isn’t experienced as isolated messages. It’s experienced as trajectories. Escalations. Conversations that unfold over the course of an entire game, across text and voice, between people who are constantly responding to one another.
Which makes me wonder whether we’ve been measuring the wrong thing all along. Not “how many toxic messages did we detect?” But “did this interaction leave someone feeling like they didn’t belong here?”
That’s a much harder question than “how many toxic messages did we detect?” But it’s the one that matches what players actually experience. Until moderation measures interactions instead of individual messages, it may keep getting better at measuring toxicity while leaving players feeling no safer.
메타데이터
- post_id
- 0c29a2ef730d
- slug
- players-experience-conversations-moderation-evaluates-messages-0c29a2ef730d
- url
- https://medium.com/@achirab/players-experience-conversations-moderation-evaluates-messages-0c29a2ef730d
- canonical_url
- https://medium.com/@achirab/players-experience-conversations-moderation-evaluates-messages-0c29a2ef730d
- author_url
- https://medium.com/@achirab
- status
- ok
- fetched_at
- 2026-07-27 00:37:34