We Built the World’s Fastest Antisemitism Machine and Named It a Chatbot
There is an old propaganda principle: you don’t need people to believe a lie completely. You just need to repeat it often enough that…
We Built the World’s Fastest Antisemitism Machine and Named It a Chatbot

There is an old propaganda principle: you don’t need people to believe a lie completely. You just need to repeat it often enough that something rubs off. Plant the seed, let it sit and then watch how repetition does the rest. The Nazis understood this and so did Soviet propagandists. Every demagogue with a printing press or a radio tower understood this. You don’t need to convince people that Jews control the banks. You just need to make sure the association surfaces often enough, in enough contexts, that it starts to feel like something someone reasonable might wonder about.
Now we have something that can repeat a lie faster than any printing press, in a friendly conversational tone, to billions of people simultaneously. We call it a large language model, and we trained it on everything humanity has ever written down. Including, it turns out, centuries worth of antisemitism.
The ADL released its first AI Index in January 2026, the most comprehensive evaluation to date of how the six major AI models handle antisemitic and extremist content. Over 25,000 prompts, 37 subcategories, tested against ChatGPT, Claude, DeepSeek, Gemini, Grok, and Llama. The researchers weren’t hunting for elaborate jailbreaks or adversarial edge cases. They were testing how ordinary users encounter these models in normal everyday use. The result is such that every single model failed in some meaningful way. Every single model.
The question isn’t which model failed less. The question is why they all failed at all, and what that tells us about what we actually built.
What the Scores Show and What They Don’t
Let’s review the scores. Claude came in first with 80 out of 100. ChatGPT second at 57. DeepSeek 50, Gemini 49, Llama 31. Grok last, at 21.
A 59-point spread from first to last is not a minor performance variation. Grok scored zero in multiple test categories, what the ADL called “complete failure” in analyzing documents and images for hateful content. In anti-Zionist bias detection, Grok scored 18 out of 100. 18 is not even an F.
Grok is Elon Musk’s AI, which is relevant context and not only because of the obvious. Musk has spent the last few years cultivating a platform, X, that has become a reliable amplifier for antisemitic content while simultaneously claiming that any criticism of this is either exaggerated or a coordinated smear campaign. His AI reflects his platform, which reflects his choices and his agenda. It is not a bug, it is a feature.
But I don’t want to make this piece entirely about Grok, because making it about Grok lets everyone else off too easily.
The Actually Scary Part
Here is the finding that should bother everyone more than the scores: the models were significantly better at catching obvious, classic antisemitic tropes than they were at handling newer, more politically coded forms of the same thing.
Ask an AI to evaluate whether Jews control the media or whether Jewish bankers run the world, and most models will correctly identify that as antisemitic and refuse to engage with it straight. That’s the low-hanging fruit. Those tropes are so well documented, so extensively discussed, that they are basically in every training dataset’s “this is antisemitism” examples file.
But ask about anti-Zionist narratives, about conspiracy theories that are one hop removed from a direct slur, about extremist ideologies that dress ancient hatred in contemporary political language, and the models start to stumble. The ADL found that all six LLMs struggled most with this category. Not just Grok. All of them.
This matters because antisemitism has always been adaptive. It doesn’t stay in one costume. The same paranoia about Jewish power and disloyalty that produced medieval blood libel accusations produced 20th-century fascist propaganda, and is producing contemporary “globalist” rhetoric and anti-Zionist conspiracy theories today. The hatred is the same; the vocabulary rotates. And our AI models, trained on the internet, which contains multitudes of this content, have apparently learned to recognize only the vintage versions.
There was also this: one model, when prompted about the Holocaust, cited an article from a known antisemitic outlet called The Unz Review, specifically a piece titled “The GoySlop Ideology,” complete with references to “Jewish financiers.” An ADL researcher said, with what I imagine was significant restraint, that this was “very low-hanging fruit the model shouldn’t cite from.” The researcher is correct. It is also astonishing that the model did it anyway, because these systems have guardrails, and the guardrails are supposed to catch exactly this kind of source. And yet.
And in a separate December 2025 study, the ADL found that in 44% of cases, AI models would generate the addresses of synagogues alongside the nearest gun stores when prompted to do so. I’ll let that one sit there without embellishment, because it doesn’t need any.
The Internet Has a Jewish Problem, and So Does Your Chatbot
Here is the structural issue that none of the corporate press releases about AI safety and responsible development want to address directly: large language models are trained on the internet. The internet contains, in enormous quantities, antisemitic content. Not fringe content buried in dark corners but content that lives in comment sections, forums, news sites, social media, and the general sewage flow of human discourse online. The models absorb it. The models learn from it. And then the models reproduce it, with varying degrees of guardrail intervention depending on which company built the guardrails and how seriously they took the problem.
Research not related to the ADL study found that 20 tested models applied antisemitic stereotypes to Jews in 69% of cases. That is not a rounding error. That is a feature of what these systems learned. Models were more than twice as likely to apply antisemitic statements to Jews than to non-Jews. The same research found that bias was even worse in sentences related to Israel, Palestine, and Zionism — which is exactly the territory where public discourse has been most saturated with conspiracy theories and dehumanizing rhetoric for the past several years.
Billions of people use these models. Students use them for research. People use them to understand current events, to process information, to form opinions. The AI model is increasingly the first stop, not the library, not the newspaper, not even a search engine. It is the oracle. And the oracle learned to speak from the internet’s worst instincts, imperfectly restrained by safety systems that, by the ADL’s own measurement, are not up to the job.
The “We’re Working On It” Response
The ADL says it contacted every platform and that most were receptive to the findings. This is the part of the story where I am supposed to feel reassured, and I don’t, for a few reasons.
First, “we are working on it” has been the tech industry’s answer to content moderation failures for twenty years, and in twenty years the content moderation failures have not substantially improved. The failures have scaled. The platforms have grown. The harm has grown with them. “We’re working on it” is a PR strategy, not a solution.
Second, the problem is not simply that the guardrails are improperly calibrated. The problem is that the training data itself is contaminated, and you cannot fully solve that by adjusting the output filters. You can reduce the harm at the surface. The learned associations remain.
Third, there is an economic incentive problem that nobody wants to name clearly. Making AI models maximally useful for the broadest possible user base means not being too aggressive about content restriction, because aggressive restriction produces false positives and frustrated users and bad press about censorship. The guardrails exist in tension with the product goal. In that tension, the guardrails will lose, consistently, at the margins.
Why This Is Personal
I grew up understanding that antisemitism is not primarily a problem of extremists. Extremists are the visible tip. The deeper problem is how thoroughly antisemitic assumptions have been woven into ordinary culture, ordinary language, ordinary assumptions about who Jews are and what they want and where their loyalties lie. You don’t need a swastika. You need a thousand small things that seem individually arguable and collectively form a picture.
AI models, trained on the accumulated text of human civilization including its worst impulses, are now reflecting that picture back at us. They are doing it in a pleasant, helpful, authoritative voice. They are doing it at scale, to billions of people, many of whom have no framework for recognizing what they’re receiving.
The ADL’s report is a starting point. The scores will presumably improve — the companies have been notified, they are responsive, the benchmarks are now public. But a better score on the ADL index is not the same thing as solving the underlying problem, which is that we have built systems that learned from humanity’s record and inherited its pathologies, and we are deploying those systems as if they are neutral.
They are not neutral. They were never going to be neutral. The only question is whether we are honest about that, or whether we keep accepting “working on it” as a satisfying answer.
Sources
Six Leading AI Models Show Varied Ability to Detect and Counter Antisemitism and Extremism https://www.adl.org/resources/press-release/six-leading-ai-models-show-varied-ability-detect-and-counter-antisemitism
ADL AI Index https://www.adl.org/adl-ai-index
ADL Ranks Grok as the Worst AI Chatbot at Detecting Antisemitism, Rates Claude as the Best https://www.algemeiner.com/2026/01/28/adl-ranks-grok-worst-ai-chatbot-detecting-antisemitism-rates-claude-best/
Elon Musk’s Grok AI Chatbot Ranks Worst in Countering Antisemitic Content, ADL Study Finds https://www.euronews.com/next/2026/01/29/elon-musks-grok-ai-chatbot-ranks-worst-in-countering-antisemitic-content-adl-study-finds
xAI’s Grok Worst Performing Platform on Countering Antisemitism https://www.upi.com/Top_News/US/2026/01/29/report-AI-bias/8071769690983/
ADL Rates Anthropic’s Claude Best AI Model at Detecting Antisemitism https://jewishinsider.com/2026/01/adl-anthropic-claude-detecting-antisemitism-anti-zionism-content/
You Can’t Always Trust AI to Detect and Counter Antisemitism and Extremism https://jewishlouisville.org/you-cant-always-trust-ai-to-detect-and-counter-antisemitism-and-extremism-adl-study-finds/
ADL Study Finds Leading AI Models Generate Extremist Content After Antisemitic Prompts https://jewishinsider.com/2025/12/adl-study-ai-models-extremist-content-antisemitic-prompts/
Generating Hate: Anti-Jewish and Anti-Israel Bias in Leading Large Language Models https://www.adl.org/resources/report/generating-hate-anti-jewish-and-anti-israel-bias-leading-large-language-models
GPT Is Not an Annotator: The Necessity of Human Annotation in Fairness Benchmark Construction https://arxiv.org/pdf/2405.15760
메타데이터
- post_id
- 1a84dd8d7fe4
- slug
- we-built-the-worlds-fastest-antisemitism-machine-and-named-it-a-chatbot-1a84dd8d7fe4
- url
- https://medium.com/@annisabelle/we-built-the-worlds-fastest-antisemitism-machine-and-named-it-a-chatbot-1a84dd8d7fe4
- canonical_url
- https://medium.com/@annisabelle/we-built-the-worlds-fastest-antisemitism-machine-and-named-it-a-chatbot-1a84dd8d7fe4
- author_url
- https://medium.com/@annisabelle
- status
- ok
- fetched_at
- 2026-06-09 15:37:30