← Back to list

AI in Writing: The Literary Witch Hunt

I started this blog so I could ramble on about my sports-related opinions. But sometimes, I feel the irrepressible urge to ramble about…

C. M. Quarterman · 2026-05-31 15:31 · 4 claps · 7.7 min read
#ai #ai-detection #ai-writing #writing #ai-literacy
Open on Medium ↗
Wiki topics: AI · AI · General LIT · Literature & Writing ✍️ · Writing & Creative 📰 · Journalism & News 🏆 · Sports · General

AI in Writing: The Literary Witch Hunt

I started this blog so I could ramble on about my sports-related opinions. But sometimes, I feel the irrepressible urge to ramble about other things. This is one of those times. And this time, I had so much to write that my original essay got a little away from me. On the recommendation of someone who had to read the original essay, I’ve decided to break this up into a more digestible four-part series. This is Part 1.

Photo by Steve A Johnson on Unsplash

Photo by Steve A Johnson on Unsplash

The Literary Witch Hunt

About three and a half years ago, my brother was falsely accused of using AI to write a short essay for a class he was enrolled in. He was able to prove his innocence (despite the inherent struggle that comes with trying to prove a negative), but his case led my family down a bit of a rabbit hole. Over the years, my family has been contacted by many students with the same story: they wrote their own essay, but an AI detector flagged their writing and they didn’t know how to prove their innocence.

More recently, there have been several stories of professional writers being accused of using AI in their stories, articles, et cetera. Whether or not these writers are using generative AI, I don’t know — and neither does anyone else (unless there’s an incredibly obvious artifact in the text, like the abrupt change to second perspective common to generative AI responses). Generative AI was trained on millions of human-written texts; its responses to prompts will mimic the human writing it was trained on. And yet, people are quick to accuse texts of being AI-generated based on vibes. They point to words and phrases commonly used by AI (read: words and phrases commonly used by humans and mimicked by AI), AI punctuation (like the em-dashes I love to use and refuse to give up), and grammatical structures (lists of three like the one I just used, etc.). Never mind the fact that studies have determined that humans can’t distinguish between AI-generated outputs and human writing (Casal & Kessler, 2023) — they’re sure of the soundness of their analysis.

Then, in order to “definitively” prove that the accused writer is in fact using AI, they pass their writing through an AI detector (itself AI that is provably inaccurate and unreliable), and often through multiple. If the results come back as even partially AI-generated, at any percentage confidence, the author is condemned. They no longer have a recourse to prove their innocence because in the minds of the accusers, the result of the AI detector is infallible. I’ve seen it over and over again. In 1692 terms, it’s a witch hunt. In modern terminology, it’s confirmation bias at its finest.

The AI detector du jour is Pangram. On its site, it claims a .01% false positive rate. I’m not paying to test it (I’ve already wasted enough money on AI detectors), but based on all the past AI detectors I’ve tested, I have my doubts about that false positive rate. Yet it has convinced other AI detector skeptics that it’s “one of the good ones”, and it’s the AI detector being used to attack professionals. Whole books have been passed through this AI detector to prove an author is using generative AI. Its obnoxious CEO, a man I hope I never have the misfortune of meeting, routinely accuses authors and journalists of using AI in their writing. Worst of all, people are buying into his bullshit.

I mentioned earlier that my family has been contacted multiple times by students who, like my brother, were falsely accused of using AI to write, or partially write, their essays. We know they were falsely accused because they were all eager to show us the evidence they had that could show their innocence: handwritten notes outlining the essay, Google Docs edits, etc. One student, a very bright young man who happened to be a non-native English speaker, had read the Liang et al. (2023) study about the higher false positives for non-native speakers, screen-recorded his writing process — an excruciatingly tedious four-hour movie of a 1000-word essay being written, edited, and re-edited. Let’s ignore for the moment that despite the screen-recording and the handwritten essay outline, the university refused to believe the student had written his own essay (Turnitin said he didn’t, so why should they believe anything else?). I was curious what Pangram had to say about his provably human-written essay. Imagine my surprise when Pangram said his provably human-written essay was in fact AI-generated.

Who Is Affected?

To me it seems that people’s belief in the reliability of AI detectors stems from a perception that the advertised false positive rate is, for lack of a better word, equitable. That is to say, a large percentage of the population seems to believe that false positives are equally distributed across the population, thus suggesting that any positives (regardless of their accuracy) are more likely to be true than false.

This seemingly pervasive assumption is false. AI detectors, like so much in this world, are built with structural biases. As a result, false positives affect certain groups at a much higher rate than others. For example, of the falsely accused students that reached out to my family for help or advice, all but one were non-native or multilingual English speakers. Despite GPTZero’s attempts to discredit the 2023 Liang et al. study, that study has in fact been replicated as recently as October 2025 (*Evaluating the Effectiveness and Ethical Implications of AI Detection Tools in Higher Education*, Das Deep et al.).

On a more personal note, when I was testing AI detectors for a study I co-authored in 2023, I discovered that myself (multilingual), my brother (multilingual), and my co-author (monolingual) all had much higher rates of false positives (ranging from 12% to 72%, depending on the person and the detector) than other authors in our corpus of texts. This led my co-author and I to hypothesize that our only common thread, neurodivergence, was also a risk factor for being flagged by an AI detector. While there as of yet has not been a study (to my knowledge) that tests specifically for neurodivergence (such a study would likely have tricky ethical questions to navigate), there is abundant anecdotal evidence on the post-Musk-takeover Twitter replacements which support our hypothesis.

The point is, some people will never have their writing flagged by an AI detector. I suspect these lucky people, who are often the ones to so quickly accuse and condemn someone who “sounds like AI”, find it hard to fathom that some people do not have the same luxury. Some people will get flagged over and over again, despite the efforts they take to avoid false positives.

Frustration

Not too long ago, I was contacted by a journalist who was doing research for an article about AI detectors and their use in academic settings. I had to admit to him that it had been a while since I had done any serious work helping students with their false accusation cases because the anger and frustration was bad for my mental health. Over and over again, my heart went out to these poor kids — students who were eager to prove their innocence only for the universities to reject clear cut evidence, and who, as a result of the accusation, were doomed to spend the rest of their college experiences paranoid that it would happen again. In fact, that short thirty-minute conversation with a journalist about this topic frustrated me so much that I felt the irrepressible urge to write close to 10,000 words about how frustrated I really am. I’ll check with my therapist, but I don’t think that’s healthy.

Just as I was ready to hit publish on the original essay, I decided to do the smart thing and have someone (my father) read the essay over to make sure it didn’t come off as some kind of unhinged manifesto. Alongside some solicited advice (the essay should be split up into parts) and some unsolicited criticism unrelated to the subject of the essay (“I’m not sure I agree with your thesis on the Giants-Dodgers rivalry”), he also recommended that I read Matteo Wong’s recent article for the Atlantic entitled “America Has a Pangram Problem”. Not sure why; I guess he felt like I hadn’t written enough.

I have to say that I am glad that Wong is pointing out that the AI in writing issue is quickly devolving into a witch hunt. I am glad that he explains that, while Pangram is more accurate than its predecessors, there is some amount of futility as generative AI advances at a much more rapid pace than AI detectors. I am glad that he highlights the dangers of false accusations stemming from the results of AI detectors, and that he brings up a high-profile example of a false accusation. I am glad that he acknowledges that there are nuances to how AI is used in writing that AI detectors cannot account for. I would have liked for him to delve deeper into all of these points, but I am glad that they are being discussed.

But I am so, so frustrated. My family and I have been yelling into the void about exactly these issues for years. We have spent three years in fear that one day we’ll wake up and read the story of a student taking irreversible action because of a false accusation. Dramatic? Maybe. But I also know that if I had been falsely accused with no visible recourse when I was suicidal in university, I might have taken drastic actions.

This ‘battle’ has been going on since my brother’s accusation. After his case was resolved, he and my father spoke to the vice head of his university’s academic integrity office, urging them to ban the use of AI detectors. The vice head told them not to worry, because they would be using Turnitin’s “much more accurate” AI detector going forward. In the following two years, my father and I have assisted in two appeals for wrongfully accused students at that university, and have been made aware of close to a dozen more.

For years, we have been shouting into the void, begging people to take this threat seriously. I co-authored a study demonstrating that some people are more affected by false positives than others. After one particularly egregious case, I wrote a letter to the UC Regents. When I didn’t hear back from them, I published it publicly on Reddit. Occasionally, a journalist finds it and asks questions, like, ‘did they ever respond?’ (No).

Now, all of a sudden, people are finally taking notice. Because it’s not just students anymore. That was fine. Justified, even. When a student’s writing is flagged by an AI detector, they probably used AI, right? Because a student couldn’t possibly write that well; a student must be lying, they must be covering their ass; regardless of the evidence, if an AI detector says they used AI, they must have used AI.

So now it’s happening to professionals. Professional writers who have been writing for years, who know they wrote their own piece, who have evidence to back it up, who would love it if people stopped accusing them of using AI in their writing just because some flawed AI detector said so. Professional writers who have themselves “analyzed thousands of posts” on Substack and determined that many are AI-written (using Pangram!) only to be falsely accused of AI writing by the same AI detector, without a single moment of public self-reflection.

Now people care.

*Part 2*


메타데이터
post_id
f12b024f1bfd
slug
ai-in-writing-the-literary-witch-hunt-f12b024f1bfd
url
https://medium.com/@cmquarterman/ai-in-writing-the-literary-witch-hunt-f12b024f1bfd
canonical_url
https://medium.com/@cmquarterman/ai-in-writing-the-literary-witch-hunt-f12b024f1bfd
author_url
https://medium.com/@cmquarterman
status
ok
fetched_at
2026-06-09 15:37:30