Guilty Until Proven Human
Can universities be sued over false AI detection results?
Guilty Until Proven Human
Can universities be sued over false AI detection results?
Photo by DuoNguyen on Unsplash
Anyone who has ever written a thesis knows how harrowing the whole experience can be.
Imagine spending six months writing a thesis. Missing sleep, and missing the kind of blissful ignorance that comes with not knowing what a literature review is. You use Google Docs, so every edit is time-stamped, and every revision is logged. At this point, Google Docs knows your pain better than your therapist does, assuming you have one.
Finally, the day comes when you do that last edit and finally submit everything. You can finally breathe, or so you think.
Then a letter arrives, and apparently, an AI detection tool reviewed your work, and it’s flagged with a 98% probability that it’s AI-generated. You happen to be on scholarship, and the board has seen the report. Unsurprisingly, the board believes the machine.
You might have bumped into this story recently, or maybe this is the first time you’re hearing about it. A student at a private New York university posted on Reddit claiming that their $45,000-per-year merit scholarship was being revoked after plagiarism software flagged their senior thesis as mostly AI-generated.
According to the post, the academic integrity board refused to investigate the matter or look at any of the timestamped Google Docs edit history showing every paragraph, every revision, or every panic-at-2 am correction. The institution decided to rely solely on the AI detection report, and that was enough for them. Case closed, because, apparently, it never makes mistakes.

Image via Reddit
The story went viral and inevitably triggered a discussion about something most universities have been getting away with for two years: punishing students based on software that, by its developers’ own admission, is not reliable enough to be the deciding vote on whether a student should get a pass or a fail.
The courts are getting involved
The clearest sign came in January 2026. Orion Newby, a first-year student at Adelphi University in New York, won his case against the school after Turnitin flagged his World Civilizations history paper as 100% AI-generated.
Orion has autism and had paid extra just so he could enroll in Adelphi’s specialist support program for neurodevelopmental students. He claimed that he wrote the paper himself with a bit of grammatical help from a university tutor.
To substantiate his claims, he submitted results from two other AI detectors, GPTZero and Grammarly, both of which classified the same essay as human-written. His university, in a now-common trend, refused to consider any of his evidence and denied his appeal.
A New York State Supreme Court judge ruled the university’s finding was “without valid basis and devoid of reason”, and proceeded to order the school to erase his record. This was a groundbreaking ruling, but it was a bit bittersweet for the family who went through all the legal trouble and had to spend north of $100,000 in legal fees to get Orion’s record expunged.
Then there’s the case that happened at the University of Michigan. A student who decided to stay anonymous filed a federal lawsuit in February 2026 with a similar case. She claimed that she was falsely accused of using AI across three separate papers in a single semester by the same instructor.
She also came forward with evidence proving her innocence, plus her documented disability, which included generalized anxiety disorder and OCD, both of which affect how she writes. Her disabilities, according to the complaint, produce “formal tone, meticulous structure, and stylistic consistency”— traits which come across as suspicious and “too perfect”. She got disciplinary probation from the university, plus a grade of “no record” on her transcript.
These are only three examples from a growing number of similar cases, and courts are increasingly skeptical of disciplinary decisions that rest on a detection score and nothing else.
How accurate are these tools, exactly?
Administrators always seem to conveniently forget to read the small print that makes it clear that the tools are not accurate enough to bear the weight being put on them.
Weber-Wulff and colleagues conducted a study where they tested 14 AI detection tools. They discovered that none of them exceeded 80% accuracy, and only five got as high as 70%. Please take note that they tested the mainstream tools that everyone uses.
Turnitin, which is in almost every academic setting, acknowledges a score variance of plus or minus 15 percentage points. This means that an AI detection result that says a piece of writing is 50% AI-generated could actually, realistically, be anywhere from 35% to 65%. In some cases, different tools can produce wildly different results. Orion in the Adelphi case got a 100% score, but two other mainstream tools had totally different results, so is this really a technology instructors can trust?
Then there’s the issue of bias. Non-native English speakers are flagged at rates up to 30% higher than native speakers on some tools. Research published in 2025 found false positive rates exceeding 20% specifically for non-native English writers and creative writing samples. This is not shocking because non-native speakers mostly rely on standardized vocabulary and grammar templates, and AI is trained on similar data.
Turnitin markets a false positive rate of under 1%. GPTZero has released benchmark figures with a recall above 99%. But those numbers come from controlled datasets, not from the real world of student writing — second-language essays, disability-affected writing styles, and careful human revision that happens to produce clean prose.
The Adelphi case made one thing clear that should have already been obvious: when three detectors test the same essay and disagree with each other, the tools are not ready to be the last word on anything.
What is it actually doing to students?
Most people don’t talk about how this affects students. New York-based education consultant Lucie Vagnerova has handled more than 100 AI-related misconduct cases since late 2023. Anxiety, she says, is the most common word she hears from students going through the process. They are not eating or sleeping, and they feel guilty for something they did not do.
Cases stretch on for weeks, sometimes months. International students face visa complications on top of everything else. Some students received misconduct letters after graduating and while already employed because a professor finally got around to running their old papers through an AI detector. The emotional disruption of proving you are human to an algorithm apparently does not end when the semester does.
The University of Michigan’s instructor in the Jane Doe case reportedly posted publicly that grading had made him “paranoid and inclined to see AI everywhere” — and then filed an academic misconduct accusation anyway. That says a lot.
Universities are split on the question
Not every institution is doubling down. The response across higher education is fractured, and that inconsistency is its own problem.
MIT’s official guidance states that AI detection tools should not be used as the sole basis for misconduct charges. The University of Waterloo discontinued Turnitin’s AI detection after the tool flagged human-written text as 100% AI-generated. Some universities, like the University of Cape Town in South Africa, have banned AI detectors entirely. But there are still others who are choosing to ignore the warnings and are expanding their use.
The result is a patchwork where the same essay, submitted to two different universities in the same week, could earn a diploma at one and trigger a disciplinary hearing at the other. The lawsuits are pushing courts to resolve this inconsistency, whether institutions are ready for the conversation or not.
The algorithm does not know you
There is an irony so perfectly shaped it almost feels engineered: institutions using AI tools to catch students using AI tools, then trusting the AI tool so completely that a student with a full revision history and six months of documented work cannot get anyone in the room to look at it.
AI detection may have a legitimate supporting role in academic integrity. But supporting role is the operative phrase. When software with a documented error margin is handed authority to revoke scholarships, mark transcripts, and end careers, the institutions deploying it have not adopted a technology. They have outsourced their judgment.
The students taking these cases to court are not asking universities to stop caring about cheating. They are asking for the right to be heard. The right to present evidence. The right not to be convicted by a tool whose own developers say should never be the sole proof of anything.
That really does not seem like a lot to ask for.
메타데이터
- post_id
- b0ce1a06a0cf
- slug
- guilty-until-proven-human-b0ce1a06a0cf
- url
- https://medium.com/ai-ai-oh/guilty-until-proven-human-b0ce1a06a0cf
- canonical_url
- https://medium.com/ai-ai-oh/guilty-until-proven-human-b0ce1a06a0cf
- author_url
- https://medium.com/@PamelaMasendeke
- status
- ok
- fetched_at
- 2026-06-16 19:09:56