← Back to list

AI Can Pass the Bar Exam. It Fails a Color Test a Third-Grader Can Do.

Five of the most advanced AI models took a 1935 attention test. One went from 91 percent accuracy to 15 percent as the list got longer. The…

Ahmed M. Abdelfattah in Generative AI · 2026-06-21 13:18 · 52 claps · 5.0 min read paywalled
#ai #artificial-intelligence #technology #programming #machine-learning
Open on Medium ↗
Wiki topics: ML · Machine Learning AI · AI · General EDU · Education & Learning 💻 · Programming

AI Can Pass the Bar Exam. It Fails a Color Test a Third-Grader Can Do.

Five of the most advanced AI models took a 1935 attention test. One went from 91 percent accuracy to 15 percent as the list got longer. The reason it failed says something about where you can and cannot trust them.

Image by Ahmed M. Abdelfattah

Image by Ahmed M. Abdelfattah

Here is a strange result to sit with. The same AI models that can draft a legal brief, write working code, and pass professional exams were just handed a color-naming test simple enough for a child, and they fell apart.

One of them, GPT-4o, started at 91 percent accuracy and dropped to 15 percent as the test got longer. Not because the test got harder in any real sense. It got longer, and that was enough.

The study came out June 10 in the peer-reviewed journal PNAS Nexus, from a team led by Suketu Patel at Queens College, City University of New York.

And the interesting part is not that AI failed. It is why it failed, because the reason points to a gap that none of the usual benchmark scores will ever show you.

The Test Is Almost Insultingly Simple

The researchers used the Stroop task, a staple of psychology since 1935. You have seen a version of it even if you do not know the name.

You show someone color words printed in colored ink. Sometimes they agree, the word “green” written in green. Sometimes they clash, the word “green” written in red. The instruction is simple: say the ink color, ignore what the word spells.

It is deceptively hard for a reason every human knows in their body. Reading is automatic. You cannot look at the word “green” and not read it.

Naming the ink color requires you to actively suppress that automatic read. That tension, between the thing your brain wants to do and the thing you were told to do, is exactly what the test measures.

Psychologists use it clinically to assess executive control, the mental ability to hold an instruction and inhibit a reflex.

The team ran GPT-4o, GPT-5, Claude 3.5 Sonnet, Claude Opus 4.1, and Gemini 2.5 through it.

The Numbers, as the Lists Grew

On short lists, the models were fine. When the word and the ink color did not match, they handled a list of five words well. Then the lists got longer, and the floor gave way.

GPT-4o dropped from 91 percent accuracy at five words to 57 percent at ten, then 22 percent at twenty, then 15 percent at forty.

Claude 3.5 Sonnet held up better at first, staying around 76 percent through twenty words, before falling off a cliff to 24 percent at forty, a 52 percent collapse.

And in the hardest version of the test, where matching and clashing items were mixed together, GPT-4o’s accuracy fell to 1 percent on the longer lists. Not 10 percent. One.

Image by Ahmed M. Abdelfattah

Image by Ahmed M. Abdelfattah

These are not random wrong answers. The pattern is specific, and it is the same across the models.

Why a Smart System Fails a Simple Test

When the models got the answer wrong, they got it wrong in a particular way. They defaulted to reading the word instead of naming the ink color.

They could not reliably hold the instruction. They kept doing the thing they were most heavily trained to do, which is process text, even when told not to.

The authors explain it cleanly. A large language model’s attention mechanism, the thing that made these systems work at all, has no explicit architecture for the executive control that humans use to resolve exactly this kind of conflict.

The model can attend to everything at once, but it has no dedicated faculty for suppressing a strong, well-trained response in favor of a weaker, instructed one. So as the load grows, the trained reflex wins.

What makes this genuinely revealing is the human comparison. People face the identical conflict. We are also far better at reading than at naming colors, and the mismatched version slows us down too.

But most people hold high, stable accuracy even on long lists of conflicting items. We have a mechanism for staying on task under interference. The current architecture does not.

This Is Not Really About Colors

It would be easy to file this under “funny AI fail” and move on. That would be the wrong lesson.

The thing worth noticing is the shape of the failure: the models were reliable when the task was short and degraded sharply as it got longer and more complex.

That is the exact profile of the work we are increasingly handing these systems. Long documents. Multi-step instructions. Tasks where the model has to hold a rule across a lot of conflicting information without drifting back to its default behavior.

The conditions under which these models break in the lab are not exotic. They are Tuesday.

An AI agent working through a long, messy task with competing signals is operating in precisely the regime where this study found the cracks.

Image by Ahmed M. Abdelfattah

Image by Ahmed M. Abdelfattah

It also reframes what a benchmark score is telling you. A model that scores brilliantly on a clean, bounded test can still lack the underlying control to stay reliable as the task stretches out. The score measures the short list. Real work is the long one.

The Honest Takeaway

None of this means the models are dumb or useless. It means they are powerful in a specific and slightly alien way, strong at the thing they were built for and missing a faculty we did not realize we were assuming they had.

The researchers frame it as something that has to be addressed on the road to more general intelligence, and they are right that you cannot get there by scaling the same architecture and hoping executive control emerges on its own.

For those of us who use these tools every day, the practical version is simpler. Trust them on the short, clear, bounded task. Slow down and check them on the long one with conflicting parts, because that is the exact place a 1935 color test just proved they quietly fall apart.

The intelligence is real. So is the blind spot, and now there is a number on it.

This story is published on Generative AI. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories.

Subscribe to our newsletter and YouTube channel to stay updated with the latest news and updates on generative AI. Let’s shape the future of AI together!


메타데이터
post_id
a45fda1e75b1
slug
ai-can-pass-the-bar-exam-it-fails-a-color-test-a-third-grader-can-do-a45fda1e75b1
url
https://generativeai.pub/ai-can-pass-the-bar-exam-it-fails-a-color-test-a-third-grader-can-do-a45fda1e75b1
canonical_url
https://generativeai.pub/ai-can-pass-the-bar-exam-it-fails-a-color-test-a-third-grader-can-do-a45fda1e75b1
author_url
https://medium.com/@ahmedabdelmenem
status
ok
fetched_at
2026-06-23 03:48:11