The interface is lying to you
AI tools are designed to feel effortless. That is precisely the problem.
The interface is lying to you
AI tools are designed to feel effortless. That is precisely the problem.
Photo by pine watt on Unsplash
In May 2023, a New York lawyer filed a brief citing six judicial decisions that did not exist. ChatGPT had produced them, complete with plausible case names and citations. The lawyer, asked later why he had not checked, replied that he had: he had asked ChatGPT whether the cases were real, and the model had assured him they were.
The episode is now folklore, but the structural lesson has barely been absorbed. The failure was not the model’s. It was the interface’s — a blank conversational box that invited a senior professional to treat a probabilistic text generator as a research database, then offered no friction when he did.
As large language models move from novelty to professional infrastructure, this design choice is becoming the central problem in AI use. The chat box is built for adoption, not for judgment. And the gap between the two is where most of the value, and most of the risk, now sits.
The cost of frictionlessness
The dominant framing of AI competence remains “prompt engineering” — the idea that better phrasing yields better answers. It is a useful skill, and an incomplete one. It locates the work entirely on the input side, as if the model were a vending machine that simply needed the right code.
A more accurate analogy is agricultural. The prompt is a seed; the context provided is soil; the output is a first growth, not a finished product. Some responses need nurturing through added precision. Others need pruning. Many need both, across several iterations, before anything usable emerges.
Current interfaces obscure this. A single text box, a single response, a thumbs-up icon. The visual grammar suggests transaction, not cultivation. Users are nudged toward accepting first drafts as final answers — and toward refining prompts when they should be interrogating outputs.
Framing and evaluation
Two distinct skills determine whether AI use produces reliable work, and current tools support neither.
The first is framing — the work done before the prompt. A question about contractual liability, a question about market positioning, and a question about code architecture each require a different frame: different assumptions about what counts as a valid answer, different boundaries on the problem, different definitions of relevance. Domain expertise is what supplies the frame. Without it, even a well-phrased prompt produces an answer that looks correct and is structurally wrong.
The second is evaluation — the work done after the response. This is more than reading critically. It means testing whether an argument actually coheres, identifying claims that the model has asserted but not supported, and stress-testing conclusions against what the user already knows. It is also the step most users skip, because the interface does not ask for it.
The shape of the problem is familiar to anyone who has worked in operations. Quality management has long been defined by the PDCA cycle — Plan, Do, Check, Act — codified by W. Edwards Deming and embedded in every serious manufacturing standard. The premise is that quality is not produced by the first two stages but by the last two: checking the output against intent, and adjusting the process accordingly. A team that only plans and does is not running a quality process. It is running execution and calling it quality.
AI use has the same structure and the same failure mode. The full loop has five operations — intent, prompt, response, evaluation, refinement — which map onto PDCA’s four stages with intent and prompt sitting inside Plan. Most chat interfaces visualise only three: prompt, response, and an implicit follow-up. Intent is collapsed into the prompt box. Evaluation and refinement are the user’s responsibility, exercised outside the interface, against no structure. The two stages that create the value are invisible, and most users stop where the interface stops.
The adoption paradox
The simplicity of these tools is not accidental. OpenAI, Anthropic, and Google are competing for users, and friction is the enemy of adoption. A product that asked users to declare their intent, specify their frame, and rate the coherence of each response would be more rigorous and substantially less popular.
But the simplicity has a cost that is now showing up in professional settings: hallucinated citations in legal filings, fabricated figures in analyst notes, confident summaries that miss the point of the underlying document. The error rate is not the issue — humans err too. The issue is that the interface gives users no signal about when to doubt and no structure for doing so. The model’s confidence is uniform; the user’s vigilance has to supply the variance.
This is the paradox at the centre of the current generation of AI tools. The systems that appear to reduce cognitive effort in fact demand more of it from anyone using them seriously. The effort has simply been moved — from composing the question to auditing the answer — and the interface pretends it has not.
What literacy actually means
A new literacy is emerging in workplaces that take this seriously, and it has little to do with prompt tricks. It is closer to a research discipline: define the intent, frame the problem, generate, evaluate against known sources or first principles, refine. Treat the model as a junior analyst whose work is fast and cheap but unverified — useful, but never the final word.
This is not a skill the tools currently teach. It has to be imported from elsewhere — from law, from journalism, from any field where the cost of an unverified claim is high enough to enforce the habit.
The pressure point in AI is no longer model capability. It is whether the people using these systems have the discipline the interface declines to require. In a market where generating answers has become trivially cheap, the ability to question them is the scarce input. The firms that recognise this will treat AI literacy as a hiring criterion, not a training module. The ones that don’t will keep filing briefs full of cases that were never decided.
메타데이터
- post_id
- c11a6a2bd1d0
- slug
- the-interface-is-lying-to-you-c11a6a2bd1d0
- url
- https://medium.com/@felicienf/the-interface-is-lying-to-you-c11a6a2bd1d0
- canonical_url
- https://medium.com/@felicienf/the-interface-is-lying-to-you-c11a6a2bd1d0
- author_url
- https://medium.com/@felicienf
- status
- ok
- fetched_at
- 2026-07-13 08:05:13