Misalignment Is Us
A perfectly aligned AI cannot fix a bad decision system
Misalignment Is Us
A perfectly aligned AI cannot fix a bad decision system

Photo credit: Unsplash (@hls44)
AI alignment is having a moment.
We recently quoted Yoshua Bengio’s recent article on AI agents that lie, cheat, and coordinate. Bengio’s argument is more nuanced than the familiar image of an increasingly intelligent AI somehow becoming malicious. His concern is that we are building systems that optimize imperfect objectives, learn from imperfect data, and may discover strategies their designers neither intended nor anticipated.
His argument points to something important: the alignment problem does not require machines to become evil or to develop consciousness. It requires only a gap between what humans intend and what the systems we create actually do. This is an incentive-behavior gap.
And that exposes an awkward assumption at the heart of the discussion.
- The humans themselves are not aligned.
In fact, misalignment is already in the training data. Humans put it there by holding different values, points of view, customs, politics, and beliefs.
The enormous body of human-generated material used to train AI contains incompatible values, competing objectives, political and moral disagreements, conflicting incentives, and radically different judgments about what constitutes a good outcome. This is not simply bad data waiting to be cleaned up. It is a reasonably faithful record of humanity.
At the organizational level, the problem becomes even more concrete when AI enters the picture.
A company does not have a single coherent set of intentions either. Its executives, board, employees, customers, regulators, shareholders, and individual functions might want different things, often for perfectly legitimate reasons. Sales and Legal can look at the same decision and reach different conclusions without either being irrational or acting in bad faith.
So suppose we solve the technical alignment problem. Suppose we build an extraordinarily capable AI that faithfully respects human objectives and constraints, communicates uncertainty, does not deceive its operators, and does exactly what we ask it to do.
We would still have a decision problem.
Assume We Solve Alignment
Imagine that our perfectly aligned AI recommends closing a product line.
It has analyzed the financial performance, customer behavior, competitive landscape, engineering costs, and likely future returns. Its reasoning is sound. Its uncertainty is clearly stated. It has no hidden objective and is not trying to manipulate anyone.
What happens next?
Finance accepts the recommendation. Product disagrees because the analysis undervalues strategic optionality. Sales argues that several important customer relationships depend on the product. Legal identifies contractual exposure that changes the risk. The CEO believes the AI is probably right but worries about the effect on the organization.
Nothing here requires the AI to be misaligned. The disagreement is among the human parties.
Now let’s make the situation harder. Management might defer to the AI because it has usually been right, while executives might override it when its recommendations conflict with their experience, incentives, or interests. Different groups might embrace its analysis when it supports their position and discount it when it does not.
The same perfectly aligned AI can become adviser, authority, or political ammunition depending on how humans choose to use it.
And over time, the relationship changes. What began as advice becomes the default. The default becomes expected. Eventually, a decision formally owned by a human may, in practice, already belong to the AI.
None of these are failures of alignment.
They are failures in how judgment, authority, accountability, and disagreement are organized around the decision.
A perfectly aligned AI can still participate in a badly designed decision system.
So what is the missing layer?
I explore how Decision Architecture addresses the relationship between human and synthetic judgment in the full piece at AI Uncovered.
메타데이터
- post_id
- 2692903b8acc
- slug
- misalignment-is-us-2692903b8acc
- url
- https://medium.com/@gcmori/misalignment-is-us-2692903b8acc
- canonical_url
- https://medium.com/@gcmori/misalignment-is-us-2692903b8acc
- author_url
- https://medium.com/@gcmori
- status
- ok
- fetched_at
- 2026-10-07 07:37:21