← Back to list

Both Sides Are Wrong About AI Coding. At Least For Now…

Over the last two years, I interviewed engineering leaders from companies like Google, Uber, SpaceX, Tesla, Microsoft, Zendesk, and a lot…

Matthieu McClintock in ITNEXT · 2026-07-10 02:21 · 11 claps · 5.9 min read
#ai-coding #claude-code #github #cursor #software-development
Open on Medium ↗
Wiki topics: LLM · Large Language Models 💻 · Programming 🔓 · Open Source 🔭 · Astronomy & Space 🧘 · Spirituality

Both Sides Are Wrong About AI Coding. At Least For Now…

The AI hype machine

The AI hype machine

Over the last two years, I interviewed engineering leaders from companies like Google, Uber, SpaceX, Tesla, Microsoft, Zendesk, and a lot of other places where software delivery is not taken lightly.

Over the last year, I built a product for those same leaders to help answer the question they kept circling:

Is AI coding actually improving software delivery?

Not demos. Not vibes. Not adoption. Not whether developers like the tools.

Actual delivery.

I was not watching this from the sidelines.

I was producing and hosting a podcast for engineering leaders. I was asking them how they were measuring AI coding ROI. Then I started building ChaosMonkey with a small group of engineering leaders giving constant feedback on whether the thing I was building would actually be useful, trustworthy, or worth buying.

Then it got weirder.

I started coding daily again for the first time in almost a decade, using the same AI coding tools ChaosMonkey was built to evaluate.

That experience left me with one conclusion:

Both sides in the AI coding debate are wrong.

At least for now.

Not because AI generated code is all slop. It is not all slop.

Not because AI coding is magic. It is definitely not a silver bullet.

But neither side has proven what engineering leaders actually need to know.

Is AI coding making the software delivery system better?

The question that kept coming back

When I started interviewing engineering leaders, I was not building ChaosMonkey yet.

I was listening.

And over time, the same unresolved problem kept showing up in different forms.

Engineering leaders knew their teams were using AI coding tools. That was not the interesting part. The harder question was whether anyone could connect AI coding behavior to downstream engineering outcomes.

Were review cycles improving? Were reopens going down? Were failures going down? Were customers getting value faster? Were expensive workflows actually better, or just more expensive? Were junior devs getting leverage, or were senior devs quietly becoming cleanup crews?

The answers were usually thoughtful.

But they were incomplete answers.

Most teams had signals. They had instincts. They had anecdotes. They had developers who loved the tools and developers who wanted to throw the tools into the ocean.

What they did not have was a clean way to connect the work happening inside AI-assisted coding workflows to what happened later in the SDLC.

I kept asking engineering leaders whether they had a real way to understand AI coding ROI, and the answer was usually some version of:

“Not yet.”

By June 2025, I started asking a smaller group of engineering leaders a more direct version of the question.

Would you use something that helped answer this? Would you trust it? Would you buy it? What would it need to show? What would make it actionable and compelling?

That group became the early feedback loop for the product.

They validated the problem before I built it. Then they kept giving feedback as the product changed, broke, improved, confused me, and occasionally made me question every decision I had ever made.

The product was entirely driven by the eventual customers who would end up using it.

Which, in my opion, is generally how software should be built.

Then the product became the lab

The strange part is that I was building a product to measure AI coding while using AI coding tools to build the product.

I used the major code editors and IDEs. I used most of the frontier coding models. I used cheap models. I used expensive models. Some were impressive. Some were overrated. Some felt like magic for thirty minutes and then quietly set a small fire in another part of the codebase.

And while all of that was happening, ChaosMonkey was collecting telemetry on my own work.

Active coding time. AI usage. Idle time. Files touched. Commits. Time to first commit. Workflow patterns. Delivery outcomes. All in the context of time, tools, repos, and the actual work moving through the system.

The product was not just something I was building.

It was something I was becoming data inside.

I could watch my own coding behavior change over time. I could see how different workflows affected the way I built. I could see where I moved faster, where I created more cleanup, and where I was letting a very confident machine produce more work for future me.

That experience changed how I think about the entire AI coding debate.

Vibe coding is not software engineering

When I first used ChatGPT in 2022, I was skeptical.

At the time, I did not think AI was a particularly good writer. I also did not think it was a particularly good software engineer. I think that was mostly right.

But 2024 and 2025 changed the conversation. The frontier coding models got better. The IDEs got better. The tools stopped feeling like autocomplete and started feeling like something stranger.

A very powerful assistant that can accelerate a skilled person, confuse an inexperienced person, and quietly wreck a codebase if nobody in the room understands what production-grade software actually requires.

That distinction matters.

Vibe coding and AI-assisted software engineering are not the same thing.

A non-technical founder producing something with a prompt is not the same thing as an engineer shipping maintainable, scalable, secure software into production.

A junior developer using AI to move faster by continuously clicking “Keep All” is not the same thing as a junior developer understanding the architecture, tradeoffs, failure modes, and long-term maintenance burden of what they just accepted.

This should not be controversial.

Somehow it became controversial.

Code generation is not software delivery

This is where I think the debate breaks.

AI made code generation easier.

It did not automatically make software delivery easier.

Those are not the same thing.

Software delivery is not just writing code. It is figuring out what should exist, understanding the problem, choosing the right level of complexity, designing something maintainable, reviewing changes, testing, deploying, operating, debugging, and handling the second-order effects of yesterday’s very impressive shortcut.

A developer can feel faster.

That does not mean the team is faster.

A team can generate more code.

That does not mean customers are getting better software.

This is the difference between activity and outcome.

AI makes activity very easy to manufacture.

That is why measurement matters more now, not less.

I recently joined Mateo Bervejillo on The Future of the Future podcast, and his framing was spot on: is AI making dev teams more productive, or is it creating new bottlenecks?

My answer is that most teams do not know yet.

“We do not know yet” is probably the honest answer for a lot of engineering organizations.

The tools are changing. The models are changing. The workflows are changing. The costs are changing. The way developers use these systems is changing. And most engineering measurement systems were built before AI coding meaningfully changed the workflow.

Why this became more than a dashboard

At first, ChaosMonkey was about visibility.

Show engineering leaders what was happening. Show them how AI coding was being used. Show them the relationship between coding behavior and delivery outcomes.

That was necessary.

It was not enough.

Engineering leaders do not need another dashboard that creates homework.

They need interpretation.

What changed? Why might it matter? How confident should we be? Is this workflow improving delivery, or creating downstream drag? Is this expensive model producing better outcomes, or just better feelings? Where should a leader look next?

That is why the product had to move toward a recommendation engine.

Not because “recommendation engine” sounds good in a pitch deck.

But because raw telemetry is not the same thing as an insight.

A metric can tell you what happened.

A system has to help you decide what to do about it.

Where I landed

I do not think the hype train has proven that every developer is 10x.

I also do not think the skeptics have proven that AI coding is slop.

I think both sides are arguing from incomplete evidence.

That is what I built ChaosMonkey to help solve.

Not to surveil developers.

Not to rank engineers with some fake productivity score.

Not to declare one model, editor, or workflow the winner.

The goal is to help engineering leaders connect AI coding behavior to downstream delivery outcomes, regardless of what AI coding tools are all the rage at any given moment.

Both sides can keep arguing about whether AI coding is magic or slop.

I am more interested in the boring operating question:

Are we using AI in a way that helps us ship better software faster, considering the cost?

If you do not know the answer yet, you are probably not alone.

That’s sort of the whole problem.

Here’s the full episode from my appearance on The Future of the Future podcast.

Watch on YouTube

Listen on Spotify

Listen on Apple Podcasts


메타데이터
post_id
630abbe153df
slug
both-sides-are-wrong-about-ai-coding-at-least-for-now-630abbe153df
url
https://itnext.io/both-sides-are-wrong-about-ai-coding-at-least-for-now-630abbe153df
canonical_url
https://itnext.io/both-sides-are-wrong-about-ai-coding-at-least-for-now-630abbe153df
author_url
https://medium.com/@mattrmclaren
status
ok
fetched_at
2026-07-13 06:23:13