What Level of AI Are You Actually At? A Practical Guide for Organizations Ready to Stop Pretending
Every company says it’s AI-forward. Almost none of them can pass the hard tests.
What Level of AI Are You Actually At? A Practical Guide for Organizations Ready to Stop Pretending
Every company says it’s AI-forward. Almost none of them can pass the hard tests.
There is a question hiding inside every AI strategy deck, every “AI-first” all-hands, every press release announcing a new partnership with a foundation model vendor. The question is not “Are we using AI?” Almost everyone is, in some form. The question is: what level of autonomy has the organization actually achieved?
That sharper framing comes from Annie Kadavy (@annimaniac), whose thread on X in May 2026 cut through the noise with unusual precision. Drawing on visits to companies ranging from eight-person startups to Ramp, a 1,500-person organization, she observed something that anyone leading engineering or product teams will recognize immediately: “AI can no longer be a personality trait of the founding team. It has to become part of the company’s DNA.”
Her insight was to borrow from a different domain — autonomous vehicles. For years, the AV industry used a levels framework (L0 through L5) to force precision in a field full of marketing language. Cruise control was not autonomy. Lane-keeping was not autonomy. The levels mattered because they ended the hand-waving.
The same discipline is overdue for organizations.
This article builds on that framework, layers in research from Credo AI, Bain/HBR, JPMorgan Chase, and Gartner, and adds practical guidance for leaders who want to stop measuring AI maturity in tools purchased and start measuring it in organizational capability earned.
Why the Binary Question Is the Wrong Question
“Is your company AI-pilled?” is currently being treated as a yes/no question. It is not. According to Kadavy’s framework, companies differ along two independent axes: intensity (how deeply AI is embedded in daily work across the organization) and technical capability (what AI is actually allowed to see, do, and change).
A company where employees use ChatGPT to summarize meetings is not in the same category as a company where agents can query systems of record, take bounded action, propagate workflows across teams, and improve the way future work gets done. Both may describe themselves as AI-forward. They are not operating at the same level.
Credo AI’s 2026 Enterprise AI Governance Maturity Model reinforces this. Their research found that fewer than one in five organizations have fully operationalized their AI practices. The majority sit in what they call a “governance implementation gap”: formal policies on paper, but shadow AI rampant, responsibilities unclear, and every new initiative feeling like a bespoke fire drill.
The four diagnostic questions Kadavy proposes are the cleanest instrument I have seen for cutting through this:
- What can AI see? Is the work of your company legible to a machine, or does it live in someone’s mind, undocumented meetings, and SaaS tools the AI can’t read?
- What can AI do? Can it act on systems of record — open PRs, update CRMs, reconcile invoices — or can it only summarize what humans already wrote down?
- Who can extend the system? Are non-engineers shipping production internal tools, or is every workflow held together by a few power users whose knowledge walks out the door when they leave?
- How has the organization changed? Or are you running a 2023 org chart with better autocomplete?
The answers to these four questions cluster into six levels. Let us walk through each one honestly.
L0: AI as Theater
What can AI see? Nothing structured. Knowledge lives in people’s heads, undocumented meetings, and SaaS tools AI cannot read.
What can AI do? Nothing of consequence. Maybe summarize a meeting if a human pastes the transcript.
Who can extend the system? No one. AI is a personal tool, ungoverned, unintegrated.
How has the org changed? It hasn’t. Same chart, same hiring plan, same handoffs, same dependence on managers as information routers.
Hard test: Can AI complete any recurring business process end-to-end?
Common false positive: A CEO who gives an excellent speech about AI transformation while still running the company through the same executive staff meetings, status updates, reporting lines, and headcount plans. Announcements are not adoption.
L0 is more common than leaders want to admit. It often looks like progress because there is visible activity: a Head of AI has been hired, licenses have been purchased, a Slack channel called #ai-tools exists and has 200 members. But if you apply the hard test — can AI complete any recurring business process without a human in the middle? — the answer is no.
L1: Personal Productivity
What can AI see? Each individual’s personal AI sees only what that person feeds it. Saved prompts, scratch files, private knowledge bases. No org-level visibility.
What can AI do? Help individuals draft, summarize, brainstorm, code. No action on systems of record.
Who can extend the system? Each user reinvents independently. Power users are heroes; their workflows leave with them.
How has the org changed? It hasn’t. Same chart. Maybe a Head of AI hire with budgetary influence and some purchased products.
Hard test: If your best AI user left tomorrow, would their workflow remain in the company?
Common false positive: “80% of employees use AI weekly.” Probably true. Also meaningless as a measure of organizational capability.
The Bain and OpenAI research published in Harvard Business Review (April 2026) put a sharper point on this: generative AI has sprinted from novelty to boardroom priority, “but despite widespread adoption, not every company is realizing bottom-line improvement commensurate with its capabilities.” L1 organizations have adoption metrics. They do not have impact metrics.
Nirmaljeet Malhotra, Executive Director of Product Management at JPMorgan Chase, described this transition problem in a February 2026 Products That Count webinar: “AI pilots optimize for technical feasibility; scaled AI requires operational accountability.” At L1, the organization has feasibility. It does not yet have accountability structures.
L2: Team Workflow
What can AI see? Teams have shared context: a team-specific system prompt, shared prompts, function-specific integrations. AI sees within team boundaries.
What can AI do? Functional workflows. AI for sales prospecting, support triage, engineering code review. Bounded actions within a team’s domain.
Who can extend the system? Within a team, non-engineers can tap shared workflows. Across teams: each function rebuilds the same thing privately.
How has the org changed? Functional efficiency within roles. A customer success manager with AI handles 200 accounts versus 50. Hiring slows, but org shape is unchanged. Role boundaries are intact.
Hard test: Does this workflow cross team boundaries, or is every function building its own private AI stack?
Common false positive: “We have AI workflows in every department.” But the workflows do not connect, so the company is a collection of AI-enhanced silos rather than an AI-native organization.
L2 is where most organizations with serious AI investment currently sit in 2026. The gains are real: productivity improvements within functions are measurable. But the structure is fragile. Each team has built its own stack, its own prompts, its own integrations. There is no horizontal transfer of capability. When the team’s AI power user leaves, the capability leaves with them — exactly the same problem as L1, just at the team level rather than the individual level.
L3: Organizational Infrastructure
What can AI see? The whole organization is queryable. Cross-functional context is accessible. Core systems of record are exposed via APIs, CLI tools, or well-defined integration layers. Agents can act on these systems, not just observe them.
What can AI do? Agents act across systems: update CRMs, open PRs, route tickets, run analyses, draft customer communications, reconcile invoices. Cross-functional, but still bounded by policy.
Who can extend the system? Non-engineers do not just consume shared workflows; they author them. A sales rep packages a call analysis pattern as a shareable skill. A customer experience engineer packages a ticket investigation workflow. Skills move horizontally across functions.
How has the org changed? The org chart looks materially different from a 2023 equivalent. The shape varies: some organizations have eliminated the traditional PM role; others have restructured it into an agent-orchestration function. The unifying signal is that the company has made an explicit structural choice about how AI changes who does what, and that choice is visible in reporting lines and job descriptions.
Hard test: Can an agent answer, across systems, what shipped last sprint, who asked for it, what broke after launch, what customers said, and what the company should do next — without convening a cross-functional meeting?
Common false positive: A large collection of meeting transcripts and dashboards that nobody synthesizes. Capture is not legibility. An inert archive is not an operating system.
L3 is the first level where AI has genuinely changed the organization. It is also the level where most enterprise transformation efforts stall. The reason is not technical: it is about systems legibility. Before agents can act on systems of record, those systems need to be readable by machines. That requires documented processes, clean data structures, API-accessible tools, and a culture that treats knowledge as organizational property rather than individual currency.
Credo AI’s framework describes this transition as moving from “Formalizing” to “Optimizing”: “You have moved beyond spreadsheets to a unified control plane that makes governance a tailwind for deployment.” Most enterprises arrive at L3 having solved the technology problem but not the institutional problem. The systems are there; the discipline to maintain them is not.
Gartner forecasts that by 2028, 33% of enterprise software applications will include agentic AI capabilities, up from less than 1% in 2024. Organizations that reach L3 before that inflection point will compound their advantage significantly.
L4: Compounding Operating System
What can AI see? Not just what happens, but the relationships between what happens. The system maintains its own context: agents update other agents, skills marketplaces propagate wins and remove duplicated effort, the system learns what to surface. Capture, synthesis, and query are continuous and automated.
What can AI do? Agents have policy-driven decision authority within scoped domains. Security agents detect an anomaly, validate it, fix it, and open a pull request with human review required only at the merge step. Custom internal tools are purpose-built for the work the company does most often.
Who can extend the system? Non-engineers ship production internal tools without filing a ticket or waiting for engineering bandwidth. A finance team member builds an automated contract reviewer. A sales AE ships a prospecting tool in under an hour. They find their own pain, prototype a fix, and pull engineering in only when it is time for production hardening.
How has the org changed? Hierarchy collapses toward what Kadavy calls “channel managers” of agent workflows. New archetypes emerge. Compensation and promotion decisions are explicitly tied to AI proficiency. The time from customer signal to shipped feature is measured in hours, not sprints.
Hard test: Show a workflow that got better because the system learned from prior runs, not because one heroic person manually improved it. Then show three production tools shipped by non-engineers in the last quarter.
Common false positive: Agent sprawl. A hundred brittle automations do not equal a compounding operating system. L4 requires managed compounding: lifecycle management, observability, and evaluation frameworks. Without what Kadavy calls “compaction discipline,” the factory clogs.
Malhotra’s framing from JPMorgan is useful here: “Scaling AI is about platformization, not duplicating isolated use cases. Reusable infrastructure — data pipelines, monitoring, governance layers — compounds ROI.” L4 organizations have not just built workflows; they have built the infrastructure that makes workflows cheaper and more reliable to build over time.
The distinction between L3 and L4 is the difference between an organization that uses AI and an organization that learns through AI. At L3, humans direct the system to improve. At L4, the system improves because it is designed to.
L5: Virtually Self-Driving
This level does not fully exist yet. Kadavy is explicit about this, and intellectual honesty requires preserving that caveat. What follows is a description of what it might look like, not a characterization of where any organization currently sits.
An L5 organization is one where the core operating loops can sense reality, diagnose issues, initiate work, execute within delegated authority, update shared memory, and improve future behavior — with humans governing strategy, taste, risk, values, and exceptions rather than running the loops themselves.
Six markers characterize the L5 state:
- The system notices something important without being asked.
- The system synthesizes across multiple sources of context.
- The system decides whether action is warranted.
- The system acts within delegated authority.
- The system escalates when uncertainty or consequence exceeds its authority.
- The system updates shared memory so future behavior improves.
Hard test: What important thing did the company notice, decide, act on, and learn from recently without a human initiating the process? Not a threshold alert. Not a configured automation. Something the system synthesized that humans had not yet framed as a question.
Common false positive: The “fake autonomy” pattern. The company claims self-driving behavior, but the system is executing preconfigured rules or surfacing threshold-based alerts. Humans are still doing all the noticing. Real generative behavior is distinguished from glorified observability: this is the open technical and organizational challenge at this level.
The L5 vision is what Steve Blank’s maxim points toward when applied to AI: a startup is not a small version of a large company, and an AI-native organization is not simply an AI-assisted version of an old company. It is an organization rebuilt around a new operating model. We are still learning what that looks like at scale.
Why Your Answers Are Asymmetric (and What That Tells You)
One of the most practically useful observations in Kadavy’s framework is this: organizations rarely answer all four questions at the same level. The asymmetry is diagnostic.
Some common patterns:
High visibility, low action: AI can see a lot but cannot do much. This is the data lake problem applied to AI: the organization has invested in knowledge capture and synthesis but has not yet built the integration layers that allow agents to act on what they know. The intervention is API and systems access, not more capture.
High action, low extensibility: AI can do a lot, but only engineers can extend the system. This is the technical moat problem: the organization has built powerful workflows but has not democratized authorship. The intervention is tooling and training that enables non-engineers to build, combined with a governance model that makes that safe.
Changed org chart, thin substrate: The org chart looks different, but the underlying systems have not changed. This is the most dangerous false positive: structural reorganization that precedes the infrastructure it depends on. The intervention is slowing the structural changes until the substrate can support them.
Good substrate, no behavioral change: The systems are legible, the APIs are there, the governance framework exists — but the organization has not made an explicit choice about who does what differently. The intervention is leadership: a clear, visible commitment to which roles change and how.
A Practical Roadmap for Moving Forward
From L0 to L1: Make the work legible
The prerequisite for any AI integration is that your organization’s work is readable by a machine. That means documented processes, written decisions, searchable knowledge bases, and a cultural norm that treats institutional knowledge as a shared asset rather than personal leverage.
Practical first steps: audit one recurring business process end-to-end and document every step, decision point, and handoff. Then ask: if a new hire needed to complete this process with no guidance, is everything they need written down somewhere a machine could find?
From L1 to L2: Build shared context at the team level
Individual workflows that live in one person’s head are fragile. The transition to L2 requires making team-level context explicit and shared: team-specific system prompts, shared prompt libraries, documented integration patterns.
Practical first steps: identify your two or three highest-volume recurring team workflows and document them as reusable patterns. Put them somewhere every team member can find and extend them. Measure whether new team members can execute the workflow without asking the power user.
From L2 to L3: Expose systems of record to agents
This is the hardest transition and the one where most enterprise transformation stalls. It requires: clean, API-accessible systems of record; governance frameworks that specify what agents are and are not permitted to do; and a skills architecture that allows workflows to move horizontally across functions rather than being rebuilt by each team.
The Credo AI governance maturity model is useful here. They recommend a “unified control plane” — a single place where AI decisions, permissions, and audit trails are visible across the organization. Without this, L3 is ungovernable.
Practical first steps: pick one system of record (your CRM, your project management tool, your support ticketing system) and build one agent integration that allows a cross-functional workflow to execute without a human handoff. Govern it explicitly: document what the agent can and cannot do, how exceptions are handled, and who is accountable when it gets something wrong.
From L3 to L4: Build the compounding infrastructure
L4 requires three things that L3 does not: observability (you can see what agents are doing and why), lifecycle management (you have a process for deprecating, improving, and replacing agent workflows), and a skills economy (non-engineers can author and share workflows without engineering bottlenecks).
The Malhotra framework is precise here: “AI systems designed for augmentation outperform those designed for full automation.” The path to L4 is not removing humans from loops; it is designing loops where humans govern at the right level of abstraction — exceptions, strategy, taste, values — while agents handle execution and learning.
Practical first steps: instrument your existing L3 workflows for observability. Build a lightweight skills marketplace where team members can find, rate, and extend agent workflows built by others. Define explicit criteria for what makes a workflow “production-ready” for non-engineering authors.
Toward L5: Invest in the substrate now
Even though L5 does not fully exist, organizations that are building toward it are making specific investments now: in memory architectures that allow agents to accumulate and retrieve organizational context across sessions; in evaluation frameworks that can distinguish genuine generative behavior from sophisticated rule-following; in governance models that can delegate bounded authority to agents without requiring human approval for every decision.
The companies that will reach L5 first are not the ones spending the most on models. They are the ones building the most legible, well-governed, well-observed substrate for agent behavior to compound on.
The Uncomfortable Answer
The hard tests at each level share a common thread. They do not primarily test AI capability. They test documentation discipline, systems legibility, and institutional knowledge management — problems that most organizations have had for decades and have consistently underinvested in solving.
The reason AI maturity is hard is not that AI is hard. It is that organizational legibility is hard. And AI makes that problem impossible to ignore, because a system that cannot read your organization cannot help it.
This is why Blank’s analogy is right: an AI-native organization is not an AI-assisted version of an old organization. It is built differently from the beginning, with the assumption that institutional knowledge must be machine-readable, that workflows must be composable, and that the people closest to the work must be empowered to author the systems that help them do it.
Those of us leading large organizations that were not built this way have a harder path. But the framework exists. The levels are clear. The interventions are knowable.
The question is whether you are willing to apply the hard test honestly.
Sources and Further Reading
- Annie Kadavy (@annimaniac), original framework thread, X, May 2026: https://x.com/annimaniac/status/2050225284277026990
- Credo AI, “The Six Levels of AI Maturity: Where Does Your Organization Rank?” February 2026: https://www.credo.ai/blog/the-six-levels-of-ai-maturity-where-does-your-organization-rank
- Arjun Dutt, Gene Rapoport, Aaron Chatterji, et al., “How to Move from AI Experimentation to AI Transformation,” Harvard Business Review, April 2026: https://hbr.org/2026/04/how-to-move-from-ai-experimentation-to-ai-transformation
- Nirmaljeet Malhotra, “From AI Pilot to Enterprise Impact: How to Scale AI the Right Way,” Products That Count, February 2026: https://productsthatcount.com/from-ai-pilot-to-enterprise-impact-how-to-scale-ai-the-right-way/
- Gartner, “Gartner Names Agentic AI Top Tech Trend for 2025,” as cited by The Journal, October 2024
- Nemko Digital, “AI Maturity Model 2025”: https://digital.nemko.com/ai-maturity-model
Sakti Bagchi is a Release Train Engineer and Senior Development Manager at Altera Digital Health, where he leads cross-geographic engineering teams across India and the US. He writes about engineering leadership, AI adoption in enterprise, and the inner game of leadership at saktibagchi.in and on Medium @sakti.bagchi.
메타데이터
- post_id
- eed4e648da51
- slug
- what-level-of-ai-are-you-actually-at-a-practical-guide-for-organizations-ready-to-stop-pretending-eed4e648da51
- url
- https://medium.com/@sakti.bagchi/what-level-of-ai-are-you-actually-at-a-practical-guide-for-organizations-ready-to-stop-pretending-eed4e648da51
- canonical_url
- https://medium.com/@sakti.bagchi/what-level-of-ai-are-you-actually-at-a-practical-guide-for-organizations-ready-to-stop-pretending-eed4e648da51
- author_url
- https://medium.com/@sakti.bagchi
- status
- ok
- fetched_at
- 2026-07-26 11:40:32