Why Counting AI Agents Is the Wrong Metric for Measuring AI Maturity
4 min read
The most dangerous place an organization can be in its AI journey is not behind—it is confidently wrong about how far ahead it actually is. Across boardrooms and strategy sessions in 2026, executives are pointing to the number of AI agents their teams have deployed as proof of maturity. It is a seductive metric. It is also deeply misleading. A rigorous AI maturity assessment does not begin with a headcount of tools. It begins with an honest examination of the friction points your organization is experiencing right now.
If we have deployed dozens of AI agents across our business units, doesn't that mean we are further along than most?
Not necessarily. Deployment volume is an activity metric, not an outcome metric. A company with fifty agents experiencing widespread trust failures, inconsistent policy enforcement, and capacity bottlenecks is not more mature than a company with five agents operating with precision, governance, and measurable business impact. What matters is not the quantity of AI in motion—it is the quality of the infrastructure, human and technical, surrounding it. Boris Cherny's five-stage framework for AI adoption reorients this conversation entirely, shifting focus from what you have deployed to what problems you are actually encountering at each stage of that deployment.
Understanding the Boris Cherny AI Stages as a Diagnostic Tool
Cherny's framework is not a ladder you climb by acquiring more technology. It is a diagnostic map that reveals where your organization's real constraints live. Each stage is defined not by capability but by challenge. In the earliest stages, organizations struggle with basic trust—employees do not trust AI outputs enough to act on them without extensive manual verification, which ironically negates much of the efficiency gain. At intermediate stages, the challenge shifts to capacity: the organization's processes and human workflows cannot absorb the volume of AI-generated output fast enough to create value. At more advanced stages, the friction moves into policy enforcement and resource management, where the governance structures needed to operate AI at scale either do not exist or are not being followed consistently.
What makes this framework particularly valuable for senior leaders is that it forces a different kind of self-awareness. Most maturity models ask you to check boxes about what you have implemented. Cherny's model asks you to identify where you are bleeding. That is a fundamentally more honest and more useful diagnostic posture.
How do we know which stage we are actually in, as opposed to where we think we are?
The answer lies in where your teams are experiencing the most recurring frustration. If your people are constantly second-guessing AI outputs and reverting to manual processes, you are in a trust-deficit stage regardless of how sophisticated your tooling is. If your workflows are generating AI-assisted work faster than your review and decision-making processes can handle, you are in a capacity-constraint stage. If you are seeing inconsistent AI behavior across departments because different teams have built their own shadow policies, you are in a governance stage. The symptom is the signal. Your job as a leader is to read those symptoms clearly rather than defaulting to the comfort of deployment statistics.
The Productivity Illusion: What a 2026 Study Reveals About AI Adoption Challenges
Perhaps the most unsettling data point shaping this conversation comes from a 2026 study of software engineers. Eighty-four percent of respondents reported feeling more productive when using AI coding tools. Yet simultaneously, objective measures of their work quality declined. This is not a paradox—it is a pattern. When people feel faster, they often feel better, even when the output tells a different story. This gap between perceived productivity and actual quality is one of the most consequential AI adoption challenges facing organizations today.
For executives, this finding should trigger an immediate review of how you are measuring AI impact. If your primary indicators are self-reported efficiency gains or velocity metrics, you may be measuring confidence rather than competence. The engineers in that study were not being dishonest. They genuinely felt more productive. But feeling productive and being productive are not the same thing, and at scale, that gap compounds into significant quality debt that eventually surfaces as customer complaints, rework cycles, or reputational risk.
What should we be measuring instead of productivity to get an accurate picture of AI implementation effectiveness?
The more revealing metrics sit at the intersection of output quality, decision accuracy, and downstream impact. How often are AI-generated outputs being revised before use? What is the error rate on AI-assisted decisions compared to human-only decisions in the same category? Are customer satisfaction scores or operational error rates trending in alignment with your AI investment? These are the questions that separate organizations with genuine AI maturity from those with a well-funded illusion of it. Improving AI implementation requires measuring what actually changed in the world, not just how your teams feel about the tools they are using.
A Self-Assessment Prompt That Drives Honest Organizational Reflection
One of the most practical tools a leader can deploy right now is a structured self-assessment prompt designed to surface the real state of AI adoption within their organization. The prompt is not a survey about tool satisfaction. It is a set of probing questions directed at the friction your teams are actually experiencing. Ask your department heads to describe the last time an AI output created a downstream problem. Ask your operations leads where AI-generated work requires the most human intervention. Ask your compliance team whether AI usage policies are being followed consistently or whether individual teams have developed their own workarounds.
The answers to those questions will tell you more about your true AI maturity stage than any vendor dashboard or deployment report. They will also reveal whether the obstacles your organization faces are primarily technical, cultural, or structural—a distinction that fundamentally shapes what your next investment should be. If the barriers are cultural, more technology will not solve them. If they are structural, governance redesign is the priority. If they are technical, then targeted capability investment makes sense.
How do we turn this self-assessment into an actionable roadmap rather than just a diagnosis?
The bridge between diagnosis and action is specificity. Once you have identified your dominant friction type—trust, capacity, policy enforcement, or resource management—you can build interventions that address the actual constraint rather than the perceived one. Organizations in trust-deficit stages benefit most from human-in-the-loop workflow redesign and AI output transparency initiatives. Those in capacity-constraint stages need process redesign before they need more agents. Those in governance stages need cross-functional policy alignment and clear accountability structures. The self-assessment prompt is not the destination. It is the map that tells you which road to take next on your organizational AI strategy journey.
Building an AI Strategy That Reflects True Organizational Maturity
The organizations that will lead in AI over the next three years are not the ones that deployed the most agents in 2025. They are the ones that built the most honest understanding of where their adoption actually stands and invested accordingly. That requires a willingness to look past vanity metrics and into the operational reality of how AI is functioning—or failing—within your specific context.
True AI maturity is not a number. It is a culture of continuous, honest self-assessment combined with the organizational agility to act on what that assessment reveals. The Boris Cherny AI stages framework gives leaders a vocabulary for that conversation. The 2026 productivity study gives them the urgency. And the self-assessment prompt gives them the starting point. What you do with those three things is the real test of your AI leadership.
Summary
- Counting AI agents deployed is an activity metric, not a maturity indicator—deployment volume does not equal effective AI implementation.
- Boris Cherny's five-stage AI adoption framework diagnoses maturity by the specific challenges an organization faces: trust, capacity, policy enforcement, and resource management.
- A 2026 study found 84% of software engineers felt more productive with AI tools while simultaneously producing lower-quality work—a critical warning about measuring self-reported productivity.
- The gap between perceived and actual productivity compounds into quality debt that surfaces as rework, customer issues, and reputational risk at scale.
- A structured self-assessment prompt focused on friction points—not tool satisfaction—reveals an organization's true AI maturity stage.
- Once the dominant friction type is identified, interventions can be matched precisely: workflow redesign for trust gaps, process redesign for capacity constraints, and governance alignment for policy failures.
- True AI maturity is a culture of honest self-evaluation combined with the agility to act on what that evaluation reveals.