As Healthcare AI Autonomy Grows, Governance Must Grow With It

The controls that work when AI assists a human are not enough when AI begins acting on the organization’s behalf
One of the most important shifts happening in artificial intelligence is not simply that AI is becoming more powerful. It is becoming more autonomous. There is a significant difference between an AI system that drafts a summary for an employee to review and an AI agent that can interpret information, decide what happens next, interact with another system, or initiate an action. That distinction should matter tremendously to healthcare leaders because as healthcare AI autonomy increases, the mechanics of governance must change with it.
At the beginning of an AI journey, humans may review virtually everything the AI produces. That can work during a pilot, but it will not work indefinitely at scale. The challenge becomes: How do we move from humans reviewing everything AI does to humans governing a system in which AI may increasingly operate independently?
That is where AI governance becomes much more sophisticated.
Start With a Simple Principle: More Autonomy Requires More Governance
Not every AI use case carries the same risk. An AI assistant helping an employee organize meeting notes is very different from an AI system analyzing patient information. An AI tool summarizing a regulation is different from an AI agent deciding how a regulatory change should be routed across an organization. An AI system recommending an action is also fundamentally different from one actually taking the action. I like to think about the progression this way: Assist → Recommend → Decide → Act
As we move from left to right, AI autonomy increases. As autonomy increases, organizations need stronger controls around authority, accountability, transparency, monitoring, escalation, intervention, and auditability. The governance model should evolve with the level of authority we give the technology.
Stage One: Pilot — Humans Review the Work
During an early pilot, AI autonomy should generally be relatively low. The AI performs a task, a human reviews the output, and that person determines whether it is accurate and appropriate before anything consequential happens.
For example, imagine an AI system helping a compliance team monitor regulatory changes.
The AI retrieves a new CMS publication, summarizes the information, and identifies what appears to have changed. The compliance professional then reviews the source, validates the interpretation, and determines whether action is necessary.
At this stage, organizations should pay close attention to output quality, accuracy, error rates, source reliability, and the types of corrections humans make. The organization is learning how the technology behaves, and human review is part of that learning.
Stage Two: Scaling — Humans Review the Exceptions
Now imagine the system performs well and the organization wants to expand it. Volume increases, more people begin using it, and more workflows become involved. At some point, reviewing every AI output becomes inefficient and may eliminate much of the value AI was supposed to create. This is where governance begins shifting from humans reviewing everything to humans reviewing what requires human attention.
That requires something very important: The organization must define what constitutes an exception.
The AI might encounter conflicting regulatory guidance, fall below an established confidence threshold, make a recommendation that could affect reimbursement, handle PHI, identify a potential overpayment, or propose an action that exceeds its authority. Those conditions should trigger escalation.
The organization therefore needs documented decision rights, risk thresholds, escalation paths, human approval points, override authority, and exception-handling procedures. And this is where orchestration becomes particularly important. The system needs to know not only what work to perform, but when it is no longer authorized to proceed.
Stage Three: Production — Humans Govern the System
This is where the leadership conversation becomes particularly interesting. When AI reaches a mature production environment, humans may no longer review every individual output. They may oversee the system instead, which is a very different governance model.
Think about other complex systems already operating in healthcare. We don't have a compliance officer personally review every claim before it is submitted. Instead, we create controls, monitor patterns, conduct audits, investigate exceptions, measure performance, identify trends, and intervene when something falls outside expected parameters. Mature AI governance will increasingly operate in a similar way.
Humans move from reviewing transactions to governing performance. That requires more sophisticated infrastructure, including audit trails, automated monitoring, performance thresholds, exception reporting, feedback loops, incident management, change controls, version tracking, access controls, override mechanisms, and clearly defined accountability.
At this level, governance is not weaker because humans review fewer individual outputs. It actually needs to be stronger.
The Critical Shift: From Output Review to Exception Management
This is one of the most important concepts for healthcare leaders to understand.
Early AI governance tends to focus on the question, "Was this answer correct?" Mature AI governance has to ask whether the system is performing as expected, where it is failing, and whether exception or override rates are changing. Leaders also need to know whether particular populations, services, payers, or workflows are producing different results; whether the data, model, or regulatory requirements have changed; and whether the agent is behaving within its authorized boundaries.
Those questions allow leadership to govern AI at scale.
What We Measure Must Change Too
The slide captures another important concept: the metrics should evolve along with autonomy.
During a pilot, we may primarily measure output quality and error rate. During scaling, we begin looking at exception rates and review cycle time. In production, we should increasingly measure business outcomes and compliance performance.
That progression matters.
Because ultimately, the purpose of AI is not to generate impressive outputs.
It is to create measurable organizational value.
For healthcare, those outcomes might include reduced administrative burden, improved audit accuracy, faster regulatory analysis, fewer denials, better documentation, a faster response to regulatory change, improved payment integrity, lower compliance risk, greater workforce capacity, and better decision quality. And we should measure unintended consequences as carefully as intended benefits.
A 2% Error Rate Can Mean Very Different Things
This is where healthcare leaders need to resist overly simplistic AI metrics. Suppose an AI system is 98% accurate. That sounds excellent, but what is happening in the other 2%? If those errors involve formatting a report incorrectly, the risk may be relatively low. If they involve incorrectly recommending that a claim is billable, failing to identify an overpayment, exposing protected information, or influencing patient care, that same 2% becomes a very different governance concern.
AI risk cannot be measured by accuracy alone. We need to understand the consequence of the error.
That means governance thresholds should be tied not only to probability, but also to impact.
Human-in-the-Loop Must Evolve Into Human-in-Governance™
This progression reinforces why I believe Human-in-Governance™ is a more useful long-term concept than simply human-in-the-loop. Human-in-the-loop often means a person reviews what the AI did.
That makes sense during early adoption and in high-risk decisions where individual review remains necessary. But it cannot be the entire governance strategy for enterprise AI.
Human-in-Governance™ means people retain authority over the system even when they are not reviewing every individual transaction. Humans determine what AI is authorized and prohibited from doing, what requires review, what constitutes an exception, which thresholds trigger escalation, what evidence must be retained, and what performance is acceptable. They also decide when autonomy can increase, when it should be reduced, and when the AI should be stopped entirely.
The human may move out of every individual loop without ever moving out of governance.
That distinction is going to become increasingly important.
Autonomy Should Be Earned
I would add another principle for healthcare organizations: AI autonomy should be earned—not assumed. We should not begin with maximum autonomy simply because the technology is technically capable of it. Organizations should start with controlled authority, measure performance, observe errors, understand exceptions, test the intervention process, and evaluate how humans interact with the system. Then, if the evidence supports it, they can expand what AI is permitted to do.
I think of this as an autonomy ladder:
AI prepares → Human decides
AI recommends → Human approves
AI acts within defined boundaries → Human reviews exceptions
AI operates within established authority → Humans monitor system performance
Importantly, that ladder should work in both directions. If performance deteriorates or risk changes, autonomy should be capable of being reduced.
Your Employees Become an Early-Warning System
This is also why staff involvement remains essential even as AI becomes more autonomous.
Employees working closest to AI-enabled workflows may notice problems before they appear on an executive dashboard.
They may see unusual recommendations, repeated overrides, missing information, workflow friction, unexpected behavior, or situations where the AI technically followed its instructions but produced an outcome that simply does not make sense. Organizations need mechanisms for those observations to travel upward.
Employees should know what to report, where to report it, whether they can override the AI or stop the process, and whether someone will investigate the concern.
The workforce is not outside the AI monitoring system. The workforce is part of the monitoring system.
This Is Where Leadership Has to Think Ahead
Healthcare leaders do not need to personally understand every model parameter or technical architecture.
But they should understand the progression of autonomy.
When evaluating an AI initiative, leadership should understand what AI is doing today, what it may eventually do, and what authority the organization is giving it. Leaders should define what evidence would justify greater autonomy, what humans would stop reviewing at scale, and what would replace that review. They also need to establish which exceptions automatically trigger intervention, how changes in performance will be detected, and who can reduce or revoke the AI's authority.
Those are governance questions.
But they are also fundamental operating-model questions.
The ProCode Perspective
AI governance cannot remain static while AI capability and autonomy continue to advance.
The governance model should mature alongside the technology. In a pilot, autonomy remains low, humans review outputs, and the organization measures quality and errors. At scale, autonomy becomes controlled, humans focus increasingly on exceptions, and leaders measure exception rates, overrides, escalation, and review efficiency.
In production, AI may operate with greater autonomy inside clearly defined boundaries while humans oversee system performance and consequential decisions. At that stage, organizations measure business outcomes, compliance performance, incidents, drift, and risk. Throughout every stage, human authority remains.
That is the distinction I believe healthcare organizations need to understand.
Responsible AI does not require a human to manually review every action forever; that would make meaningful AI scale nearly impossible. It requires something more sophisticated: a governance system that knows when humans need to be involved—and ensures humans retain the authority to intervene when it matters.
The future of AI governance is therefore not simply about keeping a human in every loop.
It is about ensuring that as AI earns greater autonomy, human accountability never disappears with it.





