Turning AI’s Promise into Performance in Health Systems | Post 3 of 3
In one health system I would describe as typical, the AI steering committee has been meeting for fourteen months. It is a serious body. General counsel sits on it. So does the chief medical information officer, the chief nursing officer, the privacy officer, and a rotating clinical representative. It meets monthly, works through a queue, and has approved four use cases.
Two floors down, a nurse manager is pasting a patient handoff summary into a consumer chatbot on her phone because she is trying to write a coherent report at the end of a twelve-hour shift and the tool is faster than the alternative. Nobody authorized it. Nobody knows it happened. There is no record that it happened.
This is the actual state of AI governance in a great many health systems: a rigorous, slow process that governs a small number of sanctioned things, running alongside a large volume of unsanctioned activity that the process cannot see and was never designed to catch.
The committee is not failing at its job. It is succeeding at a job that turned out to be the wrong one. It was built to evaluate projects, and the risk arrived as behavior.
Governance is not the brake on AI adoption. It is what allows you to take your foot off it. A health system with enforced policy boundaries and complete visibility into what its AI is doing can afford to build its data foundation domain by domain, because the blast radius of a mistake is bounded by architecture rather than by caution. The organization that waits for perfect data before deploying anything governs nothing, because its people have already started without it.
The Premise Worth Questioning
McKinsey’s June 2026 report on the health system CEO imperative contains a recommendation that reads as almost contradictory on first pass. Do not wait for the perfect data foundation. Do set up governance guardrails before you begin.
Most health systems have that sequence exactly reversed. Data readiness is treated as the prerequisite, the thing that must be substantially finished before serious AI work can begin. Governance is treated as a review gate, a compliance function consulted near the end to bless what has been built.
The premise worth questioning is that this ordering is the safe one.
The Case for Waiting on Data
Start with why health systems believe what they believe about data readiness, because they came by it honestly.
Health system data is genuinely difficult in ways that most enterprise data is not. Master patient index problems mean the same human being exists three times under slightly different names. Coding practices vary by service line and by individual physician. Legacy systems from three acquisitions ago still hold records nobody has mapped. Reference data drifts. Clinical documentation contains meaning that lives in free text and in habits of phrasing that vary by department.
Layered on top of that is a memory. Most health systems have already lived through an analytics era in which dashboards were built on data that turned out to be wrong, in which two departments brought conflicting numbers to the same meeting, and in which credibility was spent and not recovered. The instinct to fix the data first is not naive. It is scar tissue.
And the stakes are asymmetric. In most industries, a decision made on bad data produces a bad quarter. In a health system it can produce a harmed patient, an OCR investigation, and a story with the organization’s name in the headline. Given that asymmetry, waiting looks like prudence.
Why Waiting Does Not Work
Two things go wrong.
The first is arithmetic. There is no finish line. Health system data will not be clean in three years any more than it was clean three years ago, because the organization keeps acquiring practices, changing systems, and generating new documentation faster than anyone remediates the old. A prerequisite that never completes is not a prerequisite. It is a deferral.
The second problem is more interesting and less obvious. Foundations built ahead of demand get built wrong.
When a data team constructs a canonical model without a specific domain forcing the requirements, it makes hundreds of judgment calls with no evidence to guide them. Which fields matter. What grain to model at. Which of four plausible definitions of an encounter to canonicalize. Those calls get made from first principles, documented carefully, and then discovered to be misaligned the moment real work arrives. The report makes this point directly: foundations grounded in actual business needs are stronger than foundations built in a vacuum, and they arrive sooner.
Build the data capability the prioritized domain requires, and let the domain tell you what it requires. If you are rewiring the back end of revenue cycle, you need eligibility, authorization, documentation, payer policy, and claims to be reliable and connected. You do not need the ambulatory scheduling data model resolved to do that work. Solve for the domain, and let reusable capability accumulate as a byproduct of shipping rather than as a precondition for starting.
Why Governance Is the One Thing You Cannot Defer
Here is the asymmetry that makes the reversal work.
Data gaps degrade gracefully. An agent working with incomplete reference data produces a lower-quality answer, a rework queue, an exception for a human to resolve. It is a cost, and it is a recoverable one.
Governance gaps do not degrade gracefully. They fail catastrophically and often silently. An agent that reached data it should not have reached does not produce a slightly worse output. It produces a disclosure event that is discovered weeks later, if it is discovered at all. There is no partial credit and no gradual signal.
This is why the sequencing inverts. The thing that fails softly can be built as you go. The thing that fails hard has to exist before the first agent runs.
And notice what governance actually buys you once it is in place. It is not a constraint on how fast you can move. It is the condition that makes moving fast survivable. If every agent operates inside enforced access boundaries, if every action leaves an audit trail, if a policy violation is caught by the architecture rather than by someone noticing later, then a mistake in an imperfect data foundation is contained by design. You can afford to build iteratively precisely because the failure mode is bounded.
Philosophy, Not a Rulebook
The report is specific about what form governance should take, and the distinction matters more than it might appear.
Guardrails should articulate a coherent philosophy about what the organization will and will not do with AI, along with a process for evaluating new applications against it. They should not be a prescriptive rulebook enumerating approved use cases.
The reason is practical. A rulebook written against the capabilities of AI in early 2026 is substantially obsolete by late 2026. Every new capability generates a question the rules did not anticipate, which routes back to the committee, which meets monthly, which is how you arrive at fourteen months and four approvals. A rulebook creates a queue. A philosophy creates a test that a product owner can apply on Tuesday without convening anyone.
A philosophy also travels. It gives a nurse manager, a revenue cycle director, and a supply chain analyst the same basis for judgment, which is the only way principles get applied consistently across an organization too large for any committee to supervise directly.
Set the philosophy up front. Set the enforcement up front. Let the specific applications be evaluated against both as they arrive.
The Half of Governance Nobody Budgets For
There is a second recommendation in the report that most readers will file under performance management rather than governance, and I would argue it belongs in both places.
Each domain team should carry near-term and long-term metrics, reported quarterly to the executive team and the CEO. Long-term KPIs measure the outcome, something like a ten percent reduction in cost to collect. Short-term KPIs measure progress toward it, something like twenty minutes saved per appeal letter. When metrics are not being met, owners are expected to act decisively: remove the barrier, pivot, or stop the work.
That last expectation has a technical dependency that goes unstated. You cannot stop what you cannot see. A product owner can only act decisively on underperformance if there is a live, trustworthy account of what the agents actually did, which decisions they made, which data they touched, and where the process broke. Absent that, quarterly reporting becomes a slide of directional claims and the decision to stop something never gets made, because nobody can prove it is not working.
What This Looks Like Architecturally
At Datafi we built the platform on the conviction that governance is load-bearing infrastructure rather than a feature layered on afterward, and that conviction shows up in a few specific places.
Cyber embeds policy, access control, and data boundaries into the foundation. An agent does not have access because someone remembered to restrict it. It has exactly the access its policy grants, enforced at the architecture level, which means the boundary holds regardless of what the agent is asked to do or who asks it.
Control Tower provides visibility into how AI is being used across the organization, with a complete record of activity. This is the observability layer that makes the measurement discipline above executable rather than aspirational, and it is also the answer to shadow AI. Employees reach for unsanctioned tools when the sanctioned path is slower or absent. Giving them one governed platform that is genuinely faster than the alternative is the only durable solution to that problem, and Control Tower is how you verify it is working.
The global business contextual layer is what allows the foundation to be built the way McKinsey recommends. Because it federates to source systems in place rather than requiring ingestion or migration, you model the context a domain needs when that domain needs it. Your data does not leave your environment. Your EHR, ERP, and claims systems stay where they are. The context accumulates domain by domain instead of arriving as a three-year prerequisite.
On the compliance side, Datafi maintains SOC 2, ISO 27001, GDPR, and CCPA compliance, published on our public trust portal. For a health system evaluating any AI platform, that documentation should be the beginning of the conversation rather than the end of it. The right question is not whether a vendor holds certifications. It is whether the governance is architectural or procedural, because only one of those survives contact with an autonomous agent operating at two in the morning.
The Sequence, Stated Plainly
Across this series the argument has been a single one, approached from three directions.
Health systems are not short on AI. They are short on the architecture that lets AI compound, which is why fifty pilots have produced adoption without impact.
The unit of transformation is the domain rather than the task, because value in health system operations leaks at the handoffs between functions, and automating individual tasks preserves exactly those seams.
And the way to get there is not to wait. Set the guardrails and the observability before the first agent runs, then build the data foundation in service of a specific domain rather than in advance of all of them.
The report closes by calling this a leadership challenge rather than a technical challenge, and I agree with the emphasis while wanting to add one qualification. It is a leadership challenge that has a technical prerequisite. A CEO can set the ambition, pick the domain, stand up the cross-functional pod, and demand the metrics, and still get fifty disconnected pilots if the underlying architecture cannot hold shared context, coordinate work across functions, and enforce policy without human supervision at every step.
Vision and architecture are not alternatives. The organizations that will reverse healthcare’s productivity trajectory are the ones that treat them as the same project.
This concludes the series. If you are a health system leader working through where to start, the conversation I find most useful is the specific one: which domain, what the handoffs look like today, and what would have to be true for the work to run differently ninety days from now.
Datafi is the operating system for business AI. We help mid-enterprise organizations unify their data, deploy governed AI agents, and deliver measurable outcomes in weeks rather than years, without migrating off the systems that already run the business.
Turning AI’s Promise into Performance in Health Systems
Part 1: Health Systems Aren’t Short on AI. They’re Short on Impact.
Part 2: Rewire the Domain, Not the Task
Part 3: Guardrails First, Foundations As You Go
Written by Vaughan Emery, Co-Founder & Chief Product Officer at Datafi
Source: “The health system CEO imperative: Turning AI’s promise into performance,” McKinsey & Company Healthcare Practice, June 2026.

