Uncategorized

93% of Hospitals Are Already Running AI—Only 44% Have Tested It Properly, New Report Finds

A new UPMC and KLAS Research survey finds 93% of health systems have deployed third-party AI, but only 44% have a dedicated environment to test it—exposing a widening gap between adoption and governance.

93% of Hospitals Are Already Running AI—Only 44% Have Tested It Properly, New Report Finds

A new survey of health system leaders lands on a blunt conclusion: American hospitals have raced ahead with artificial intelligence far faster than they have built the guardrails to manage it. The report, titled “Validation and Trust: How Health Systems Are Testing and Governing Analytics and AI Solutions,” was published in early August 2026 by the Center for Connected Medicine at UPMC together with KLAS Research, based on interviews and survey responses from more than two dozen health system executives. Its headline number is stark: 93 percent of health systems have already deployed third-party AI tools into clinical or administrative workflows, yet only 44 percent maintain a dedicated data environment—a sandbox—to test those models for accuracy, safety, and drift before turning them loose on real patients.

A Widening Gap Between Adoption and Oversight

The report frames the imbalance as the defining risk of this phase of health-system AI. Sixty-three percent of the surveyed organizations describe their overall AI strategy as still “developing” or largely ad hoc, and only 4 percent call their approach “advanced.” That gap matters because AI models, unlike traditional software, can degrade silently as patient populations, documentation habits, or lab equipment change underneath them—a phenomenon researchers call model drift. Without a sandbox to periodically re-test a deployed algorithm, a hospital may not notice for months that a tool’s accuracy has quietly slipped.

Where the AI Is Actually Being Used

The survey found clinical documentation and ambient scribing tools are the single most common AI use case, cited by 52 percent of respondents, ahead of revenue-cycle and coding automation at 36 percent, medical imaging analysis at 32 percent, and EHR-embedded clinical decision support also at 32 percent. That ordering reflects where health systems have found the least resistance: documentation AI touches billing and paperwork rather than direct treatment decisions, making it an easier entry point than, say, an algorithm that recommends a diagnosis or a drug dose.

Why Governance Keeps Losing the Race

Executives interviewed for the report pointed to a familiar list of obstacles: constrained budgets and staff time, unclear lines of accountability for AI performance, the sheer volume of vendor pitches competing for attention, and a lack of consensus on how to even measure whether an AI tool is working. Ninety-two percent of organizations say they do conduct some form of pre-deployment testing, but the report’s authors note that testing once before go-live is a different discipline than continuously validating a model after it is embedded in daily clinical workflows—and it is the latter that most systems are missing.

The View From Health System Leadership

Officials involved with the report frame the findings less as an indictment and more as a call to formalize what has largely been improvised. The Center for Connected Medicine has positioned the study as the first in a planned series aimed at giving hospital boards and C-suites a benchmark for what mature AI governance should look like, rather than leaving each system to reinvent validation processes independently. Vendors, for their part, have generally welcomed the scrutiny publicly, arguing that clearer, shared standards for testing and monitoring AI would reduce the reputational risk that comes when any single high-profile AI failure—a missed diagnosis, a biased billing recommendation—taints the credibility of the entire category.

What Critics Say Is Missing

Patient-safety advocates and some health IT researchers argue the report likely understates the problem, since it relies on self-reported survey data from institutions that volunteered to participate—systems mature enough to answer detailed governance questions may be the ones already ahead of the curve, while laggards go uncounted. Others note that “pre-deployment testing” can mean wildly different things across organizations: some systems rigorously validate a model’s outputs against their own patient population before go-live, while others rely almost entirely on a vendor’s marketing claims and published accuracy statistics from a different hospital’s dataset entirely.

What Comes Next

The report’s authors expect that governance frameworks, rather than model performance claims, will increasingly become the deciding factor in which AI vendors win hospital contracts, as boards grow wary of being blindsided by an algorithm that quietly stopped working. Expect more health systems to stand up formal AI governance committees over the next year, along with growing demand for third-party auditing services that can independently verify a vendor’s accuracy claims. For patients, the practical stakes are largely invisible day to day, but the report’s authors argue that the difference between a hospital with a real validation sandbox and one without could eventually show up in something as consequential as whether a documentation error, a missed billing flag, or a decision-support recommendation reaches a clinician’s desk uncorrected.