Learn where hospitals operational gaps lose value, and what your teams can do about it

Read the report
scroll to top


Machine-Generated Data: The Answer to Improving the Efficiency of Healthcare AI

Today, sixty-five percent of hospitals now use predictive AI models, the vast majority of which were built by the same company that developed the hospital’s electronic health record.

Clearly, AI has arrived in healthcare. Whether it is actually working is a separate question, and the evidence suggests the answer is complicated.

A survey of 67 health systems found that 77% of respondents identified immature AI tools as their single greatest obstacle to widespread adoption. That finding points to a problem that is common to AIs across any industry: algorithms are only as capable as their data inputs, and right now, there’s not enough high-quality data to feed the machines.

The problem with EHR data

When health systems talk about leveraging AI, the conversation almost always turns to the EHRs, the largest repository of patient and operational data within most hospitals. As a result, they are the logical starting point for any data-driven initiative.

Unfortunately, training AIs solely on EHR data is insufficient, simply because EHRs were never designed as AI inputs. Instead, they were created to document care, support billing, and satisfy regulatory requirements. These purposes produce data with very different characteristics than what AI models need to perform well.

A comprehensive review published in ACM Computing Surveys examined EHR data quality across more than a decade of research, and found that raw EHR records are routinely prone to inaccuracies, inconsistencies, sparsity, and bias, due to several factors:

  • Clinicians, pressed for time, type quickly and make errors.
  • Notes get copied forward from previous visits without review.
  • The same patient receives different diagnostic codes across different encounters.
  • Data recorded across departments and facilities rarely conforms to a common standard.

These factors all share the same root cause: the difficulty of clinical documentation under real-world conditions, when nurses or physicians are pressed for time, overwhelmed by paperwork, and data is siloed by departments or facilities.

Whatever the cause, however, the results are the same: feeding this data into an AI model does not produce smarter outputs, but unreliable ones. In cases where the training data carries embedded bias, it produces outputs that actively mislead the clinicians relying on them.

Requiring more documentation is the wrong answer

The reflexive response to unreliable data is to ask for better, more complete documentation. But registered nurses already spend nearly 25% of every shift on EHR documentation and review, while physicians spend close to 49% of their working hours on administrative tasks; in both cases, this time comes directly at the expense of patient care.

These figures reflect a workforce that has reached the limit of what human abilities can sustain. To ask clinicians to generate more data, of higher quality and at greater volume, is not the path to better AI. After all, the documentation burden already contributes to burnout, which then leads to high turnover, costing the average hospital up to $6.2 million annually in recruitment and replacement costs for nurses alone.

The ceiling on EHR data quality is a human capacity problem; it cannot be overcome by placing additional demands on people already stretched past their limits.

How machine-generated data makes AI successful

Instead of adding more water to a glass that is already full, hospitals can determine how to offload this burden onto technology. For instance, what operational data could be collected continuously, accurately, and at scale without placing more demands on clinical staff?

Sensors, location solutions, and automated data pipelines already exist in many hospital environments. For hospitals, the next step is to use this infrastructure to generate the kind of ambient, structured, continuous data that AI models need in order to function well.

The contrast with human-generated documentation is significant across three dimensions.

Accuracy increases when collection is automated, because sensors do not mistype, copy forward, or record data hours after the fact under time pressure. The data produced by sensors will reflect what actually happened, as well as when and where it happened.

Consistency also improves when collection is programmatic, because sensors can be configured to capture data at defined intervals and in specific formats, producing the structured, regular inputs that predictive modeling depends on.

Lastly, volume increases without adding any burden on clinical staff, because a single location sensor passively generates continuous data across an entire shift, without requiring a clinician to ever touch a clipboard or punch in extra information into their EHR.

What this looks like in practice

Across all of a hospital’s operational use cases, devices are already recording important data to improve AI performance.

Asset location sensors track the movement of infusion pumps, wheelchairs, and monitoring equipment across a facility continuously, generating data on utilization, distribution, and availability that no nurse charting system captures. This also greatly reduces search times; right now, nurses spend up to 60 minutes per shift searching for equipment, time which could be better used for patient care.

Staff badges record location data, documenting where care interactions are occurring and how long they take, filling gaps in the clinical record or correcting skewed EHR timestamps.

Room occupancy sensors capture patient flow data in real time, creating the kind of granular, continuous journey model that discharge planning algorithms require to function accurately.

Hand hygiene monitoring systems record compliance events at a scale no human observer can match; one study captured over 630,000 events in a single study period, compared with 480 recorded manually. Given that hospital-acquired infections (HAIs) account for 72,000 deaths and 687,000 infections annually.

These models can also decrease length of stay, and improve hospital revenue. In 2023, a three-month study across 50 New York hospitals found 992 patients experiencing discharge delays of over two weeks, resulting in an estimated $167 million in costs. Yet even small interventions have outsized impacts: a recent study found that reducing unnecessary stays by a single day could regain up to $2,373 per patient.

Each of these data streams feeds AI with inputs that are more accurate, more consistent, and more complete than what clinical documentation alone provides, and none of them require a nurse, physician, or any other licensed clinician to document anything.

The data beneath the model

The conversation about AI in healthcare has focused heavily on algorithms, model architecture, and integration infrastructure, all of which are crucial. But AI outputs will only be as good as their inputs; flawed, incomplete data generated by overstressed clinicians in difficult circumstances will negatively affect AI results.

Instead, health systems that want better AI performance need to look upstream, at where the training data comes from and what conditions it was created under. For optimal outputs, EHRs alone are not enough; health systems also need machine-generated data.


● KIO AI Assistant

Intelligently orchestrate
your hospital operations.

Ask KIO