scroll to top

Not All AI Is Created Equal: Why Predictive Models, Not Language Models, Run Hospital Operations

  • Key Takeaways
    65% of hospitals use AI, but 77% of health systems believe their AI tools are immature. This disconnect arises from treating AI tools interchangeably, rather than using the right AI tools for the right task.
  • Large language models (LLMs) generate the statistically likely next word; because their output is prose, the error is harder to identify and correct than a wrong prediction.
  • Predictive analytics learn from structured data, producing numbers a hospital can check against benchmarks.
  • Predictive analytics have a proven track record of forecasting hospital operational needs, including discharge timing, ICU flow, and bed demand, with 80% accuracy across multiple studies.

Today, 65% of hospitals use predictive AI models for tasks like discharge forecasting, bed capacity, and staffing. Most of these models were built by the same vendor that supplies the hospital’s electronic health record. At the same time, 77% of health systems name immature AI tools as their single biggest obstacle to further adoption.

This paradox can be easily explained: AI is a broad umbrella that covers tools with diverse architectures, limitations, and failure modes. As a result, health systems may be conflating very different tools, and using the wrong tool for the wrong task. In order to effectively implement an AI product, health system leaders must have a clear understanding of their goals and workflows, as well as the various kinds of AI and where they excel.

How do large language models and predictive analytics differ?

LLMs predict the most statistically likely sequence of words, while predictive analytics forecast measurable outcomes from structured data – a difference that determines which hospital tasks each can safely run.

This quirk of design is precisely why LLMs can sound so confident, even while stating something blatantly false: the model is optimizing for plausible language, not verified facts.

A 2026 study tested nine large language models on two hospital administration tasks: counting patient records and filtering them by criteria, using 50,000 real emergency department visits. Prompting a model to answer directly produced poor results across all nine models. Chain-of-thought reasoning, where the model shows its work before answering, helped only a little, and accuracy collapsed as the record set grew larger. Even GPT-4o, the strongest model tested, fell from roughly 95% accuracy on small tables to under 60% on larger ones.

Predictive analytics models function differently. They take structured, quantifiable data, admission timestamps, room sensor pings, lab values as discrete fields, and learn statistical relationships between defined variables.

Because their outputs are numbers or classification, not a sentence, they cannot hallucinate a citation or invent a symptom. Predictive analytics can still make mistakes, but their errors are the kind a hospital can measure, calibrate, and correct against a benchmark, instead of being buried in pages of confident-sounding copy.

Importantly, predictive analytics are already common across various industries and use cases, such as optimizing maintenance or forecasting drill failure events for oil and gas companies. As a result, leaders can adapt this proven technology to healthcare settings with more confidence.

Where predictive analytics excels in hospital operations

Predictive analytics consistently outperforms LLMs on core hospital forecasting tasks – admissions, discharges, ICU flow, and bed demand – with accuracy above 80% across multiple peer-reviewed studies.

Predictive analytics has an extensive, research-backed record of success in forecasting and coordinating hospital operations, whether it’s a head-to-head test against an LLM, predicting ICU flow and demand, or booking inpatient beds in advance.

Researchers at Mount Sinai tested GPT-4 against a traditional machine learning solution on a purely operational question: which emergency department patients would end up admitted as inpatients? Using records from more than 864,000 ED visits across seven New York City hospitals, the ML solution, built from XGBoost and Bio-Clinical-BERT, reached an accuracy of 82.9% and an AUC of 0.88.

GPT-4, reasoning on its own from the same records, performed considerably worse. It only improved its accuracy once researchers stopped asking it to reason unassisted, and instead fed it retrieved real-world examples and the ML model’s own probability scores to lean on. Left to reason from scratch, on a task that decides how many beds a hospital needs to hold open, the language model was the weaker tool.

Other researchers have found similar success using machine learning. Researchers built an interpretable predictive model from 63,432 hospital admissions and used it to forecast short-term discharges, flag long-stay patients, and anticipate ICU flow with accuracy above 80%.

In 2026, scientists assessed the performance of ML models for predicting elective surgical ICU admission at a UK NHS Trust. In a prospective evaluation, only 71.6% of patients who required an ICU bed had one electronically requested in advance. The researchers then tested the same model against the hospital’s existing electronic system, with both forecasting how many ICU beds the whole hospital would need the next day. The new model’s outputs came closer to what actually happened, meaning it made smaller errors than the system already in place.

Where LLMs succeed in healthcare (as a language layer)

LLMs perform best in hospitals as a language and coordination layer – translating questions, summarizing outputs, and routing tasks – not as the engine doing the forecasting itself.

While LLMs are poorly suited for predicting, orchestrating, and optimizing hospital operations on their own, they can be effective when incorporated into a larger system.

In one study, researchers built a patient scheduling system to coordinate appointments across 12 anesthesiologists and surgeons, 12 operating rooms, 15 patients, and three surgery types. While their solution included a LLM, its tasks were limited to translating user queries and algorithmic outputs for a separate model (not an LLM) that actually ingested data and carried out the core analysis work. This integrated system landed within 1% of the optimal solution found by branch-and-bound search, and outperformed two of the hospital’s existing scheduling approaches.

Another study attempted something similar, using seven LLMs, including ChatGPT-4o, LaMDA 2, PaLM 2, and Qwen, on allocating operating rooms, postoperative beds, and surgeons. Each LLM was part of a larger framework that combined prompt engineering, hospital data retrieval, and API calls to dedicated tools, and handled orchestration, rather than computation.

In this experiment, researchers leaned heavily on the strengths of LLMs: summarizing and interpreting complex data, and natural language processing. Rather than crunching data and creating predictions, LLMs simply served as a coordination and communication layer, helping administrators, clinicians, and patients understand, improve, and navigate the care process.

The right AI for every hospital task

Hospitals get the most from AI when predictive models handle the forecasting and LLMs handle the communication -treating them as interchangeable is what creates friction.

From these studies, one conclusion emerges: language models, left to reason on their own about admissions, schedules, or bed allocation, underperform predictive algorithms that were built for purpose. Given a narrower niche, such as translating outputs into natural language for end users or calling the right tool, LLMs perform much better.

After all, a model trained on structured data can be checked against a benchmark and corrected, but LLMs designed to generate language cannot be optimized the same way. If leaders treat these two AI types as interchangeable, then their deployments will encounter friction. Only by keeping predictions and language processing separate, can leaders play to the strengths of AI.

Kontakt.io separates its products in a similar manner. Patient Journey Analytics, the predictive foundation of Kontakt.io’s AI products, fuses EHR and RTLS data to forecast hospital operations and create a digital twin of hospitals and health systems. On top of it sits Curiosity Engine, a LLM layer that turns plain language questions, such as which departments have the highest patient traffic over the weekend, into a precise query for Patient Journey Analytics and other AI agents to answer, and then translates the resulting output into natural language.

Learn more about how Patient Journey Analytics forecasts demand and Curiosity Engine


● KIO AI Assistant

Intelligently orchestrate
your hospital operations.

Ask KIO