It is AI that can actually be used in hospital settings. In other words, AI that puts patient safety above all else.
What We Found by Analyzing 3 Real Medical AI Cases
Even AI developers who are board-certified physicians run into performance problems they can't solve.
Even professionals equipped with both medical knowledge and development expertise run into performance problems that persist no matter how carefully they design the model architecture or fine-tune the parameters.
So what sets apart the medical AIs that deliver results in real clinical settings? To answer this question, we examined three real-world medical AI cases. Our analysis revealed that all three shared one thing in common: the underlying medical data was well-prepared. No matter how much budget and time you invest, if the foundational data is in disarray, that investment will not translate into the results you expect.
Drawing on three real-world case studies, this article shares practical insights for anyone wrestling with medical AI development.
Analysis of 3 Real-World Medical AI Cases
1. The Brain Hemorrhage AI Study at Geisinger Health System in the U.S.
A brain hemorrhage diagnosis that took over 8 hours, reduced to 19 minutes by AI.
Geisinger is a major healthcare system based in Pennsylvania. The Geisinger research team hypothesized that automatically analyzing head CT scans and reprioritizing the radiology worklist could reduce the time to brain hemorrhage diagnosis, and conducted their study accordingly.
Notably, this research didn't stay in the lab. The model's performance was validated directly in a real clinical environment.
How Was the Data Analyzed and Improved?
The research team collected 46,583 head CT exams, approximately 2 million images, taken at multiple Geisinger facilities between 2007 and 2017. Of these, 37,074 exams were used to train the deep learning model, and 9,499 exams the model had never seen were used to evaluate its performance.
Category | Details |
Data source | Multiple imaging facilities within the Geisinger system |
Collection period | 2007 to 2017 (10 years) |
Total data | 46,583 head CT exams (approx. 2 million images) |
Training data | 37,074 exams |
Evaluation data | 9,499 exams (not used in training) |
Clinical deployment | Applied to the live radiology worklist for 3 months |
What did the clinical deployment phase look like? For three months, the research team connected the model to the actual radiology worklist. When a brain hemorrhage was detected in a head CT classified as "routine," the exam was escalated to "emergency" priority. The team then compared diagnosis turnaround times between the reprioritized exams and routine exams.
Out of 347 routine exams, the AI flagged 94 as brain hemorrhages, and about two-thirds of those were confirmed by radiologists. Among them were 5 patients newly diagnosed with brain hemorrhages, and their average time to diagnosis dropped from roughly 8.5 hours to 19 minutes.
The study is also a useful example of how to handle apparent model errors. Here, a "false positive" was an exam the model flagged for hemorrhage but the interpreting radiologist read as normal. Because the radiologist's read served as ground truth, some false positives may not have been model errors at all. They may have been hemorrhages the radiologist missed.
To check, the team had a neuroradiologist who was blinded to the study review the false-positive cases. Instead of writing off every disagreement as a model error, they verified whether the ground truth labels themselves were correct.
2. The Korean Medical LLM Case at Seoul National University Hospital in Korea
An open-source model that surpassed the accuracy of physicians?
In 2025, Seoul National University Hospital (SNUH) developed Korea's first Korean medical large language model (LLM). In experiments on the past three years of the Korean Medical Licensing Examination, the model achieved 86.2% accuracy, becoming the first Korean open-source model to surpass the average accuracy of actual physicians (79.7%).
The performance figures are certainly striking, but what deserves even more attention is how that performance was achieved. When a medical AI is introduced into a new country or hospital, even the most capable general-purpose model inevitably runs into barriers around local language, documentation practices, and regulations. The Korean medical field was no exception. Korean clinicians document in a mix of Korean and English, and they constantly use abbreviations and shorthand.
To solve this, SNUH reorganized clinical data according to the field's own documentation standards. This case isn't specific to one country. Any company pursuing localization, anywhere, can learn from SNUH's approach.
How Was the Data Analyzed and Improved?
SNUH leveraged an enormous body of clinical text from within the hospital, including inpatient admission notes, outpatient records, and surgical, prescription, and nursing records, amounting to more than 38 million documents. From this text, the hospital built a Korean medical text corpus. After pseudonymizing and de-identifying personal information, the corpus was released for safe use within the hospital.
The team then integrated Korean medical laws, Korean-language paper abstracts, and clinical practice guidelines from academic societies, built a dictionary of medical abbreviations, and carried out terminology standardization. The same condition is abbreviated and documented differently by different clinicians and different departments. Without consolidating these into a single standard, an AI will recognize the same entity as multiple different ones.
The team went one step further. They created department-specific instruction-tuning datasets that reflect actual clinical workflows, and developed knowledge graph-based
retrieval-augmented generation (RAG) along with a multidisciplinary multi-agent framework. They did not stop at unifying terminology; they structured the relationships between medical concepts so the model could reference them.
3. The Sepsis Risk Prediction AI Case Based on the MIMIC-IV Database in the U.S.
This study developed an AI that predicts the risk of sepsis in patients admitted to the ICU with non-traumatic subarachnoid hemorrhage, using MIMIC-IV, the critical care database of Beth Israel Deaconess Medical Center in the U.S.
When developing medical AI, even the early step of defining the study population demands careful attention. Choose the wrong population, and the AI learns the patterns of an irrelevant patient group as ground truth, ultimately producing misguided predictions for the very patients who need help most. So what process did this study follow to analyze the data and select its subjects?
How Was the Data Analyzed and Improved?
The study defined its population using the following criteria.
Category | Criteria |
Inclusion | Age 18 or older |
Inclusion | Admitted to the ICU during the first hospitalization |
Inclusion | ICU stay of 24 hours or longer |
Exclusion | ICU stay of less than 24 hours |
Exclusion | More than 20% of the total records missing |
Let us look at the reasons for exclusion.
Patients who stayed in the ICU for less than 24 hours simply do not provide enough time to observe the onset and progression of sepsis.
For patients with more than 20% of their records missing, allowing the AI to arbitrarily fill in those gaps risks learning nonexistent patterns about sepsis.
Applying these criteria to admission records from 2008 to 2022, a final total of 1,052 patients were included in the study.
Documenting these criteria clearly is just as important. Because the exclusion criteria and the patients excluded under them are both explicitly documented, the AI model's judgments can be traced back and reviewed later. In other words, filtering data with clear criteria and transparently documenting those criteria comes before simply securing more data.
What Matters Most When Analyzing Data in Healthcare?
1. The AI Must Recognize the Same Thing as the Same Thing
The terminology standardization we highlighted in the SNUH case is a step every medical data project must go through.
The same test is recorded under different names and codes across hospitals and systems.
The same drug may be documented by its brand name or by its ingredient name.
The same diagnosis is abbreviated differently by different clinicians.
To a human, these are obviously the same thing. To an AI, they are three entirely different things.
What happens if training begins in this state? In reality, one patient received the same test three times, but the AI perceives three different tests each performed once. The volume of data stays the same, while the signal within it is fragmented and scattered.
1. Standard Terminology Systems
This is why unifying test names, drug codes, and diagnosis names under a single standard is essential. The foundational starting point is a standard terminology system: mapping hospital data to established international and domestic standards, such as LOINC for tests, KCD for diagnoses, and active ingredient codes for drugs.
2. Ontology
However, there comes a point where unifying codes alone is not enough. "Pneumonia" and "bacterial pneumonia" carry different codes but share a hierarchical relationship, and a brand name and an ingredient name can refer to the same drug despite looking entirely different. Only when these relationships are also taught to a medical AI can it distinguish them as clearly as a clinician would.
Structuring the relationships between concepts in this way is called an ontology. The knowledge graph we saw in the SNUH case follows the same principle. It goes beyond listing terms like a dictionary, defining in advance which concept is a subcategory of which, and which concepts are equivalent.
2. A Deep Understanding of Clinical Practice
Data can only be analyzed correctly when its structure is designed with an understanding of clinical workflows. Building accurate logic requires understanding how, and why, hospitals actually record data the way they do.
So if you are considering an external specialized solution, verify the depth of its domain knowledge. Have an in-depth conversation, and judge based on whether they focus only on data processing technology or start by asking about your hospital's documentation practices and clinical workflows.
No matter how skilled a medical AI data expert may be, they cannot know every detail of a given hospital's clinical workflows and systems. That is why deep communication with the people who know the field best is absolutely essential. The deeper that dialogue goes, the stronger the medical AI's performance.
3. Data Structures Must Be Designed So the AI's Decision-Making Process Can Be Clearly Explained
Medical AI development is subject to far denser regulation than other industries: the data carries health and safety stakes as well as sensitive personal information, and if it's breached, that risk falls on patients.
The Korean government has, in fact, codified this into law. The AI Framework Act (Framework Act on the Advancement of Artificial Intelligence and Establishment of Trust), which took effect in January 2026, classifies healthcare as a "high-impact AI" domain that directly affects human life and physical safety, and imposes obligations for safety and transparency.
This calls for two conditions.
When a medical AI makes a judgment, it must be possible to transparently verify why it reached that judgment.
That transparency comes from the data. To logically explain an AI's judgment from evidence to conclusion, the underlying dataset itself must be logically organized.
In Data Clinic, Pebblous' data quality management solution, every stage is logged: which data was excluded under which criteria, how the metrics changed before and after processing, and more.
When Pebblous builds data for medical AI, we pursue one goal, and it isn't impressive accuracy figures or massive data volumes.
💡
The medical data your organization holds today may have quality issues hiding beneath the surface. Data Clinic will diagnose those issues precisely and turn the data into something your team can trust and use in real hospital settings.