2026-03-17
This paper distills decades of experience in clinical epidemiology and prediction modeling into a set of recurring problems, provocatively called “enemies”, that undermine prediction models long before they ever reach the bedside.
Reading the paper resonated strongly with concerns I raised in an earlier blog post about AI/ML in medicine. In many projects, we spend enormous effort asking how to build models, while the important question of why this model should exist at all is often overlooked.
What this paper shows clearly is that the biggest obstacles to reliable clinical AI are rarely algorithmic.
Most problems appear before the model is even trained.
The Real Problem: Many Models, Little Clinical Impact
Prediction modeling is one of the most active areas of biomedical research. New models are proposed for disease diagnosis, prognosis prediction, high-risk patient identification, and treatment recommendation. However, systematic reviews consistently reveal a troubling pattern:
many models are developed
few are externally validated
even fewer are evaluated for clinical impact
almost none become routine tools in healthcare
In other words, the field produces many models but relatively little usable knowledge.
To explain why, the authors organize the challenges into twelve “enemies,” grouped into four broad categories:
the nature of clinical problems
the quality of available data
methodological choices in model development
the culture and incentives of biomedical research
Together, these enemies illustrate why reliable prediction modeling is much harder than simply applying machine learning to medical data.
Enemy 1: Medical Data Often Contain Weak Signals
Biological systems are highly complex. Most diseases arise from multiple interacting pathways involving genetics, environment, lifestyle, and chance. As a result, predictors rarely cleanly separate patients into clear outcome groups. For example, two patients with identical risk factors for cardiovascular disease may still experience very different outcomes. In such settings, flexible machine learning models can easily mistake random noise for a meaningful signal, especially when datasets are small. Interestingly, systematic comparisons often show that complex machine learning models do not consistently outperform simpler statistical models in clinical prediction tasks. More sophisticated algorithms cannot compensate for weak underlying signals.
Enemy 2: Individual Risk Is Inherently Uncertain
Prediction models frequently generate individualized risk estimates. For example: “Your risk of developing disease X within five years is 20%.”
However, a patient belongs to many possible reference groups depending on which predictors are considered. Two equally valid models might assign very different risk scores to the same individual while still performing similarly overall.
This highlights an often overlooked limitation: prediction models estimate probabilities, not certainties.
Enemy 3: Healthcare Systems Are Not Stable
Clinical environments vary across hospitals, countries, and time. A model trained in one setting may perform poorly in another because:
patient populations differ
diagnostic practices vary
treatments evolve
measurement procedures change
For example, a model developed using laboratory data from one hospital may rely on tests that are measured differently elsewhere. Even within the same institution, models may degrade over time as clinical practices change, a phenomenon sometimes called data drift.
Enemy 4: Clinical Data Are Messy
In medicine, the biggest challenge is often data quality. Electronic health records contain missing values, inconsistent coding, irregular measurements, and complex clinical workflows. For example, certain laboratory tests are only ordered when physicians suspect severe disease. When those tests appear “missing” in other patients, the missingness itself reflects clinical decision-making rather than random absence. Ignoring this structure can distort model predictions.
Enemy 5: Diagnostic Labels Are Imperfect
Prediction models rely on labeled outcomes. Yet in medicine, diagnostic labels are often uncertain. Some diseases can only be confirmed through selective, invasive procedures. Others depend on subjective clinical judgment or evolving diagnostic criteria. As a result, models may be trained on labels that are themselves imperfect approximations of reality.
Enemy 6: Treatment Can Change the Outcome
Prediction models can influence clinical decisions. This creates a paradox.
Suppose a model predicts that a patient has a high risk of a heart attack. Doctors intervene with preventive treatment, and the heart attack never occurs. When evaluated retrospectively, the model may appear inaccurate even though it helped prevent the event. This phenomenon is sometimes called the treatment paradox.
Prediction models interact with clinical decisions, which means they can change the very outcomes they aim to predict.
Enemy 7: Methodological Shortcuts
Many prediction models suffer from avoidable methodological problems. Common examples include:
small datasets with few outcome events
excessive model complexity
inappropriate variable selection
arbitrary categorization of continuous predictors. For instance, converting continuous variables such as age or blood pressure into simple thresholds (“above 60” vs “below 60”) discards valuable information.
Such practices can produce models that appear accurate during development but fail during validation.
Enemy 8: Poor Reporting
Even well-developed models become difficult to use when poorly reported. Researchers sometimes fail to publish:
the full model equation
predictor definitions
implementation details
Without these elements, other researchers cannot validate or apply the model. Scientific progress depends on reproducibility.
Enemy 9: Clinical Impact Is Rarely Studied
A model might accurately predict outcomes yet fail to improve patient care if clinicians do not use it or if it does not alter decisions.
Evaluating clinical impact requires careful studies, sometimes randomized trials, to determine whether model use actually improves outcomes. Such studies are rare.
Enemy 10: Scientific Incentives Favor Novelty
Academic incentives reward new discoveries. Publishing a new model is often easier than validating an existing one. As a result, the literature accumulates hundreds of similar models for the same clinical problem while rigorous validation studies remain scarce. The authors argue that the field would benefit from fewer models and more validation.
Enemy 11: Data Science Education Can Be Too Narrow
Many data science training programs emphasize algorithms and programming, but devote less attention to:
study design
domain knowledge
causal reasoning
clinical context
Students learn how to build models, but not always when the act of building a model is scientifically justified. Prediction modeling is not purely a computational activity; it is part of clinical research.
Enemy 12: Commercial Opacity
Healthcare AI is increasingly commercialized. Some companies treat predictive models as proprietary products, restricting access to their algorithms or training data. Without transparency, independent validation becomes difficult. In several cases where independent evaluations were conducted, commercial models performed significantly worse than advertised.
Why These Lessons Matter Now
These issues are not merely theoretical. Hospitals around the world are actively exploring how to integrate AI into clinical workflows, from triage systems to diagnostic support tools. However, deploying models without addressing the challenges described above risks introducing unreliable tools into healthcare.
The key lesson from the prediction modeling literature is simple:
The biggest risks in clinical AI rarely come from algorithms.
They come from poorly defined problems, weak datasets, and inadequate validation.
Beyond Hospitals: Consumer Health and Omics Platforms
These challenges also apply to companies offering predictive health insights based on biological measurements.
Many platforms now analyze:
genomic variants
metabolomic profiles
microbiome composition
multi-omics data
These services often promise to estimate health risks or provide personalized recommendations. Such predictions rely on models trained on datasets linking biological measurements to health outcomes. For genomic platforms, predictors are largely static. Even there, however, models must contend with population differences and uncertain interpretation of genetic variants.
For metabolomics or microbiome-based services, the challenges may be even greater. These biological signals vary over time and are influenced by diet, medications, environment, and lifestyle.
As a result, many of the “enemies” described earlier appear in this space as well:
weak signal-to-noise ratios
heterogeneous populations
uncertain labels
limited external validation
This does not mean these technologies lack promise. But it does mean that strong predictive claims require strong evidence.
What Good Prediction Model Research Should Look Like
If so many models fail, what distinguishes good prediction modeling? Several principles emerge from the literature.
First, models should start with a clear clinical question. Researchers should be able to explain what decision the model will improve and why existing approaches are insufficient.
Second, datasets should reflect the real context in which the model will be used, with representative populations and well-defined outcomes.
Third, models should follow sound methodological practices, including appropriate sample sizes and rigorous validation.
Fourth, models should be validated across multiple independent settings to ensure they generalize beyond the original dataset.
Fifth, researchers should study whether the model actually improves clinical decisions or patient outcomes.
Finally, models should be transparent and reproducible, allowing independent researchers to evaluate and improve them.
The Bigger Lesson
Prediction modeling in medicine is not primarily a computational challenge. It is a scientific and methodological challenge.
Building a model is relatively easy. Building one that is reliable, validated, and genuinely useful in practice is much harder.
As AI becomes more prominent in healthcare, the real task ahead is not simply to produce more models. It is to produce fewer, better models.
The activation of the cGAS-STING pathway is a protective mechanism of the body, primarily designed to respond to viral infections and DNA damage by triggering an immune response. This pathway plays a critical role in innate immunity, where it helps to detect and respond to pathogenic DNA within cells. The primary role of cGAS (cyclic GMP-AMP synthase) is to sense cytoplasmic DNA, which is often a sign of viral infection or cellular damage. Upon detecting DNA, cGAS produces cGAMP, which activates STING (stimulator of interferon genes), leading to the production of type I interferons. These interferons are potent antiviral molecules that help to control infections and signal neighboring cells to fortify their defenses.
While the cGAS-STING pathway serves a crucial protective function in response to acute infections and damage, its chronic activation within the brain, especially in the context of neurodegenerative diseases, can lead to harmful effects on neurons. Interferons can negatively impact neuronal transcription networks, such as those regulated by MEF2C. MEF2C is essential for neuronal health, synaptic function, and plasticity. Interferon-mediated disruption of MEF2C function can lead to synaptic loss and cognitive decline, exacerbating Alzheimer’s pathology.
Recent findings indicate that targeting this pathway could offer therapeutic benefits. Inhibiting cGAS in tauopathy models preserves synaptic integrity and cognitive function without affecting tau pathology itself, highlighting a novel approach to bolstering brain resilience against Alzheimer’s disease.
References:
Udeochu et. al. (2023), Tau activation of microglial cGAS–IFN reduces MEF2C-mediated cognitive resilience, Nature Neuroscience.
In a recent study, Biskaduros et al. (2024) examined the effects of hypertension management on Alzheimer's disease-related biomarkers in cerebrospinal fluid (CSF). Published in Alzheimer's and Dementia, this research revealed that pharmaceutical interventions in hypertension might influence the leakage of tau proteins into the CSF, potentially preventing vascular injuries associated with uncontrolled blood pressure levels.
The study longitudinally observed three groups: cognitively healthy normotensive subjects, individuals with controlled hypertension, and those with uncontrolled hypertension. Intriguingly, while CSF biomarkers like total tau (T-tau) and phospho-tau181 (P-tau 181) generally increased over time, they did not in the group with controlled hypertension. This suggests that managing hypertension could play a role in mitigating brain pathology.
However, the scenario was different for those with uncontrolled hypertension, where increases in T-tau and P-tau 181 coincided with reductions in blood pressure. This paradox might be explained by a speculative model where cumulative vascular injury renders the brain susceptible to relative hypoperfusion with blood pressure reduction. Essentially, as blood pressure decreases, certain brain areas might experience reduced perfusion, exacerbating biomarker leakage into the CSF.
These findings echo those of a 2002 trial that suggested hypertension intervention might protect against dementia development. Yet, the mechanisms remain elusive and warrant further investigation. For instance, understanding how body mass index (BMI) and LDL cholesterol levels interact with blood pressure changes could provide deeper insights into these dynamics.
In a related study by Jukkola et al. (2024) published in Fluids and Barriers of the CNS, it was shown that lowering blood pressure could enhance CSF efflux to the systemic circulation via the lymphatic drainage. This supports the hypothesis that blood pressure management might activate mechanisms that help clear CSF, potentially influencing biomarker levels.
These studies collectively highlight the complex relationship between hypertension, blood pressure management, and Alzheimer's disease biomarkers. They point to a nuanced understanding that managing hypertension could influence the progression of neurodegenerative diseases, though more research is needed to fully elucidate these relationships.
References:
Biskaduros et. al. (2024) Longitudinal trajectories of Alzheimer's disease CSF biomarkers and blood pressure in cognitively healthy subjects, Alzheimer's and Dementia
Jukkola et. al. (2024) Blood pressure lowering enhances cerebrospinal fluid efflux to the systemic circulation primarily via the lymphatic vasculature, Fluids and Barriers of the CNS
Have you ever considered the crucial role metabolism plays in our lives? Typically, when we think about metabolism, it's in the context of food, nutrients, diet, exercise, and lifestyle diseases. However, it turns out that metabolism has also shaped who we are—not just as individuals but as a species.
A dedicated community of scientists has been exploring how we evolved to have such large brains and long lifespans. They propose that many paths of human evolution are intricately linked with metabolic adaptations. For instance, research shows that "human TEE exceeded that of chimpanzees and bonobos, gorillas and orangutans by approximately 400, 635 and 820 kcal day−1, respectively, readily accommodating the cost of humans’ greater brain size and reproductive output. Much of the increase in TEE is attributable to humans’ greater basal metabolic rate (kcal day−1), indicating increased organ metabolic activity. Humans also had the greatest body fat percentage. An increased metabolic rate, along with changes in energy allocation," While it's not the sole reason, metabolism is undoubtedly a pivotal factor in shaping us as humans.
Moreover, when considering metabolism, do we often think of the brain as a metabolic organ? Typically, organs like the liver, gastrointestinal tract, and pancreas get most of the credit for metabolic activities. Yet, our brain also plays a crucial role in regulating our metabolism. It is in constant communication with our organ system, sensing signals, and coordinating the metabolic network to meet our body's needs. This brain-body interaction is crucial, as "The brain integrates internal (sensory, hormonal) and external signals and generates behaviors and physiological responses by modulating the activity of autonomic neurons. These efferent neuronal signals control metabolism as well as hormone secretion in peripheral organs."
So, the next time you think about metabolism, remember that it's not just about how your body processes a meal but also about the profound impacts it has had on our development as a species and how our brains function to keep everything in balance.
References:
Pontzer et. al. (2016), Metabolic acceleration and the evolution of human brain size and life history, Nature
Furlan et. al. (2023), Brain–body communication in metabolic control, Trends in Endocrinology & Metabolism.