AI Bias and Fairness

© Fraunhofer MEVIS / chest X-ray images courtesy of the NIH CXR8 database

Does Your AI Work for Every Patient? 

Medical AI systems are increasingly used to support clinical decisions, from triaging radiology findings to predicting patient risk. But a system that performs well on average may still fail systematically for specific patient groups: patients of different sexes or genders, elderly or younger patients, individuals of different skin tones or ethnic backgrounds, or patients from underrepresented clinical sites. These disparities are a fundamental challenge to the safe and equitable deployment of AI in healthcare; addressing requires an intentional and methodologically rigorous approach.

 

Gender-Sensitive Medicine and Broader Equity 

Historically, medical research — and the clinical data on which AI models are trained — has underrepresented non-male patients, older patients, and populations outside Western Europe and North America. AI systems trained on such data risk perpetuating or amplifying these biases. Gender-sensitive medicine is increasingly recognized as a critical dimension of clinical quality, a standard that AI systems must meet alongside their human counterparts.

Reliable performance across sexes and genders, skin tones, clinical sites, scanner vendors, and age groups is the mark of a well-engineered medical AI product.

What Is AI Bias, and Why Is It Difficult to Measure? 

An AI system can be biased in many different senses. A model may achieve equal accuracy across groups while producing very different true- and false-negative rates, with serious clinical consequences. Conversely, enforcing equal error rates may require the model to treat groups differently in other respects. Different notions of fairness are mathematically incompatible in general, and choosing the right one requires understanding the clinical context and the downstream effects of model errors.

Correctly attributing a detected performance gap to actual model bias — rather than to legitimate confounders, differences in disease prevalence, or data quality disparities — demands rigorous statistical methodology. Deconfounding, subgroup-stratified evaluation, propensity weighting, and appropriate metric selection are all highly non-trivial in practice, and the wrong choices can lead to misleading conclusions in either direction.

 

From Detection to Mitigation via Root Cause Analysis 

Identifying a bias is the beginning, not the end. Once a performance disparity is characterized and understood, the appropriate response depends on its root cause: why is the model under-performing in this population? Answering this question often requires a substantial amount of experimental and analytical detective work, and the correct mitigation approach will depend on this answer. Targeted data augmentation or preprocessing, resampling or reweighting strategies, fairness-aware training objectives, or recalibration for specific subpopulations may all be relevant, and the most appropriate approach depends on the root cause of the observed bias. Our experts can help you uncover the root causes of observed performance disparities, implement and validate appropriate mitigation strategies, and then re-evaluate rigorously to confirm that the intervention worked without introducing new problems elsewhere.

This iterative evaluation → mitigation → re-evaluation loop is what enables developing truly robust models.

 

Regulatory Requirements and the EU AI Act 

Medical AI systems classified as high-risk under the EU AI Act are subject to explicit requirements for bias testing, fairness documentation, and ongoing monitoring. This includes assessing model performance across relevant patient subgroups and establishing processes for post-market surveillance to detect emerging bias over time.

Meeting these requirements demands a structured, methodologically sound evaluation framework that can withstand regulatory scrutiny. Our experts at Fraunhofer MEVIS have worked on precisely such approaches, processes, and tooling for years and can give you a head start.

 

© Fraunhofer MEVIS / breast MRI images adapted from the ODELIA Challenge dataset, provided by the ODELIA consortium, licensed under CC-BY-NC-4.0

© Fraunhofer MEVIS

Highlight Publications 

  • SEG4SEG: Identifying Systematic Failure Modes in Segmentation by Subgroup Discovery Methods. Weng, Petersen, et al., 2026. Article 
  • meval: A Statistical Toolbox for Fine-Grained Model Performance Analysis. Sutariya & Petersen, 2025. Article, Toolbox
  • The cause and effect of an MR image: Robustness and generalizability. Feragen, Petersen, Ganz-Benjaminsen, 2025. Book chapter
  • Slicing Through Bias: Explaining Performance Gaps in Medical Image Analysis Using Slice Discovery Methods. Olesen, Weng, Feragen, Petersen, 2024. Article
  • The path toward equal performance in medical machine learning. Petersen et al., 2023. Article
  • Are Sex-Based Physiological Differences the Cause of Gender Bias for Chest X-Ray Diagnosis? Weng, Bigdeli, Petersen, Feragen, 2023. Article
  • On (assessing) the fairness of risk score models. Petersen, Ganz, Holm, Feragen, 2023. Article
  • Responsible and Regulatory Conform Machine Learning for Medicine: A Survey of Challenges and Solutions. Petersen et al. 2022. Article

Our Expertise 

Our team has deep expertise in bias evaluation methodology for medical AI, spanning the full analytical pipeline: defining appropriate fairness criteria for the clinical use case, designing statistically robust evaluation studies, selecting and implementing suitable metrics, correcting for selection and confounding biases in observational data, interpreting results in a clinically meaningful way, and mapping outcomes onto regulatory validation and post-market surveillance reporting requirements.

Our open-source meval toolbox — developed and maintained at Fraunhofer MEVIS — provides a rigorous statistical foundation for fine-grained subgroup performance analysis, including studentized permutation testing, confidence interval computations, and correction for multiple comparisons. It is designed to support the kind of transparent, reproducible evaluation that is required for good scientific practice and regulatory compliance alike. → meval on GitHub

In addition, we are co-organizers of the international FAIMI initiative (Fairness of AI in Medical Imaging), which brings together the research community to advance methodology and awareness in this area. → faimi.org

 

Our Offer 

We partner with medtech companies, clinical institutions, and AI developers to conduct rigorous bias and fairness evaluations, whether it is purely for research purposes, to improve general model robustness, or for regulatory submissions under EU AI Act, MDR, or FDA guidance. We identify root causes of detected performance disparities and develop targeted mitigation strategies; we implement and validate post-market surveillance frameworks for ongoing bias monitoring; and we advise on fairness metric selection, evaluation study design, and regulatory documentation.

Learn more about the services we offer

© Fraunhofer MEVIS