Blog · ppj.es
Synthetic medical data: the FDA proposes 7 criteria to use them without fooling ourselves
Published on 19 August 2026
🇪🇸 Leer este artículo en español
Title: Synthetic medical data: the FDA proposes 7 criteria to use them without fooling ourselves
Type: C (AI / data governance / LOPD)
Tags: synthetic data, medical AI, FDA, privacy, evaluation, mammography
PMID: 42602824
DOI: 10.1093/bjrai/ubag005
Journal: BJR Artificial Intelligence (2026) — FDA (Division of Imaging, Diagnostics, and Software Reliability)
Category: Oncology / AI
⚖️ Transparency notice: this article was written with AI assistance and reviewed by the author, a medical oncologist.
The problem nobody wants to admit
Synthetic medical data (SMDs) — mammograms, scans, notes — generated by AI are sold as the solution to data scarcity and privacy. But nobody had defined how to evaluate their quality. The FDA (Zamzmi et al., BJR AI 2026) proposes a framework of 7 criteria: Congruence, Coverage, Constraint, Consistency, Comprehension, Compliance, Completeness [PMID 42602824].
The 7 criteria in practice
| Criterion | What it measures |
|---|---|
| Congruence | Statistical fidelity vs real data |
| Coverage | Representation of subgroups (breast density, age) |
| Constraint | Respects physical/anatomical rules |
| Consistency | Internal coherence (same patient, multiple views) |
| Comprehension | Interpretable by humans/clinicians |
| Compliance | Complies with privacy/regulation (LOPD/GDPR) |
| Completeness | No missing critical dimensions |
Key finding
Applied to 7 digital mammography datasets: those generated by generative AI had high Congruence but failed in Constraint and Completeness vs knowledge-based methods. And quality varied by subgroup — dense/heterogeneous breasts showed poorer quality. This is critical: a synthetic dataset that works on fatty breasts may fail on dense ones [PMID 42602824].
Critical reading
- It is a framework, not a binding standard. Useful as a checklist, but the FDA does not enforce it yet.
- Silent subgroup bias. The finding that dense breasts come out worse is a real warning: models trained on synthetics may amplify disparities.
- Relevance for your hospital AI Dossier: if the hospital uses synthetic data to train or validate AI (MARTE, COG-LLM), this framework is the reference to audit them before using them with real patients.
💡 Implications for the consultation / AI Dossier
The Dossier you prepare for IT/DPO must cite such frameworks: synthetic data ≠ automatically secure private data. Compliance (LOPD) and Completeness (subgroups) are the points where a hospital AI project can sink in an audit. This paper provides the language to demand quality in synthetic data.
Reference: Zamzmi G, et al. BJR Artif Intell. 2026;3(1):ubag005. doi:10.1093/bjrai/ubag005. PMID: 42602824.
— This analysis was generated by ANGIE (Always Next to Guide, Inspire and Empower), an artificial intelligence system with SOUL profiles, designed by Dr. Javier Pumares Pérez.
Disclaimer: this article is educational and informational in nature and reflects the personal opinion of the author. It does not constitute medical advice nor replace the assessment of a healthcare professional. If you have a health concern, consult your physician.
