Care delivery needs the Development, Implementation, And MONitoring for Dependable AI model framework.
Development requires clear use cases, reliable data, standardized predictors, and rigorous validation.
Implementation requires workflow integration, transparent governance, fairness, and accountability.
Monitoring requires ongoing evaluation, model updates, and cost-effectiveness assessment.
The proposed DIAMOND checklist offers a practical path from promising algorithms to dependable care.
| [1] | Wei Q., Cui M., Liu Z., et al. (2025). Integrating statistical design and inference: A roadmap for robust and trustworthy medical AI. Innov. Med. 3:100145. DOI:10.59717/j.xinn-med.2025.100145 |
| [2] | Hippisley-Cox J., Coupland C., Vinogradova Y., et al. (2007). Derivation and validation of QRISK, a new cardiovascular disease risk score for the United Kingdom: Prospective open cohort study. BMJ 335:136. DOI:10.1136/bmj.39261.471806.55 |
| [3] | Hobbs F.D. (2015). Prevention of cardiovascular diseases. BMC Med. 13:261. DOI:10.1186/s12916-015-0507-0 |
| [4] | Vasey B., Clifton D.A., Collins G.S., et al. (2021). DECIDE-AI: New reporting guidelines to bridge the development-to-implementation gap in clinical artificial intelligence. Nat. Med. 27:186−187. DOI:10.1038/s41591-021-01229-5 |
| [5] | Collins G.S., Moons K.G.M., Dhiman P., et al. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 385:e078378. DOI:10.1136/bmj-2023-078378 |
| [6] | Ansari S., Baur B., Singh K., et al. (2025). Challenges in the postmarket surveillance of clinical prediction models. NEJM AI 2:AIp2401116. DOI:10.1056/AIp2401116 |
| [7] | Van Calster B., Collins G.S., Vickers A.J., et al. (2025). Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: Overview and guidance. Lancet Digit. Health 7:100916. DOI:10.1016/j.landig.2025.100916 |
| [8] | Lekadir K., Frangi A.F., Porras A.R., et al. (2025). FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ 388:e081554. DOI:10.1136/bmj-2024-081554 |
| [9] | Saelmans A., Seinen T., Pera V., et al. (2025). Implementation and updating of clinical prediction models: A systematic review. Mayo Clin. Proc. Digit. Health 3:100228. DOI:10.1016/j.mcpdig.2025.100228 |
| [10] | Efthimiou O., Seo M., Chalkou K., et al. (2024). Developing clinical prediction models: A step-by-step guide. BMJ 386:e078276. DOI:10.1136/bmj-2023-078276 |
| [11] | Liu X., Peng Y., Li N., et al. (2025). The necessity and feasibility assessment tool of the clinical prediction model for individual prognosis before its startup: A multi-sectoral Delphi consensus study. J. Evid. Based Med. e70106. DOI:10.1111/jebm.70106. |
| [12] | Zhou Y., Nie H., Gong X., et al. (2025). A three-tier AI solution for equitable glaucoma diagnosis across China’s hierarchical healthcare system. NPJ Digit. Med. 8:400. DOI:10.1038/s41746-025-01835-4 |
| [13] | El-Sherbini A.H., Hassan Virk H.U., Wang Z., et al. (2023). Machine-learning-based prediction modelling in primary care: State-of-the-art review. AI 4:437−460. DOI:10.3390/ai4020024 |
| [14] | Abbasi A.B., Curtis L.H. and Califf R.M. (2025). The promise of real-world data for research: What are we missing. N. Engl. J. Med. 393:318−321. DOI:10.1056/NEJMp2416479 |
| [15] | Liu H. and Ding G. (2024). Scientific wellness in China: Innovations and implementation of data- and AI-driven health. Innov. Med. 2:100103. DOI:10.59717/j.xinn-med.2024.100103 |
| [16] | Field M., Thwaites D.I., Carolan M., et al. (2022). Infrastructure platform for privacy-preserving distributed machine learning development of computer-assisted theragnostics in cancer. J. Biomed. Inform. 134:104181. DOI:10.1016/j.jbi.2022.104181 |
| [17] | Chekroud A.M., Hawrilenko M., Loho H., et al. (2024). Illusory generalizability of clinical prediction models. Science 383:164−167. DOI:10.1126/science.adg8538 |
| [18] | Yang J., Dung N.T., Thach P.N., et al. (2024). Generalizability assessment of AI models across hospitals in a low-middle and high-income country. Nat. Commun. 15:8270. DOI:10.1038/s41467-024-52618-6 |
| [19] | de Hond A.A.H., Leeuwenberg A.M., Hooft L., et al. (2022). Guidelines and quality criteria for artificial intelligence-based prediction models in healthcare: A scoping review. NPJ Digit. Med. 5:2. DOI:10.1038/s41746-021-00549-7 |
| [20] | Feng J., Phillips R.V., Malenica I., et al. (2022). Clinical artificial intelligence quality improvement: Towards continual monitoring and updating of AI algorithms in healthcare. NPJ Digit. Med. 5:66. DOI:10.1038/s41746-022-00611-y |
| [21] | Xu H., Feng G., Ma C., et al. (2023). AMHconverter: An online tool for converting results between the different anti-Müllerian hormone assays of Roche Elecsys®, Beckman Access, and Kangrun. PeerJ 11:e15301. DOI:10.7717/peerj.15301 |
| [22] | Pasqualetti S., Mussap M., Monteverde E., et al. (2024). C-reactive protein and brain natriuretic peptides harmonization. Clin. Chim. Acta 562:119848. DOI:10.1016/j.cca.2024.119848 |
| [23] | Collins G.S., Dhiman P., Ma J., et al. (2024). Evaluation of clinical prediction models (part 1): From development to external validation. BMJ 384:e074819. DOI:10.1136/bmj-2023-074819 |
| [24] | Riley R.D., Archer L., Snell K.I.E., et al. (2024). Evaluation of clinical prediction models (part 2): How to undertake an external validation study. BMJ 384:e074820. DOI:10.1136/bmj-2023-074820 |
| [25] | Riley R.D., Snell K.I.E., Archer L., et al. (2024). Evaluation of clinical prediction models (part 3): Calculating the sample size required for an external validation study. BMJ 384:e074821. DOI:10.1136/bmj-2023-074821 |
| [26] | Wong A., Otles E., Donnelly J.P., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern. Med. 181:1065−1070. DOI:10.1001/jamainternmed.2021.2626 |
| [27] | Han R., Acosta J.N., Shakeri Z., et al. (2024). Randomised controlled trials evaluating artificial intelligence in clinical practice: A scoping review. Lancet Digit. Health 6:e367−e373. DOI:10.1016/S2589-7500(24)00047-5 |
| [28] | Rajkomar A., Dean J. and Kohane I. (2019). Machine learning in medicine. N. Engl. J. Med. 380:1347−1358. DOI:10.1056/NEJMra1814259 |
| [29] | Liu X., Faes L., Kale A.U., et al. (2019). A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: A systematic review and meta-analysis. Lancet Digit. Health 1:e271−e297. DOI:10.1016/S2589-7500(19)30123-2 |
| [30] | Mayfield J.D. and Romero J. (2025). Establishing a chain of evidence for AI in radiology: Sham AI and randomized controlled trials. Radiol. Artif. Intell. 7:e250334. DOI:10.1148/ryai.250334 |
| [31] | Shi Z., Hu B., Lu M., et al. (2025). Development and validation of a sham-AI model for intracranial aneurysm detection at CT angiography. Radiol. Artif. Intell. 7:e240140. DOI:10.1148/ryai.240140 |
| [32] | Hubbard R.A., Gatsonis C.A., Hogan J.W., et al. (2024). Target trial emulation for observational studies: Potential and pitfalls. N. Engl. J. Med. 391:1975−1977. DOI:10.1056/NEJMp2407586 |
| [33] | Kandaswamy S., Muthu N., Braykov N., et al. (2025). Human performance evaluation of a pediatric artificial intelligence sepsis model. J. Am. Med. Inform. Assoc. 32:1552−1561. DOI:10.1093/jamia/ocaf106 |
| [34] | Saha S., Ross H., Velmovitsky P.E., et al. (2025). Machine learning-enhanced expert system for detecting heart failure decompensation using patient-reported vitals and electronic health records. Sci. Rep. 15:30979. DOI:10.1038/s41598-025-16376-9 |
| [35] | Nong P., Raj M. and Platt J. (2022). Integrating predictive models into care: Facilitating informed decision-making and communicating equity issues. Am. J. Manag. Care 28:18−24. DOI:10.37765/ajmc.2022.88812 |
| [36] | Amann J., Blasimme A., Vayena E., et al. (2020). Explainability for artificial intelligence in healthcare: A multidisciplinary perspective. BMC Med. Inform. Decis. Mak. 20:310. DOI:10.1186/s12911-020-01332-6 |
| [37] | Muralidharan V., Adewale B.A., Huang C.J., et al. (2024). A scoping review of reporting gaps in FDA-approved AI medical devices. NPJ Digit. Med. 7:273. DOI:10.1038/s41746-024-01270-x |
| [38] | Singh R., Paxton M. and Auclair J. (2025). Regulating the AI-enabled ecosystem for human therapeutics. Commun. Med. 5:181. DOI:10.1038/s43856-025-00910-x |
| [39] | Williams M., Karim W., Gelman J., et al. (2024). Ethical data acquisition for LLMs and AI algorithms in healthcare. NPJ Digit. Med. 7:377. DOI:10.1038/s41746-024-01399-9 |
| [40] | Bruns A. and Winkler E.C. (2024). Dynamic consent: A royal road to research consent? J. Med. Ethics jme-2024-110153. DOI:10.1136/jme-2024-110153. |
| [41] | Mascalzoni D., Melotti R., Pattaro C., et al. (2022). Ten years of dynamic consent in the CHRIS study: Informed consent as a dynamic process. Eur. J. Hum. Genet. 30:1391−1397. DOI:10.1038/s41431-022-01160-4 |
| [42] | Kang D.Y., DeYoung P.N., Tantiongloc J., et al. (2021). Statistical uncertainty quantification to augment clinical decision support: A first implementation in sleep medicine. NPJ Digit. Med. 4:142. DOI:10.1038/s41746-021-00515-3 |
| [43] | Kompa B., Snoek J. and Beam A.L. (2021). Second opinion needed: Communicating uncertainty in medical machine learning. NPJ Digit. Med. 4:4. DOI:10.1038/s41746-020-00367-3 |
| [44] | Tsaneva-Atanasova K., Pederzanil G. and Laviola M. (2025). Decoding uncertainty for clinical decision-making. Philos. Trans. A Math. Phys. Eng. Sci. 383:20240207. DOI:10.1098/rsta.2024.0207 |
| [45] | Widner K., Virmani S., Krause J., et al. (2023). Lessons learned from translating AI from development to deployment in healthcare. Nat. Med. 29:1304−1306. DOI:10.1038/s41591-023-02293-9 |
| [46] | Yang J., Soltan A.A.S., Eyre D.W., et al. (2023). Algorithmic fairness and bias mitigation for clinical machine learning with deep reinforcement learning. Nat. Mach. Intell. 5:884−894. DOI:10.1038/s42256-023-00697-3 |
| [47] | Ebad S.A., Alhashmi A., Amara M., et al. (2025). Artificial intelligence-based software as a medical device (AI-SaMD): A systematic review. Healthcare (Basel) 13. DOI:10.3390/healthcare13070817. |
| [48] | Wachter R.M. (2013). Personal accountability in healthcare: Searching for the right balance. BMJ Qual. Saf. 22:176−180. DOI:10.1136/bmjqs-2012-001227 |
| [49] | Yapps B., Shin S., Bighamian R., et al. (2017). Hypotension in ICU patients receiving vasopressor therapy. Sci. Rep. 7:8551. DOI:10.1038/s41598-017-08137-0 |
| [50] | Lee B., Patel S., Favorito C., et al. (2025). Development and commercialization pathways of AI medical devices in the United States: Implications for safety and regulatory oversight. NEJM AI 2:AIra2500061. DOI:10.1056/AIra2500061 |
| [51] | Nouis S.C., Uren V. and Jariwala S. (2025). Evaluating accountability, transparency, and bias in AI-assisted healthcare decision-making: A qualitative study of healthcare professionals’ perspectives in the UK. BMC Med. Ethics 26:89. DOI:10.1186/s12910-025-01243-z |
| [52] | Habli I., Lawton T. and Porter Z. (2020). Artificial intelligence in health care: Accountability and safety. Bull. World Health Organ. 98:251−256. DOI:10.2471/blt.19.237487 |
| [53] | Talbert M. (2016). Moral responsibility: An introduction (John Wiley & Sons). |
| [54] | LeCun Y., Bengio Y. and Hinton G. (2015). Deep learning. Nature 521:436−444. DOI:10.1038/nature1459 |
| [55] | Cleland G.M., Sujan M.A., Habli I., et al. (2012). Evidence: Using safety cases in industry and healthcare (The Health Foundation). |
| [56] | Denney E., Pai G. and Habli I. (2015). Dynamic safety cases for through-life safety assurance. 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering. |
| [57] | Hwang T.J., Kesselheim A.S. and Vokinger K.N. (2019). Lifecycle regulation of artificial intelligence- and machine learning-based software devices in medicine. JAMA 322:2285−2286. DOI:10.1001/jama.2019.16842 |
| [58] | Vasey B., Nagendran M., Campbell B., et al. (2022). Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 28:924−933. DOI:10.1038/s41591-022-01772-9 |
| [59] | Zhang L., Wang X., Chen Q., et al. (2025). Lung cancer risk assessment by prediction model: A global perspective. Thorax 80:890−899. DOI:10.1136/thorax-2023-221253 |
| [60] | Tammemägi M.C., Katki H.A., Hocking W.G., et al. (2013). Selection criteria for lung-cancer screening. N. Engl. J. Med. 368:728−736. DOI:10.1056/NEJMoa1211776 |
| [61] | Liu X., Glocker B., McCradden M.M., et al. (2022). The medical algorithmic audit. Lancet Digit. Health 4:e384−e397. DOI:10.1016/s2589-7500(22)00003-6 |
| [62] | Giddings R., Joseph A., Callender T., et al. (2024). Factors influencing clinician and patient interaction with machine learning-based risk prediction models: A systematic review. Lancet Digit. Health 6:e131−e144. DOI:10.1016/s2589-7500(23)00241-8 |
| [63] | Kattan M.W. and Gerds T.A. (2020). A framework for the evaluation of statistical prediction models. Chest 158:S29−S38. DOI:10.1016/j.chest.2020.03.005 |
| [64] | Alba A.C., Agoritsas T., Walsh M., et al. (2017). Discrimination and calibration of clinical prediction models: Users’ guides to the medical literature. JAMA 318:1377−1384. DOI:10.1001/jama.2017.12126 |
| [65] | Pan Z., Zhang R., Shen S., et al. (2023). OWL: An optimized and independently validated machine learning prediction model for lung cancer screening based on the UK Biobank, PLCO, and NLST populations. eBioMedicine 88:104443. DOI:10.1016/j.ebiom.2023.104443 |
| [66] | Ye Z., Sun Y., Yin Y., et al. (2025). Assessment and recalibration of seventeen lung cancer risk prediction models in approximately one million Chinese population utilising healthcare big data: A retrospective cohort analysis. Lancet Reg. Health West. Pac. 58:101575. DOI:10.1016/j.lanwpc.2025.101575 |
| [67] | Feng X., Goodley P., Alcala K., et al. (2024). Evaluation of risk prediction models to select lung cancer screening participants in Europe: A prospective cohort consortium analysis. Lancet Digit. Health 6:e614−e624. DOI:10.1016/S2589-7500(24)00123-7 |
| [68] | Zhang T., Wang Y., Chen X., et al. (2025). Cost-effectiveness of risk model-based lung cancer screening in smokers and nonsmokers in China. BMC Med. 23:315. DOI:10.1186/s12916-025-04065-3 |
| [69] | Toumazis I., Cao P., de Nijs K., et al. (2023). Risk model-based lung cancer screening: A cost-effectiveness analysis. Ann. Intern. Med. 176:320−332. DOI:10.7326/m22-2216 |
| [70] | Hippisley-Cox J., Coupland C. and Brindle P. (2017). Development and validation of QRISK3 risk prediction algorithms to estimate future risk of cardiovascular disease: Prospective cohort study. BMJ 357:j2099. DOI:10.1136/bmj.j2099 |
| [71] | Hippisley-Cox J., Coupland C., Vinogradova Y., et al. (2008). Performance of the QRISK cardiovascular risk prediction algorithm in an independent UK sample of patients from general practice: A validation study. Heart 94:34−39. DOI:10.1136/hrt.2007.134890 |
| [72] | Hippisley-Cox J., Coupland C.A., Bafadhel M., et al. (2024). Development and validation of a new algorithm for improved cardiovascular risk prediction. Nat. Med. 30:1440−1447. DOI:10.1038/s41591-024-02905-y |
| [73] | Collins G.S., Reitsma J.B., Altman D.G., et al. (2015). Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): The TRIPOD statement. Br. J. Surg. 102:148−158. DOI:10.1002/bjs.9736 |
| [74] | Samarasekera E.J., Clark C.E., Kaur S., et al. (2023). Cardiovascular disease risk assessment and reduction: Summary of updated NICE guidance. BMJ 381:p1028. DOI:10.1136/bmj.p1028 |
| Wei Q., Cui M., Qi Y., et al. (2026). From code to continuous care: Key considerations for developing and implementing dependable AI prediction models in healthcare delivery. The Innovation Medicine 4:100229. https://doi.org/10.59717/j.xinn-med.2026.100229 |
To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.
Hierarchical evidence strength pyramid for clinical prediction model validation.
Governance structure and responsibilities of key stakeholders in clinical prediction model
Cross-regional divergence in disease incidence trend — lung cancer as example
DIAMOND checklist: key considerations of prediction model development, implementation, and monitoring after implementation.