Article Contents
REPORT   Open Access     Cite

From code to continuous care: Key considerations for developing and implementing dependable AI prediction models in healthcare delivery

    Show all affliationsShow less
More Information
  • DownLoad: Full size image
    1. Care delivery needs the Development, Implementation, And MONitoring for Dependable AI model framework.

      Development requires clear use cases, reliable data, standardized predictors, and rigorous validation.

      Implementation requires workflow integration, transparent governance, fairness, and accountability.

      Monitoring requires ongoing evaluation, model updates, and cost-effectiveness assessment.

      The proposed DIAMOND checklist offers a practical path from promising algorithms to dependable care.

  • Predictive models in healthcare are widely published, yet few achieve routine clinical use due to gaps in methodological rigor, workflow integration, and governance. Existing guidelines primarily focus on clinical settings, with few addressing broader healthcare delivery contexts. We propose a practical framework for translating code to continuous care: Development, Implementation, And MONitoring for Dependable AI prediction model (DIAMOND). Models must be built for explicit clinical use cases, supported by interoperable data, standardized predictors, and rigorous validation with prospective designs. Translation into practice requires workflow integration, proportionate regulatory oversight of intended use, transparency, uncertainty, bias, and accountability, and continued post-deployment evaluation—testing transportability across settings, monitoring and updating for data shift and performance degradation, and assessing health-economic impact to inform iterative refinement. By systematically linking these stages, the DIAMOND framework provides a structured pathway for advancing AI predictive models from promising algorithms to dependable clinical tools.
  • 加载中
  • [1] Wei Q., Cui M., Liu Z., et al. (2025). Integrating statistical design and inference: A roadmap for robust and trustworthy medical AI. Innov. Med. 3:100145. DOI:10.59717/j.xinn-med.2025.100145

    View in Article CrossRef Google Scholar

    [2] Hippisley-Cox J., Coupland C., Vinogradova Y., et al. (2007). Derivation and validation of QRISK, a new cardiovascular disease risk score for the United Kingdom: Prospective open cohort study. BMJ 335:136. DOI:10.1136/bmj.39261.471806.55

    View in Article CrossRef Google Scholar

    [3] Hobbs F.D. (2015). Prevention of cardiovascular diseases. BMC Med. 13:261. DOI:10.1186/s12916-015-0507-0

    View in Article CrossRef Google Scholar

    [4] Vasey B., Clifton D.A., Collins G.S., et al. (2021). DECIDE-AI: New reporting guidelines to bridge the development-to-implementation gap in clinical artificial intelligence. Nat. Med. 27:186−187. DOI:10.1038/s41591-021-01229-5

    View in Article CrossRef Google Scholar

    [5] Collins G.S., Moons K.G.M., Dhiman P., et al. (2024). TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 385:e078378. DOI:10.1136/bmj-2023-078378

    View in Article CrossRef Google Scholar

    [6] Ansari S., Baur B., Singh K., et al. (2025). Challenges in the postmarket surveillance of clinical prediction models. NEJM AI 2:AIp2401116. DOI:10.1056/AIp2401116

    View in Article CrossRef Google Scholar

    [7] Van Calster B., Collins G.S., Vickers A.J., et al. (2025). Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: Overview and guidance. Lancet Digit. Health 7:100916. DOI:10.1016/j.landig.2025.100916

    View in Article CrossRef Google Scholar

    [8] Lekadir K., Frangi A.F., Porras A.R., et al. (2025). FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ 388:e081554. DOI:10.1136/bmj-2024-081554

    View in Article CrossRef Google Scholar

    [9] Saelmans A., Seinen T., Pera V., et al. (2025). Implementation and updating of clinical prediction models: A systematic review. Mayo Clin. Proc. Digit. Health 3:100228. DOI:10.1016/j.mcpdig.2025.100228

    View in Article CrossRef Google Scholar

    [10] Efthimiou O., Seo M., Chalkou K., et al. (2024). Developing clinical prediction models: A step-by-step guide. BMJ 386:e078276. DOI:10.1136/bmj-2023-078276

    View in Article CrossRef Google Scholar

    [11] Liu X., Peng Y., Li N., et al. (2025). The necessity and feasibility assessment tool of the clinical prediction model for individual prognosis before its startup: A multi-sectoral Delphi consensus study. J. Evid. Based Med. e70106. DOI:10.1111/jebm.70106.

    View in Article Google Scholar

    [12] Zhou Y., Nie H., Gong X., et al. (2025). A three-tier AI solution for equitable glaucoma diagnosis across China’s hierarchical healthcare system. NPJ Digit. Med. 8:400. DOI:10.1038/s41746-025-01835-4

    View in Article CrossRef Google Scholar

    [13] El-Sherbini A.H., Hassan Virk H.U., Wang Z., et al. (2023). Machine-learning-based prediction modelling in primary care: State-of-the-art review. AI 4:437−460. DOI:10.3390/ai4020024

    View in Article CrossRef Google Scholar

    [14] Abbasi A.B., Curtis L.H. and Califf R.M. (2025). The promise of real-world data for research: What are we missing. N. Engl. J. Med. 393:318−321. DOI:10.1056/NEJMp2416479

    View in Article CrossRef Google Scholar

    [15] Liu H. and Ding G. (2024). Scientific wellness in China: Innovations and implementation of data- and AI-driven health. Innov. Med. 2:100103. DOI:10.59717/j.xinn-med.2024.100103

    View in Article CrossRef Google Scholar

    [16] Field M., Thwaites D.I., Carolan M., et al. (2022). Infrastructure platform for privacy-preserving distributed machine learning development of computer-assisted theragnostics in cancer. J. Biomed. Inform. 134:104181. DOI:10.1016/j.jbi.2022.104181

    View in Article CrossRef Google Scholar

    [17] Chekroud A.M., Hawrilenko M., Loho H., et al. (2024). Illusory generalizability of clinical prediction models. Science 383:164−167. DOI:10.1126/science.adg8538

    View in Article CrossRef Google Scholar

    [18] Yang J., Dung N.T., Thach P.N., et al. (2024). Generalizability assessment of AI models across hospitals in a low-middle and high-income country. Nat. Commun. 15:8270. DOI:10.1038/s41467-024-52618-6

    View in Article CrossRef Google Scholar

    [19] de Hond A.A.H., Leeuwenberg A.M., Hooft L., et al. (2022). Guidelines and quality criteria for artificial intelligence-based prediction models in healthcare: A scoping review. NPJ Digit. Med. 5:2. DOI:10.1038/s41746-021-00549-7

    View in Article CrossRef Google Scholar

    [20] Feng J., Phillips R.V., Malenica I., et al. (2022). Clinical artificial intelligence quality improvement: Towards continual monitoring and updating of AI algorithms in healthcare. NPJ Digit. Med. 5:66. DOI:10.1038/s41746-022-00611-y

    View in Article CrossRef Google Scholar

    [21] Xu H., Feng G., Ma C., et al. (2023). AMHconverter: An online tool for converting results between the different anti-Müllerian hormone assays of Roche Elecsys®, Beckman Access, and Kangrun. PeerJ 11:e15301. DOI:10.7717/peerj.15301

    View in Article CrossRef Google Scholar

    [22] Pasqualetti S., Mussap M., Monteverde E., et al. (2024). C-reactive protein and brain natriuretic peptides harmonization. Clin. Chim. Acta 562:119848. DOI:10.1016/j.cca.2024.119848

    View in Article CrossRef Google Scholar

    [23] Collins G.S., Dhiman P., Ma J., et al. (2024). Evaluation of clinical prediction models (part 1): From development to external validation. BMJ 384:e074819. DOI:10.1136/bmj-2023-074819

    View in Article CrossRef Google Scholar

    [24] Riley R.D., Archer L., Snell K.I.E., et al. (2024). Evaluation of clinical prediction models (part 2): How to undertake an external validation study. BMJ 384:e074820. DOI:10.1136/bmj-2023-074820

    View in Article CrossRef Google Scholar

    [25] Riley R.D., Snell K.I.E., Archer L., et al. (2024). Evaluation of clinical prediction models (part 3): Calculating the sample size required for an external validation study. BMJ 384:e074821. DOI:10.1136/bmj-2023-074821

    View in Article CrossRef Google Scholar

    [26] Wong A., Otles E., Donnelly J.P., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern. Med. 181:1065−1070. DOI:10.1001/jamainternmed.2021.2626

    View in Article CrossRef Google Scholar

    [27] Han R., Acosta J.N., Shakeri Z., et al. (2024). Randomised controlled trials evaluating artificial intelligence in clinical practice: A scoping review. Lancet Digit. Health 6:e367−e373. DOI:10.1016/S2589-7500(24)00047-5

    View in Article CrossRef Google Scholar

    [28] Rajkomar A., Dean J. and Kohane I. (2019). Machine learning in medicine. N. Engl. J. Med. 380:1347−1358. DOI:10.1056/NEJMra1814259

    View in Article CrossRef Google Scholar

    [29] Liu X., Faes L., Kale A.U., et al. (2019). A comparison of deep learning performance against health-care professionals in detecting diseases from medical imaging: A systematic review and meta-analysis. Lancet Digit. Health 1:e271−e297. DOI:10.1016/S2589-7500(19)30123-2

    View in Article CrossRef Google Scholar

    [30] Mayfield J.D. and Romero J. (2025). Establishing a chain of evidence for AI in radiology: Sham AI and randomized controlled trials. Radiol. Artif. Intell. 7:e250334. DOI:10.1148/ryai.250334

    View in Article CrossRef Google Scholar

    [31] Shi Z., Hu B., Lu M., et al. (2025). Development and validation of a sham-AI model for intracranial aneurysm detection at CT angiography. Radiol. Artif. Intell. 7:e240140. DOI:10.1148/ryai.240140

    View in Article CrossRef Google Scholar

    [32] Hubbard R.A., Gatsonis C.A., Hogan J.W., et al. (2024). Target trial emulation for observational studies: Potential and pitfalls. N. Engl. J. Med. 391:1975−1977. DOI:10.1056/NEJMp2407586

    View in Article CrossRef Google Scholar

    [33] Kandaswamy S., Muthu N., Braykov N., et al. (2025). Human performance evaluation of a pediatric artificial intelligence sepsis model. J. Am. Med. Inform. Assoc. 32:1552−1561. DOI:10.1093/jamia/ocaf106

    View in Article CrossRef Google Scholar

    [34] Saha S., Ross H., Velmovitsky P.E., et al. (2025). Machine learning-enhanced expert system for detecting heart failure decompensation using patient-reported vitals and electronic health records. Sci. Rep. 15:30979. DOI:10.1038/s41598-025-16376-9

    View in Article CrossRef Google Scholar

    [35] Nong P., Raj M. and Platt J. (2022). Integrating predictive models into care: Facilitating informed decision-making and communicating equity issues. Am. J. Manag. Care 28:18−24. DOI:10.37765/ajmc.2022.88812

    View in Article CrossRef Google Scholar

    [36] Amann J., Blasimme A., Vayena E., et al. (2020). Explainability for artificial intelligence in healthcare: A multidisciplinary perspective. BMC Med. Inform. Decis. Mak. 20:310. DOI:10.1186/s12911-020-01332-6

    View in Article CrossRef Google Scholar

    [37] Muralidharan V., Adewale B.A., Huang C.J., et al. (2024). A scoping review of reporting gaps in FDA-approved AI medical devices. NPJ Digit. Med. 7:273. DOI:10.1038/s41746-024-01270-x

    View in Article CrossRef Google Scholar

    [38] Singh R., Paxton M. and Auclair J. (2025). Regulating the AI-enabled ecosystem for human therapeutics. Commun. Med. 5:181. DOI:10.1038/s43856-025-00910-x

    View in Article CrossRef Google Scholar

    [39] Williams M., Karim W., Gelman J., et al. (2024). Ethical data acquisition for LLMs and AI algorithms in healthcare. NPJ Digit. Med. 7:377. DOI:10.1038/s41746-024-01399-9

    View in Article CrossRef Google Scholar

    [40] Bruns A. and Winkler E.C. (2024). Dynamic consent: A royal road to research consent? J. Med. Ethics jme-2024-110153. DOI:10.1136/jme-2024-110153.

    View in Article Google Scholar

    [41] Mascalzoni D., Melotti R., Pattaro C., et al. (2022). Ten years of dynamic consent in the CHRIS study: Informed consent as a dynamic process. Eur. J. Hum. Genet. 30:1391−1397. DOI:10.1038/s41431-022-01160-4

    View in Article CrossRef Google Scholar

    [42] Kang D.Y., DeYoung P.N., Tantiongloc J., et al. (2021). Statistical uncertainty quantification to augment clinical decision support: A first implementation in sleep medicine. NPJ Digit. Med. 4:142. DOI:10.1038/s41746-021-00515-3

    View in Article CrossRef Google Scholar

    [43] Kompa B., Snoek J. and Beam A.L. (2021). Second opinion needed: Communicating uncertainty in medical machine learning. NPJ Digit. Med. 4:4. DOI:10.1038/s41746-020-00367-3

    View in Article CrossRef Google Scholar

    [44] Tsaneva-Atanasova K., Pederzanil G. and Laviola M. (2025). Decoding uncertainty for clinical decision-making. Philos. Trans. A Math. Phys. Eng. Sci. 383:20240207. DOI:10.1098/rsta.2024.0207

    View in Article CrossRef Google Scholar

    [45] Widner K., Virmani S., Krause J., et al. (2023). Lessons learned from translating AI from development to deployment in healthcare. Nat. Med. 29:1304−1306. DOI:10.1038/s41591-023-02293-9

    View in Article CrossRef Google Scholar

    [46] Yang J., Soltan A.A.S., Eyre D.W., et al. (2023). Algorithmic fairness and bias mitigation for clinical machine learning with deep reinforcement learning. Nat. Mach. Intell. 5:884−894. DOI:10.1038/s42256-023-00697-3

    View in Article CrossRef Google Scholar

    [47] Ebad S.A., Alhashmi A., Amara M., et al. (2025). Artificial intelligence-based software as a medical device (AI-SaMD): A systematic review. Healthcare (Basel) 13. DOI:10.3390/healthcare13070817.

    View in Article Google Scholar

    [48] Wachter R.M. (2013). Personal accountability in healthcare: Searching for the right balance. BMJ Qual. Saf. 22:176−180. DOI:10.1136/bmjqs-2012-001227

    View in Article CrossRef Google Scholar

    [49] Yapps B., Shin S., Bighamian R., et al. (2017). Hypotension in ICU patients receiving vasopressor therapy. Sci. Rep. 7:8551. DOI:10.1038/s41598-017-08137-0

    View in Article CrossRef Google Scholar

    [50] Lee B., Patel S., Favorito C., et al. (2025). Development and commercialization pathways of AI medical devices in the United States: Implications for safety and regulatory oversight. NEJM AI 2:AIra2500061. DOI:10.1056/AIra2500061

    View in Article CrossRef Google Scholar

    [51] Nouis S.C., Uren V. and Jariwala S. (2025). Evaluating accountability, transparency, and bias in AI-assisted healthcare decision-making: A qualitative study of healthcare professionals’ perspectives in the UK. BMC Med. Ethics 26:89. DOI:10.1186/s12910-025-01243-z

    View in Article CrossRef Google Scholar

    [52] Habli I., Lawton T. and Porter Z. (2020). Artificial intelligence in health care: Accountability and safety. Bull. World Health Organ. 98:251−256. DOI:10.2471/blt.19.237487

    View in Article CrossRef Google Scholar

    [53] Talbert M. (2016). Moral responsibility: An introduction (John Wiley & Sons).

    View in Article Google Scholar

    [54] LeCun Y., Bengio Y. and Hinton G. (2015). Deep learning. Nature 521:436−444. DOI:10.1038/nature1459

    View in Article CrossRef Google Scholar

    [55] Cleland G.M., Sujan M.A., Habli I., et al. (2012). Evidence: Using safety cases in industry and healthcare (The Health Foundation).

    View in Article Google Scholar

    [56] Denney E., Pai G. and Habli I. (2015). Dynamic safety cases for through-life safety assurance. 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering.

    View in Article Google Scholar

    [57] Hwang T.J., Kesselheim A.S. and Vokinger K.N. (2019). Lifecycle regulation of artificial intelligence- and machine learning-based software devices in medicine. JAMA 322:2285−2286. DOI:10.1001/jama.2019.16842

    View in Article CrossRef Google Scholar

    [58] Vasey B., Nagendran M., Campbell B., et al. (2022). Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 28:924−933. DOI:10.1038/s41591-022-01772-9

    View in Article CrossRef Google Scholar

    [59] Zhang L., Wang X., Chen Q., et al. (2025). Lung cancer risk assessment by prediction model: A global perspective. Thorax 80:890−899. DOI:10.1136/thorax-2023-221253

    View in Article CrossRef Google Scholar

    [60] Tammemägi M.C., Katki H.A., Hocking W.G., et al. (2013). Selection criteria for lung-cancer screening. N. Engl. J. Med. 368:728−736. DOI:10.1056/NEJMoa1211776

    View in Article CrossRef Google Scholar

    [61] Liu X., Glocker B., McCradden M.M., et al. (2022). The medical algorithmic audit. Lancet Digit. Health 4:e384−e397. DOI:10.1016/s2589-7500(22)00003-6

    View in Article CrossRef Google Scholar

    [62] Giddings R., Joseph A., Callender T., et al. (2024). Factors influencing clinician and patient interaction with machine learning-based risk prediction models: A systematic review. Lancet Digit. Health 6:e131−e144. DOI:10.1016/s2589-7500(23)00241-8

    View in Article CrossRef Google Scholar

    [63] Kattan M.W. and Gerds T.A. (2020). A framework for the evaluation of statistical prediction models. Chest 158:S29−S38. DOI:10.1016/j.chest.2020.03.005

    View in Article CrossRef Google Scholar

    [64] Alba A.C., Agoritsas T., Walsh M., et al. (2017). Discrimination and calibration of clinical prediction models: Users’ guides to the medical literature. JAMA 318:1377−1384. DOI:10.1001/jama.2017.12126

    View in Article CrossRef Google Scholar

    [65] Pan Z., Zhang R., Shen S., et al. (2023). OWL: An optimized and independently validated machine learning prediction model for lung cancer screening based on the UK Biobank, PLCO, and NLST populations. eBioMedicine 88:104443. DOI:10.1016/j.ebiom.2023.104443

    View in Article CrossRef Google Scholar

    [66] Ye Z., Sun Y., Yin Y., et al. (2025). Assessment and recalibration of seventeen lung cancer risk prediction models in approximately one million Chinese population utilising healthcare big data: A retrospective cohort analysis. Lancet Reg. Health West. Pac. 58:101575. DOI:10.1016/j.lanwpc.2025.101575

    View in Article CrossRef Google Scholar

    [67] Feng X., Goodley P., Alcala K., et al. (2024). Evaluation of risk prediction models to select lung cancer screening participants in Europe: A prospective cohort consortium analysis. Lancet Digit. Health 6:e614−e624. DOI:10.1016/S2589-7500(24)00123-7

    View in Article CrossRef Google Scholar

    [68] Zhang T., Wang Y., Chen X., et al. (2025). Cost-effectiveness of risk model-based lung cancer screening in smokers and nonsmokers in China. BMC Med. 23:315. DOI:10.1186/s12916-025-04065-3

    View in Article CrossRef Google Scholar

    [69] Toumazis I., Cao P., de Nijs K., et al. (2023). Risk model-based lung cancer screening: A cost-effectiveness analysis. Ann. Intern. Med. 176:320−332. DOI:10.7326/m22-2216

    View in Article CrossRef Google Scholar

    [70] Hippisley-Cox J., Coupland C. and Brindle P. (2017). Development and validation of QRISK3 risk prediction algorithms to estimate future risk of cardiovascular disease: Prospective cohort study. BMJ 357:j2099. DOI:10.1136/bmj.j2099

    View in Article CrossRef Google Scholar

    [71] Hippisley-Cox J., Coupland C., Vinogradova Y., et al. (2008). Performance of the QRISK cardiovascular risk prediction algorithm in an independent UK sample of patients from general practice: A validation study. Heart 94:34−39. DOI:10.1136/hrt.2007.134890

    View in Article CrossRef Google Scholar

    [72] Hippisley-Cox J., Coupland C.A., Bafadhel M., et al. (2024). Development and validation of a new algorithm for improved cardiovascular risk prediction. Nat. Med. 30:1440−1447. DOI:10.1038/s41591-024-02905-y

    View in Article CrossRef Google Scholar

    [73] Collins G.S., Reitsma J.B., Altman D.G., et al. (2015). Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): The TRIPOD statement. Br. J. Surg. 102:148−158. DOI:10.1002/bjs.9736

    View in Article CrossRef Google Scholar

    [74] Samarasekera E.J., Clark C.E., Kaur S., et al. (2023). Cardiovascular disease risk assessment and reduction: Summary of updated NICE guidance. BMJ 381:p1028. DOI:10.1136/bmj.p1028

    View in Article CrossRef Google Scholar

  • Cite this article:

    Wei Q., Cui M., Qi Y., et al. (2026). From code to continuous care: Key considerations for developing and implementing dependable AI prediction models in healthcare delivery. The Innovation Medicine 4:100229. https://doi.org/10.59717/j.xinn-med.2026.100229
    Wei Q., Cui M., Qi Y., et al. (2026). From code to continuous care: Key considerations for developing and implementing dependable AI prediction models in healthcare delivery. The Innovation Medicine 4:100229. https://doi.org/10.59717/j.xinn-med.2026.100229

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(4)     Tables(1)

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(541) PDF downloads(281)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint