Reference health

A Preliminary Checklist (METRICS) to Standardize the Design and Reporting of Studies on Generative Artificial Intelligence–Based Models in Health Care Education and Practice: Development Study Involving a Literature Review

https://doi.org/10.2196/54704
CiteStamped reference-health badge
94/94 checkable references clean · checked 2026-08-16

Every reference with a DOI in the deposited reference list resolved to a known work in Crossref or DataCite at the dated check, and none carried a retraction, withdrawal, or removal notice.

6 without a DOI — not checked. A reference deposited without a DOI is never matched by title or guessed at; it stays outside the checked set, and this line discloses that.

The 94 checked references that resolve
resolves10.3390/healthcare11060887
ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns
resolves10.34172/hpp.2023.22
Exploring the role of ChatGPT in patient care (diagnosis and treatment) and medical research: A systematic review
resolves10.3389/fmed.2023.1279707
Integrating AI in medical education: embracing ethical usage and critical understanding
resolves10.1016/j.cmpb.2024.108013
ChatGPT in healthcare: A taxonomy and systematic review
resolves10.2196/48392
The ChatGPT (Generative Artificial Intelligence) Revolution Has Made Artificial Intelligence Approachable for Medical Professionals
resolves10.37074/jalt.2023.6.1.23
War of the chatbots: Bard, Bing Chat, ChatGPT, Ernie and beyond. The new AI gold rush and its impact on higher education
resolves10.2196/47184
Investigating the Impact of User Trust on the Adoption and Use of ChatGPT: Survey Analysis
resolves10.2196/47564
User Intentions to Use ChatGPT for Self-Diagnosis and Health-Related Purposes: Cross-sectional Survey Study
resolves10.7861/fhj.2021-0095
Artificial intelligence in healthcare: transforming the practice of medicine
resolves10.2196/48659
Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow: Development and Usability Study
resolves10.2196/47737
Performance of ChatGPT on UK Standardized Admission Tests: Insights From the BMAT, TMUA, LNAT, and TSA Examinations
resolves10.1186/s12909-023-04698-z
Revolutionizing healthcare: the role of artificial intelligence in clinical practice
resolves10.2196/49963
A Future of Smarter Digital Health Empowered by Generative Pretrained Transformer
resolves10.3389/fpubh.2021.755808
A Framework of AI-Based Approaches to Improving eHealth Literacy and Combating Infodemic
resolves10.2196/48433
Examining Real-World Medication Consultations and Drug-Herb Interactions: ChatGPT Performance Evaluation
resolves10.1057/s41599-023-02079-x
Ethics and discrimination in artificial intelligence-enabled recruitment practices
resolves10.2196/48009
Ethical Considerations of Using ChatGPT in Health Care
resolves10.1038/s41537-023-00379-4
ChatGPT: these are not hallucinations – they’re fabrications and falsifications
resolves10.2196/49368
A SWOT (Strengths, Weaknesses, Opportunities, and Threats) Analysis of ChatGPT in the Medical Literature: Concise Review
resolves10.2967/jnumed.123.265687
An Opinion on ChatGPT in Health Care—Written by Humans Only
resolves10.1186/s12910-021-00687-3
Privacy and artificial intelligence: challenges for protecting health information in a new era
resolves10.58496/mjcs/2023/004
ChatGPT: Exploring the Role of Cybersecurity in the Protection of Medical Information
resolves10.2196/45312
How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment
resolves10.7759/cureus.35029
ChatGPT Output Regarding Compulsory Vaccination and COVID-19 Vaccine Conspiracy: A Descriptive Study at the Outset of a Paradigm Shift in Online Search for Information
resolves10.2196/49877
ChatGPT Interactive Medical Simulations for Early Clinical Education: Case Study
resolves10.1016/j.iotcps.2023.04.003
ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope
resolves10.3390/clinpract13050104
AI-Powered Renal Diet Support: Performance of ChatGPT, Bard AI, and Bing Chat
resolves10.7759/cureus.45473
Efficacy of AI Chats to Determine an Emergency: A Comparison Between OpenAI’s ChatGPT, Google Bard, and Microsoft Bing AI Chat
resolves10.1016/j.resuscitation.2023.110009
Can novel multimodal chatbots such as Bing Chat Enterprise, ChatGPT-4 Pro, and Google Bard correctly interpret electrocardiogram images?
resolves10.3390/jpm13101457
Navigating the Landscape of Personalized Medicine: The Relevance of ChatGPT, BingChat, and Bard AI in Nephrology Literature Searches
resolves10.1016/j.tmrv.2023.150753
Battle of the (Chat)Bots: Comparing Large Language Models to Practice Guidelines for Transfusion-Associated Graft-Versus-Host Disease Prevention
resolves10.2196/46885
The Role of ChatGPT, Generative Language Models, and Artificial Intelligence in Medical Education: A Conversation With ChatGPT and a Call for Papers
resolves10.2196/50638
Prompt Engineering as an Important Emerging Skill for Medical Professionals: Tutorial
resolves10.2196/48966
ChatGPT vs Google for Queries Related to Dementia and Other Cognitive Decline: Comparison of Results
resolves10.1136/bmj.n71
The PRISMA 2020 statement: an updated guideline for reporting systematic reviews
resolves10.1136/bmj.39335.541782.AD
Strengthening the reporting of observational studies in epidemiology (STROBE) statement: guidelines for reporting observational studies
resolves10.2147/DHPS.S425858
Evaluating the Sensitivity, Specificity, and Accuracy of ChatGPT-3.5, ChatGPT-4, Bing AI, and Bard Against Conventional Drug-Drug Interactions Clinical Tools
resolves10.1007/s10439-023-03338-3
Sailing the Seven Seas: A Multinational Comparison of ChatGPT’s Performance on Medical Licensing Examinations
resolves10.1111/eje.12937
ChatGPT—A double‐edged sword for healthcare education? Implications for assessments of dental students
resolves10.7759/cureus.45043
ChatGPT Conquers the Saudi Medical Licensing Exam: Exploring the Accuracy of Artificial Intelligence in Medical Knowledge Assessment and Implications for Modern Medical Education
resolves10.7759/cureus.40351
Snakebite Advice and Counseling From Artificial Intelligence: An Acute Venomous Snakebite Consultation With ChatGPT
resolves10.2196/51421
Exploring the Possible Use of AI Chatbots in Public Health Education: Feasibility Study
resolves10.1111/opo.13207
Assessing the utility of ChatGPT as an artificial intelligence‐based large language model for information to answer questions on myopia
resolves10.1136/bmjno-2023-000530
Assessment of ChatGPT’s performance on neurology written board examination questions
resolves10.3390/vaccines11071217
Artificial Intelligence and Public Health: Evaluating ChatGPT Responses to Vaccination Myths and Misconceptions
resolves10.7759/cureus.37023
Evaluating ChatGPT's Ability to Solve Higher-Order Questions on the Competency-Based Medical Education Curriculum in Medical Biochemistry
resolves10.1136/bmjno-2023-000451
Evaluating the limits of AI in medical specialisation: ChatGPT’s performance on the UK Neurology Specialty Certificate Examination
resolves10.1590/1806-9282.20230848
Performance of ChatGPT-4 in answering questions from the Brazilian National Examination for Medical Degree Revalidation
resolves10.7759/cureus.40135
Radiology Gets Chatty: The ChatGPT Saga Unfolds
resolves10.1016/j.wneu.2023.08.042
GPT-4 Artificial Intelligence Model Outperforms ChatGPT, Medical Students, and Neurosurgery Residents on Neurosurgery Written Board-Like Questions
resolves10.7759/cureus.38784
Exploring ChatGPT’s Potential in Facilitating Adaptation of Clinical Guidelines: A Case Study of Diabetic Ketoacidosis Guidelines
resolves10.1007/s00405-023-08051-4
ChatGPT’s quiz skills in different otolaryngology subspecialties: an analysis of 2576 single-choice and multiple-choice board certification preparation questions
resolves10.7759/cureus.36272
The Capability of ChatGPT in Predicting and Explaining Common Drug-Drug Interactions
resolves10.1097/JS9.0000000000000571
ChatGPT encounters multiple opportunities and challenges in neurosurgery
resolves10.7759/cureus.43861
Large Language Models in Hematology Case Solving: A Comparative Study of ChatGPT-3.5, Google Bard, and Microsoft Bing
resolves10.2106/JBJS.OA.23.00056
Evaluating ChatGPT Performance on the Orthopaedic In-Training Examination
resolves10.3389/fmed.2023.1240915
Evaluating the performance of ChatGPT-4 on the United Kingdom Medical Licensing Assessment
resolves10.1186/s42492-023-00136-5
Translating radiology reports into plain language using ChatGPT and GPT-4 with prompt learning: results, limitations, and potential
resolves10.3390/children10101634
Can ChatGPT Guide Parents on Tympanostomy Tube Insertion?
resolves10.7759/cureus.45911
Bias and Inaccuracy in AI Chatbot Ophthalmologist Recommendations
resolves10.1097/MD.0000000000034673
ChatGPT performance in the medical specialty exam: An observational study
resolves10.1016/j.cgh.2023.08.033
Accuracy, Reliability, and Comprehensibility of ChatGPT-Generated Medical Responses for Patients With Nonalcoholic Fatty Liver Disease
resolves10.52225/narra.v3i1.103
ChatGPT applications in medical, dental, pharmacy, and public health education: A descriptive study highlighting the advantages and limitations
resolves10.1093/asjof/ojad084
Comparing the Efficacy of Large Language Models ChatGPT, BARD, and Bing AI in Providing Information on Rhinoplasty: An Observational Study
resolves10.7759/cureus.43958
Artificial Intelligence (AI) in Radiology: A Deep Dive Into ChatGPT 4.0's Accuracy with the American Journal of Neuroradiology's (AJNR) "Case of the Month"
resolves10.2196/47479
Reliability of Medical Information Provided by ChatGPT: Assessment Against Clinical Guidelines and Patient Information Quality Instrument
resolves10.1016/j.ijmedinf.2023.105173
Performance and exploration of ChatGPT in medical examination, records and education in Chinese: Pave the way for medical AI
resolves10.1097/JCMA.0000000000000942
Performance of ChatGPT on the pharmacist licensing examination in Taiwan
resolves10.1007/s00590-023-03742-4
Evaluating ChatGPT responses in the context of a 53-year-old male with a femoral neck fracture: a qualitative analysis
resolves10.1007/s11042-023-15295-z
Recent advances in deep learning models: a systematic literature review
resolves10.2196/48808
ChatGPT-Generated Differential Diagnosis Lists for Complex Case–Derived Clinical Vignettes: Diagnostic Accuracy Evaluation
resolves10.2196/51232
Suicide Risk Assessments Through the Eyes of ChatGPT-3.5 Versus ChatGPT-4: Vignette Study
resolves10.2196/48039
Performance of ChatGPT on the Peruvian National Licensing Medical Examination: Cross-Sectional Study
resolves10.2196/49995
Comparison of Diagnostic and Triage Accuracy of Ada Health and WebMD Symptom Checkers, ChatGPT, and Physicians for Patients in an Emergency Department: Clinical Data Analysis Study
resolves10.2196/50514
Assessment of Resident and AI Chatbot Performance on the University of Toronto Family Medicine Residency Progress Test: Comparative Study
resolves10.1186/s40537-021-00444-8
Review of deep learning: concepts, CNN architectures, challenges, applications, future directions
resolves10.1016/j.inffus.2023.101805
Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence
resolves10.1007/s10439-023-03272-4
Prompt Engineering with ChatGPT: A Guide for Academic Writers
resolves10.2196/47049
The Potential and Concerns of Using AI in Scientific Research: ChatGPT Performance Evaluation
resolves10.1016/j.jjimei.2023.100165
How can we manage biases in artificial intelligence systems – A systematic literature review
resolves10.2196/48023
Accuracy of ChatGPT on Medical Questions in the National Medical Licensing Examination in Japan: Evaluation Study
resolves10.2196/47305
Performance of the Large Language Model ChatGPT on the National Nurse Examinations in Japan: Evaluation Study
resolves10.7759/cureus.49373
Pilot Testing of a Tool to Standardize the Assessment of the Quality of Health Information Generated by Artificial Intelligence-Based Models
resolves10.2196/49240
Clinical Accuracy of Large Language Models and Google Search Responses to Postpartum Depression Questions: Cross-Sectional Study
resolves10.2196/49324
Large Language Models for Therapy Recommendations Across 3 Clinical Specialties: Comparative Study
resolves10.2196/49280
Evaluation of ChatGPT Dermatology Responses to Common Patient Queries
resolves10.2196/46599
Trialling a Large Language Model (ChatGPT) in General Practice With the Applied Knowledge Test: Observational Study Demonstrating Opportunities and Limitations in Primary Care
resolves10.2196/50409
Assessing the Accuracy and Comprehensiveness of ChatGPT in Offering Clinical Guidance for Atopic Dermatitis and Acne Vulgaris
resolves10.2196/48978
Performance of ChatGPT on the Situational Judgement Test—A Professional Dilemmas–Based Examination for Doctors in the United Kingdom
resolves10.2196/51300
An AI Dietitian for Type 2 Diabetes Mellitus Management Based on Large Language and Image Recognition Models: Preclinical Concept Validation Study
resolves10.2196/49936
The Potential Influence of AI on Population Mental Health
resolves10.2196/54704
A Preliminary Checklist (METRICS) to Standardize the Design and Reporting of Studies on Generative Artificial Intelligence–Based Models in Health Care Education and Practice: Development Study Involving a Literature Review
resolves10.3389/feduc.2023.1333415
Below average ChatGPT performance in medical microbiology exam compared to university students
resolves10.7759/cureus.50629
ChatGPT Performance in Diagnostic Clinical Microbiology Laboratory-Oriented Case Scenarios
The 6 references without a DOI — listed, not checked
no DOI — not checkedChoudhuryAElkefiSTounsiAExploring factors influencing user perspective of ChatGPT as a technology that assists in healthcare decision making: A cross sectional survey studymedRxiv20232024-02-02https://www.medrxiv.org/content/10.1101/2023.12.07.23299685v1.full
no DOI — not checkedref24
no DOI — not checkedDoshiRAminKKhoslaPBajajSChheangSFormanHUtilizing Large Language Models to Simplify Radiology Reports: a comparative analysis of ChatGPT3.5, ChatGPT4.0, Google Bard, and Microsoft BingmedRxiv20232024-02-02https://www.medrxiv.org/content/10.1101/2023.06.04.23290786v2
no DOI — not checkedref39
no DOI — not checkedCASP Qualitative Studies ChecklistCritical Appraisal Skills Programme2023-11-10https://casp-uk.net/casp-tools-checklists/
no DOI — not checkedref51
What this badge says. CiteStamped means the CHECKABLE references of this work were clean at the dated check: each resolved to a known work in a public registry, and none carried a retraction notice at that time. It says nothing about the quality, findings, or importance of the work itself, and nothing about references deposited without a DOI.

checked 2026-08-16 — re-checked daily as this page is visited; titles and statuses come from Crossref and DataCite and are not part of the signed record

Embed this badge

Both snippets point at the live badge image and link back to this page. The badge re-renders from the daily check, so an embed never goes stale by more than a day of visits.

<a href="https://citestamp.com/citestamped/10.2196/54704"><img src="https://citestamp.com/citestamped/10.2196/54704/badge.svg" alt="CiteStamped reference-health badge" width="460" height="64"></a>
[![CiteStamped reference-health badge](https://citestamp.com/citestamped/10.2196/54704/badge.svg)](https://citestamp.com/citestamped/10.2196/54704)