Accessibility settings

Published on in Vol 27 (2025)

Preprints (earlier versions) of this paper are available at https://preprints.jmir.org/preprint/84120, first published .
Doctor reviewing patient chart and writing notes with stethoscope

Knowledge-Practice Performance Gap in Clinical Large Language Models: Systematic Review of 39 Benchmarks

Knowledge-Practice Performance Gap in Clinical Large Language Models: Systematic Review of 39 Benchmarks

Journals

  1. Lin X, Yang Y, Ren Y. Making Chatbots more human: deep reasoning large language models in ophthalmology. Frontiers in Medicine 2026;12 View
  2. Spieser J, Balapour A, Meller J, Patra K, Shamsaei B. A Review of Multi-Agent AI Systems for Biological and Clinical Data Analysis. Methods and Protocols 2026;9(2):33 View
  3. Zhu Q, Li Q, Zan Y, Lu Y, Xia L, Xia Y, Xu T. Patient-centered gastrointestinal function assessment technologies: a paradigm shift from traditional approaches to non-invasive innovations. Frontiers in Physiology 2026;17 View
  4. Prause M. No skin in the game: why agentic AI requires principal-agent governance. AI and Ethics 2026;6(2) View
  5. Eltaybani S. Knowledge Cut‐Off in Large Language Models: Implications for Critical Care Nursing. Nursing in Critical Care 2026;31(3) View
  6. Lee W, Kim J, Leem J, Lee B, Lee S, Kim Y. Benchmark Evaluation of a Tool-Augmented Large Language Model Agent Using Traditional Asian Medicine Metadata. Applied Sciences 2026;16(7):3377 View
  7. Mine Y, Taji T, Okazaki S, Takeda S, Shimoe S, Kaku M, Nikawa H, Kakimoto N, Murayama T. Beyond exam accuracy: Tracking a persistent-failure set reveals visual dental reasoning gaps in multimodal LLMs. Journal of Dentistry 2026;170:106675 View
  8. Wang X, Yin C, He H, Guo J, Fu X, Bai F. Benchmarking public large language model responses to patient-facing inflammatory bowel disease questions: informational quality, transparency proxies, and readability. Frontiers in Public Health 2026;14 View
  9. Keshav T, Chow D, Kippenberger T, Livezey J, Aranda M. Evaluating Large-Language Models Against Providers on Surgical Diagnostic Reasoning Tasks. Journal of Surgical Research 2026;322:259 View
  10. Rajwal S, Pandey A, Zhang Z, Chen Y, Liu M, Das S, Rogers H, Sarker A, Xiao Y. Applications of Natural Language Processing and Large Language Models for Social Determinants of Health: Systematic Review. Journal of Medical Internet Research 2026;28:e83793 View
  11. Bajwa M, Hoyt R, Knight D, Haider M. The Performance of DeepSeek R1 and Gemini 3 in Complex Medical Scenarios: Comparative Study. JMIRx Med 2026;7:e76822 View
  12. Chang Y, Hsieh M, Ju P, Liu Y, Chang C. Clinical Plausibility in Large Language Model Robustness Testing for Medicine: A Scoping Review. Journal of Medical Systems 2026;50(1) View
  13. Karunanayake N. When Chatbots Become Agents: The Next Phase of Healthcare AI. Journal of Medical Systems 2026;50(1) View
  14. Yeh Y, Shih M, De Backer D, Celi L, See K, Fujii T, Ling L, Mongkolpun W, Hu H, Chen H, Chen W, Cholley B, Fong K, Ryu H, Na S, Egi M, Chan W, Chen K, Kamaleswaran R, Chuang Y, Yang C, Hsiao W, Lai S, Ku D, Jahan A, Martin G. The IMPACT framework for evaluating generative AI in critical care: development and multinational consensus validation. Annals of Intensive Care 2026;16:100078 View
  15. Khosravi M, Zamaninasab Z, Khosravi F, Attar M, Arab‐Zozani M. Performance of Large Language Models in Answering Healthcare Delivery Questions: A Quantitative Cross‐Sectional Study. Health Science Reports 2026;9(6) View
  16. Khosravi M, Dindar E, Sayar B. Evaluating large language model`s performance in answering principles of health course questions. Scientific Reports 2026;16(1) View
  17. Karataş S, Öner S. Pre-deployment safety and governance assessment of LLM-based clinical decision support systems: A health technology assessment-oriented evaluation framework. Health Policy and Technology 2026;15(9):101281 View
  18. Zhang W, Xu J, Dong T, Hao X, Yang Q, Zhang J, Han Y. Exploring large language models as a prescription decision support tool for rational antibiotic use: A dual-framework analysis using standardized examinations and real-world clinical cases. Exploratory Research in Clinical and Social Pharmacy 2026;23:100821 View
  19. Ucdal M, Ekingen E, Kurtcebe A. Benchmark Performance of a Neurosymbolic Multi Model Large Language Reasoning Pipeline Versus Board Certified Specialists and Single Model Baselines on Septic Arthritis: A Five Center Prospective Benchmarking Study with Item Level and Question Subtype Analysis (Preprint). JMIR AI 2026 View
  20. Zhang Z, Chen L, Lv Z, Lv H, Sheng W, Wei Z, Wang B, Shen Y, Tian Y, Hu J, Shen Z, Lv L. Discordance Between Textual Reasoning and Visual Interpretation in Large Language Models for Low Back Pain: Cross-Sectional Quantitative Evaluation and Exploratory Multimodal Stress Test. JMIR Medical Informatics 2026;14:e93522 View
  21. Dong Y, Cheng J, Ding C, Lu R. Governing Clinical Readiness Claims Derived from Medical AI Benchmark Results. Journal of Medical Systems 2026;50(1) View
  22. Guo Z, Lai A, Korakas E, Vagenas A, Ahamed I, Albor C, Zhang H, Healy J, Li K. Retrieval-Augmented Large Language Model Counseling for Continuous Glucose Monitoring in Diabetes: Source-Masked Multirater Comparative Evaluation. Journal of Medical Internet Research 2026;28:e98519 View
  23. Eapen K, Krothapalli V, Vaid R, Dhinagar M, Raju L, Goyal R, Irodi A. Patterns of Errors and Hallucinations Among ChatGPT, Perplexity, Qwen, and Copilot in Answering ACR DXIT Radiology Questions. Indian Journal of Radiology and Imaging 2026 View
  24. Diyasena H, Ahmad A, Weerasinghe T, Crawshaw O, O'Brien A. Comparative Error Analysis of GPT-5.2 Reasoning Modes on Performance in Postgraduate Internal Medicine Single-Best-Answer Questions. Cureus 2026 View
  25. Damor P, Gill S, Wani W, Arora G, Charan M. A Psychometric Comparison of Faculty-Authored and Large Language Model-Generated Multiple-Choice Questions in Endodontics. Cureus 2026 View
  26. Güler I, Kraus A, Grieb G, Stelling H. Burn Extent and Fitzpatrick Skin Tone Assessment from Clinical Photographs: Systematic and Random Error in Multimodal Large Language Models. Bioengineering 2026;13(9):1000 View
  27. Morris K, Öttl F, Pruneski J, Zsidai B, Hirschmann M, Samuelsson K, Kraeutler M. Select large language models outperform hip preservation experts on consensus‐based hip preservation questionnaire. Knee Surgery, Sports Traumatology, Arthroscopy 2026 View
  28. Nonye Tochi Aghanya , Osinakachi Akuma Kalu . Digital Distrust: Patient-Centred Communication and the Calibration of Epistemic Reliance in the Era of Online Self-Diagnosis and AI Symptom Checkers. Journal of Medical Clinical Case Reports 2026 View
  29. Liu L, Dong C, Ba Y, Yan H, Tang X, Sun Y, Chen H, Shi B, Yu Q, Zhang S. Dynamic alignment of large language models for evidence-grounded heart failure decision support. DIGITAL HEALTH 2026;12 View

Books/Policy Documents

  1. Jha A, Bhatele K, Mihir P. Medical LLMs for Clinical Safety Assessment. View

Conference Proceedings

  1. Kumar A, Joshi S, Sachdeva S. 2026 International Conference on Signal Processing and Electronics Design (ICSPED). JsonUtil: An Open-Source RESTful JSON-Based Dynamic Form Generation Framework validation with OpenEHR ORBDA Benchmarking Dataset View