Leveraging Large Language Models for Improved Patient Access and Self-Management: Assessor-Blinded Comparison Between Expert- and AI-Generated Content

Xiaolei Lv; Xiaomeng Zhang; Yuan Li; Xinxin Ding; Hongchang Lai; Junyu Shi

doi:10.2196/55847

Leveraging Large Language Models for Improved Patient Access and Self-Management: Assessor-Blinded Comparison Between Expert- and AI-Generated Content

J Med Internet Res. 2024 Apr 25:26:e55847. doi: 10.2196/55847.

Authors

Xiaolei Lv^{1

2

3

4

5

6}, Xiaomeng Zhang^{1

2

3

4

5

6}, Yuan Li^{1

2

3

4

5

6}, Xinxin Ding^{1

2

3

4

5

6}, Hongchang Lai^{1

2

3

4

5

6}, Junyu Shi^{1

2

3

4

5

6}

Affiliations

¹ Department of Oral and Maxillofacial Implantology, Shanghai PerioImplant Innovation Center, Shanghai Ninth People's Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.
² College of Stomatology, Shanghai Jiao Tong University, Shanghai, China.
³ National Center for Stomatology, Shanghai, China.
⁴ National Clinical Research Center for Oral Diseases, Shanghai, China.
⁵ Shanghai Key Laboratory of Stomatology, Shanghai, China.
⁶ Shanghai Research Institute of Stomatology, Shanghai, China.

PMID: 38663010
PMCID: PMC11082737
DOI: 10.2196/55847

Abstract

Background: While large language models (LLMs) such as ChatGPT and Google Bard have shown significant promise in various fields, their broader impact on enhancing patient health care access and quality, particularly in specialized domains such as oral health, requires comprehensive evaluation.

Objective: This study aims to assess the effectiveness of Google Bard, ChatGPT-3.5, and ChatGPT-4 in offering recommendations for common oral health issues, benchmarked against responses from human dental experts.

Methods: This comparative analysis used 40 questions derived from patient surveys on prevalent oral diseases, which were executed in a simulated clinical environment. Responses, obtained from both human experts and LLMs, were subject to a blinded evaluation process by experienced dentists and lay users, focusing on readability, appropriateness, harmlessness, comprehensiveness, intent capture, and helpfulness. Additionally, the stability of artificial intelligence responses was also assessed by submitting each question 3 times under consistent conditions.

Results: Google Bard excelled in readability but lagged in appropriateness when compared to human experts (mean 8.51, SD 0.37 vs mean 9.60, SD 0.33; P=.03). ChatGPT-3.5 and ChatGPT-4, however, performed comparably with human experts in terms of appropriateness (mean 8.96, SD 0.35 and mean 9.34, SD 0.47, respectively), with ChatGPT-4 demonstrating the highest stability and reliability. Furthermore, all 3 LLMs received superior harmlessness scores comparable to human experts, with lay users finding minimal differences in helpfulness and intent capture between the artificial intelligence models and human responses.

Conclusions: LLMs, particularly ChatGPT-4, show potential in oral health care, providing patient-centric information for enhancing patient education and clinical care. The observed performance variations underscore the need for ongoing refinement and ethical considerations in health care settings. Future research focuses on developing strategies for the safe integration of LLMs in health care settings.

Keywords: artificial intelligence; health care access; large language model; patient education; public oral health.

©Xiaolei Lv, Xiaomeng Zhang, Yuan Li, Xinxin Ding, Hongchang Lai, Junyu Shi. Originally published in the Journal of Medical Internet Research (https://www.jmir.org), 25.04.2024.

Publication types

Comparative Study

MeSH terms

Artificial Intelligence
Health Services Accessibility
Humans
Language
Oral Health
Self-Management* / methods