Accuracy of ChatGPT-3.5 and -4 in providing scientific references in otolaryngology-head and neck surgery

Jerome R Lechien; Giovanni Briganti; Luigi A Vaira

doi:10.1007/s00405-023-08441-8

Accuracy of ChatGPT-3.5 and -4 in providing scientific references in otolaryngology-head and neck surgery

Eur Arch Otorhinolaryngol. 2024 Apr;281(4):2159-2165. doi: 10.1007/s00405-023-08441-8. Epub 2024 Jan 11.

Authors

Jerome R Lechien^{1

2

3

4

5}, Giovanni Briganti⁶, Luigi A Vaira^{7

8}

Affiliations

¹ Division of Laryngology and Broncho-Esophagology, Department of Otolaryngology-Head Neck Surgery, EpiCURA Hospital, UMONS Research Institute for Health Sciences and Technology, University of Mons (UMons), Mons, Belgium. Jerome.Lechien@umons.ac.be.
² Department of Otorhinolaryngology and Head and Neck Surgery, School of Medicine, Phonetics and Phonology Laboratory (UMR 7018, Foch Hospital, CNRS, Université Sorbonne Nouvelle/Paris 3), Paris, France. Jerome.Lechien@umons.ac.be.
³ Department of Otorhinolaryngology and Head and Neck Surgery, School of Medicine, CHU de Bruxelles, CHU Saint-Pierre, Université Libre de Bruxelles, Brussels, Belgium. Jerome.Lechien@umons.ac.be.
⁴ Polyclinique Elsan de Poitiers, Poitiers, France. Jerome.Lechien@umons.ac.be.
⁵ Department of Human Anatomy and Experimental Oncology, Faculty of Medicine, UMONS Research Institute for Health Sciences and Technology, Avenue du Champ de Mars, 6, 7000, Mons, Belgium. Jerome.Lechien@umons.ac.be.
⁶ Chair of AI and Digital Medicine, Faculty of Medicine, University of Mons, Mons, Belgium.
⁷ Maxillofacial Surgery Operative Unit, Department of Medicine, Surgery and Pharmacy, University of Sassari, Sassari, Italy.
⁸ Biomedical Sciences Department, PhD School of Biomedical Science, University of Sassari, Sassari, Italy.

PMID: 38206389
DOI: 10.1007/s00405-023-08441-8

Abstract

Introduction: Chatbot generative pre-trained transformer (ChatGPT) is a new artificial intelligence-powered language model of chatbot able to help otolaryngologists in practice and research. We investigated the accuracy of ChatGPT-3.5 and -4 in the referencing of manuscripts published in otolaryngology.

Methods: ChatGPT-3.5 and ChatGPT-4 were interrogated for providing references of the top-30 most cited papers in otolaryngology in the past 40 years including clinical guidelines and key studies that changed the practice. The responses were regenerated three times to assess the accuracy and stability of ChatGPT. ChatGPT-3.5 and ChatGPT-4 were compared for accuracy of reference and potential mistakes.

Results: The accuracy of ChatGPT-3.5 and ChatGPT-4.0 ranged from 47% to 60%, and 73% to 87%, respectively (p < 0.005). ChatGPT-3.5 provided 19 inaccurate references and invented 2 references throughout the regenerated questions. ChatGPT-4.0 provided 13 inaccurate references, while it proposed only one invented reference. The stability of responses throughout regenerated answers was mild (k = 0.238) and moderate (k = 0.408) for ChatGPT-3.5 and 4.0, respectively.

Conclusions: ChatGPT-4.0 reported higher accuracy than the free-access version (3.5). False references were detected in both 3.5 and 4.0 versions. Practitioners need to be careful regarding the use of ChatGPT in the reach of some key reference when writing a report.

Keywords: Artificial intelligence; ChatGPT; Chatbot; Head neck surgery; Otolaryngology; Reference.

MeSH terms

Artificial Intelligence*
Humans
Language
Otolaryngologists
Otolaryngology*
Software