Quantification of BERT Diagnosis Generalizability Across Medical Specialties Using Semantic Dataset Distance

Mihir P Khambete; William Su; Juan C Garcia; Marcus A Badgeley

Quantification of BERT Diagnosis Generalizability Across Medical Specialties Using Semantic Dataset Distance

AMIA Jt Summits Transl Sci Proc. 2021 May 17:2021:345-354. eCollection 2021.

Authors

Mihir P Khambete^{1

2}, William Su^{1

3}, Juan C Garcia¹, Marcus A Badgeley^{1

4}

Affiliations

¹ nference LLC, Cambridge, MA.
² Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA.
³ Department of Radiation Oncology, Penn Medicine, University of Pennsylvania Health System, Philadelphia, PA.
⁴ Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology, Cambridge, MA.

PMID: 34457149
PMCID: PMC8378651

Abstract

Deep learning models in healthcare may fail to generalize on data from unseen corpora. Additionally, no quantitative metric exists to tell how existing models will perform on new data. Previous studies demonstrated that NLP models of medical notes generalize variably between institutions, but ignored other levels of healthcare organization. We measured SciBERT diagnosis sentiment classifier generalizability between medical specialties using EHR sentences from MIMIC-III. Models trained on one specialty performed better on internal test sets than mixed or external test sets (mean AUCs 0.92, 0.87, and 0.83, respectively; p = 0.016). When models are trained on more specialties, they have better test performances (p < 1e-4). Model performance on new corpora is directly correlated to the similarity between train and test sentence content (p < 1e-4). Future studies should assess additional axes of generalization to ensure deep learning models fulfil their intended purpose across institutions, specialties, and practices.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Deep Learning*
Humans
Language
Medicine*
Semantics