Generalization of finetuned transformer language models to new clinical contexts

Kevin Xie; Samuel W Terman; Ryan S Gallagher; Chloe E Hill; Kathryn A Davis; Brian Litt; Dan Roth; Colin A Ellis

doi:10.1093/jamiaopen/ooad070

Generalization of finetuned transformer language models to new clinical contexts

JAMIA Open. 2023 Aug 16;6(3):ooad070. doi: 10.1093/jamiaopen/ooad070. eCollection 2023 Oct.

Authors

Kevin Xie^{1

2}, Samuel W Terman³, Ryan S Gallagher^{2

4}, Chloe E Hill³, Kathryn A Davis^{2

4

5}, Brian Litt^{1

2

4

5}, Dan Roth⁶, Colin A Ellis^{2

4

5}

Affiliations

¹ Department of Bioengineering, University of Pennsylvania, Philadelphia, Pennsylvania 19104, USA.
² Center for Neuroengineering and Therapeutics, University of Pennsylvania, Philadelphia, Pennsylvania 19104, USA.
³ Department of Neurology, University of Michigan, Ann Arbor, Michigan 48109, USA.
⁴ Perelman School of Medicine, University of Pennsylvania, Philadelphia, Pennsylvania 19104, USA.
⁵ Department of Neurology, University of Pennsylvania, Philadelphia, Pennsylvania 19104, USA.
⁶ Department of Computer and Information Science, University of Pennsylvania, Philadelphia, Pennsylvania 19104, USA.

Abstract

Objective: We have previously developed a natural language processing pipeline using clinical notes written by epilepsy specialists to extract seizure freedom, seizure frequency text, and date of last seizure text for patients with epilepsy. It is important to understand how our methods generalize to new care contexts.

Materials and methods: We evaluated our pipeline on unseen notes from nonepilepsy-specialist neurologists and non-neurologists without any additional algorithm training. We tested the pipeline out-of-institution using epilepsy specialist notes from an outside medical center with only minor preprocessing adaptations. We examined reasons for discrepancies in performance in new contexts by measuring physical and semantic similarities between documents.

Results: Our ability to classify patient seizure freedom decreased by at least 0.12 agreement when moving from epilepsy specialists to nonspecialists or other institutions. On notes from our institution, textual overlap between the extracted outcomes and the gold standard annotations attained from manual chart review decreased by at least 0.11 F₁ when an answer existed but did not change when no answer existed; here our models generalized on notes from the outside institution, losing at most 0.02 agreement. We analyzed textual differences and found that syntactic and semantic differences in both clinically relevant sentences and surrounding contexts significantly influenced model performance.

Discussion and conclusion: Model generalization performance decreased on notes from nonspecialists; out-of-institution generalization on epilepsy specialist notes required small changes to preprocessing but was especially good for seizure frequency text and date of last seizure text, opening opportunities for multicenter collaborations using these outcomes.

Keywords: clinical informatics; electronic health records; epilepsy; natural language processing.

Abstract

Grants and funding