Using a Large Open Clinical Corpus for Improved ICD-10 Diagnosis Coding

Anastasios Lamproudis; Therese Olsen Svenning; Torbjørn Torsvik; Taridzo Chomutare; Andrius Budrionis; Phuong Dinh Ngo; Thomas Vakili; Hercules Dalianis

Using a Large Open Clinical Corpus for Improved ICD-10 Diagnosis Coding

AMIA Annu Symp Proc. 2024 Jan 11:2023:465-473. eCollection 2023.

Authors

Anastasios Lamproudis¹, Therese Olsen Svenning¹, Torbjørn Torsvik¹, Taridzo Chomutare^{1

2}, Andrius Budrionis^{1

3}, Phuong Dinh Ngo^{1

3}, Thomas Vakili⁴, Hercules Dalianis^{1

4}

Affiliations

¹ Norwegian Centre for E-health Research, Tromsø, Norway.
² Department of Computer Science, UiT - The Arctic University of Norway, Tromsø, Norway.
³ Department of Physics and Technology, UiT - The Arctic University of Norway, Tromsø, Norway.
⁴ Department of Computer and Systems Science (DSV), Stockholm University, Kista, Sweden.

PMID: 38222373
PMCID: PMC10785868

Abstract

With the recent advances in natural language processing and deep learning, the development of tools that can assist medical coders in ICD-10 diagnosis coding and increase their efficiency in coding discharge summaries is significantly more viable than before. To that end, one important component in the development of these models is the datasets used to train them. In this study, such datasets are presented, and it is shown that one of them can be used to develop a BERT-based language model that can consistently perform well in assigning ICD-10 codes to discharge summaries written in Swedish. Most importantly, it can be used in a coding support setup where a tool can recommend potential codes to the coders. This reduces the range of potential codes to consider and, in turn, reduces the workload of the coder. Moreover, the de-identified and pseudonymised dataset is open to use for academic users.

MeSH terms

Clinical Coding
Humans
International Classification of Diseases*
Natural Language Processing
Patient Discharge*