Advancing Italian biomedical information extraction with transformers-based models: Methodological insights and multicenter practical application

Claudio Crema; Tommaso Mario Buonocore; Silvia Fostinelli; Enea Parimbelli; Federico Verde; Cira Fundarò; Marina Manera; Matteo Cotta Ramusino; Marco Capelli; Alfredo Costa; Giuliano Binetti; Riccardo Bellazzi; Alberto Redolfi

doi:10.1016/j.jbi.2023.104557

Advancing Italian biomedical information extraction with transformers-based models: Methodological insights and multicenter practical application

J Biomed Inform. 2023 Dec:148:104557. doi: 10.1016/j.jbi.2023.104557. Epub 2023 Nov 25.

Authors

Affiliations

¹ Laboratory of Neuroinformatics, IRCCS Istituto Centro San Giovanni di Dio Fatebenefratelli, Brescia, Italy. Electronic address: ccrema@fatebenefratelli.eu.
² Dept. of Electrical, Computer and Biomedical Engineering, University of Pavia, Pavia, Italy. Electronic address: buonocore.tms@gmail.com.
³ Molecular Markers Laboratory, IRCCS Istituto Centro San Giovanni di Dio Fatebenefratelli, Brescia, Italy. Electronic address: sfostinelli@fatebenefratelli.eu.
⁴ Dept. of Electrical, Computer and Biomedical Engineering, University of Pavia, Pavia, Italy. Electronic address: enea.parimbelli@unipv.it.
⁵ Department of Neurology and Laboratory of Neuroscience, IRCCS Istituto Auxologico Italiano, Milan, Italy; Department of Pathophysiology and Transplantation, Dino Ferrari Center, Università degli Studi di Milano, Milan, Italy. Electronic address: f.verde@auxologico.it.
⁶ Neurophysiopatology Unit, IRCCS Istituti Clinici Scientifici Maugeri, Pavia, Italy. Electronic address: cira.fundaro@icsmaugeri.it.
⁷ Psychology Unit, IRCCS Istituti Clinici Scientifici Maugeri, Pavia, Italy. Electronic address: marina.manera@icsmaugeri.it.
⁸ Unit of Behavioral Neurology, IRCCS Mondino Foundation Pavia, and Dept. of Brain and Behavioral Sciences, University of Pavia, Pavia, Italy. Electronic address: matteo.cottaramusino@mondino.it.
⁹ Unit of Behavioral Neurology, IRCCS Mondino Foundation Pavia, and Dept. of Brain and Behavioral Sciences, University of Pavia, Pavia, Italy. Electronic address: marco.capelli@mondino.it.
¹⁰ Unit of Behavioral Neurology, IRCCS Mondino Foundation Pavia, and Dept. of Brain and Behavioral Sciences, University of Pavia, Pavia, Italy. Electronic address: alfredo.costa@mondino.it.
¹¹ Molecular Markers Laboratory, IRCCS Istituto Centro San Giovanni di Dio Fatebenefratelli, Brescia, Italy. Electronic address: gbinetti@fatebenefratelli.eu.
¹² Dept. of Electrical, Computer and Biomedical Engineering, University of Pavia, Pavia, Italy. Electronic address: riccardo.bellazzi@unipv.it.
¹³ Laboratory of Neuroinformatics, IRCCS Istituto Centro San Giovanni di Dio Fatebenefratelli, Brescia, Italy. Electronic address: aredolfi@fatebenefratelli.eu.

PMID: 38012982
DOI: 10.1016/j.jbi.2023.104557

Abstract

The introduction of computerized medical records in hospitals has reduced burdensome activities like manual writing and information fetching. However, the data contained in medical records are still far underutilized, primarily because extracting data from unstructured textual medical records takes time and effort. Information Extraction, a subfield of Natural Language Processing, can help clinical practitioners overcome this limitation by using automated text-mining pipelines. In this work, we created the first Italian neuropsychiatric Named Entity Recognition dataset, PsyNIT, and used it to develop a Transformers-based model. Moreover, we collected and leveraged three external independent datasets to implement an effective multicenter model, with overall F1-score 84.77 %, Precision 83.16 %, Recall 86.44 %. The lessons learned are: (i) the crucial role of a consistent annotation process and (ii) a fine-tuning strategy that combines classical methods with a "low-resource" approach. This allowed us to establish methodological guidelines that pave the way for Natural Language Processing studies in less-resourced languages.

Keywords: Biomedical text mining; Deep learning; Language model; Natural language processing; Transformer.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Data Mining* / methods
Electronic Health Records
Humans
Italy
Language*
Multicenter Studies as Topic
Natural Language Processing