A machine learning approach to detect potentially harmful and protective suicide-related content in broadcast media

Hannah Metzler; Hubert Baginski; David Garcia; Thomas Niederkrotenthaler

doi:10.1371/journal.pone.0300917

A machine learning approach to detect potentially harmful and protective suicide-related content in broadcast media

PLoS One. 2024 May 14;19(5):e0300917. doi: 10.1371/journal.pone.0300917. eCollection 2024.

Authors

Hannah Metzler^{1

2

3

4}, Hubert Baginski^{3

5}, David Garcia^{1

3

6

7}, Thomas Niederkrotenthaler^{2

8}

Affiliations

¹ Section for Science of Complex Systems, Center for Medical Data Science, Medical University of Vienna, Vienna, Austria.
² Unit Public Mental Health Research, Department of Social and Preventive Medicine, Center for Public Health, Medical University of Vienna, Vienna, Austria.
³ Complexity Science Hub, Vienna, Austria.
⁴ Institute for Globally Distributed Open Research and Education, Austria.
⁵ Institute of Information Systems Engineering, Vienna University of Technology, Vienna, Austria.
⁶ Department of Politics and Public Administration, University of Konstanz, Konstanz, Germany.
⁷ Institute of Interactive Systems and Data Science, Department of Computer Science and Biomedical Engineering, Graz University of Technology, Graz, Austria.
⁸ Wiener Werkstaette for Suicide Research, Vienna, Austria.

Abstract

Suicide-related media content has preventive or harmful effects depending on the specific content. Proactive media screening for suicide prevention is hampered by the scarcity of machine learning approaches to detect specific characteristics in news reports. This study applied machine learning to label large quantities of broadcast (TV and radio) media data according to media recommendations reporting suicide. We manually labeled 2519 English transcripts from 44 broadcast sources in Oregon and Washington, USA, published between April 2019 and March 2020. We conducted a content analysis of media reports regarding content characteristics. We trained a benchmark of machine learning models including a majority classifier, approaches based on word frequency (TF-IDF with a linear SVM) and a deep learning model (BERT). We applied these models to a selection of more simple (e.g., focus on a suicide death), and subsequently to putatively more complex tasks (e.g., determining the main focus of a text from 14 categories). Tf-idf with SVM and BERT were clearly better than the naive majority classifier for all characteristics. In a test dataset not used during model training, F1-scores (i.e., the harmonic mean of precision and recall) ranged from 0.90 for celebrity suicide down to 0.58 for the identification of the main focus of the media item. Model performance depended strongly on the number of training samples available, and much less on assumed difficulty of the classification task. This study demonstrates that machine learning models can achieve very satisfactory results for classifying suicide-related broadcast media content, including multi-class characteristics, as long as enough training samples are available. The developed models enable future large-scale screening and investigations of broadcast media.

Copyright: © 2024 Metzler et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

MeSH terms

Deep Learning
Humans
Machine Learning*
Mass Media*
Oregon
Suicide
Suicide Prevention
Washington

Grants and funding

This work was funded by a granta from Vibrant Emotional Health (award number: 2020-n.a.) (https://www.vibrant.org/) to TN. This research has also been funded by the Vienna Science and Technology Fund (WWTF) [10.47379/VRG16005] (https://www.wwtf.at) to DG. The funders did not have any role in the study design; data collection and analysis; or preparation of the manuscript. Vibrant Emotional Health approved the decision to submit for publication.