COCO: an annotated Twitter dataset of COVID-19 conspiracy theories

Johannes Langguth; Daniel Thilo Schroeder; Petra Filkuková; Stefan Brenner; Jesper Phillips; Konstantin Pogorelov

doi:10.1007/s42001-023-00200-3

COCO: an annotated Twitter dataset of COVID-19 conspiracy theories

J Comput Soc Sci. 2023 Apr 4:1-42. doi: 10.1007/s42001-023-00200-3. Online ahead of print.

Authors

Johannes Langguth^{1

2}, Daniel Thilo Schroeder^{1

3}, Petra Filkuková¹, Stefan Brenner⁴, Jesper Phillips⁵, Konstantin Pogorelov¹

Affiliations

¹ Simula Research Lab, Kristian Augusts Gate 23, Oslo, Norway.
² Norwegian Business School, Nydalsveien 37, Oslo, Norway.
³ Department of Journalism and Media Studies, Oslo Metropolitan University, Pilestredet Park 0890, 0176 Oslo, Norway.
⁴ Stuttgart Media University, Nobelstraße 10, Stuttgart, Germany.
⁵ Bates College, Andrews Rd 2, Lewiston, ME USA.

Abstract

The COVID-19 pandemic has been accompanied by a surge of misinformation on social media which covered a wide range of different topics and contained many competing narratives, including conspiracy theories. To study such conspiracy theories, we created a dataset of 3495 tweets with manual labeling of the stance of each tweet w.r.t. 12 different conspiracy topics. The dataset thus contains almost 42,000 labels, each of which determined by majority among three expert annotators. The dataset was selected from COVID-19 related Twitter data spanning from January 2020 to June 2021 using a list of 54 keywords. The dataset can be used to train machine learning based classifiers for both stance and topic detection, either individually or simultaneously. BERT was used successfully for the combined task. The dataset can also be used to further study the prevalence of different conspiracy narratives. To this end we qualitatively analyze the tweets, discussing the structure of conspiracy narratives that are frequently found in the dataset. Furthermore, we illustrate the interconnection between the conspiracy categories as well as the keywords.

Keywords: BERT; Conspiracy theories; Misinformation; Twitter.