Implicit data crimes: Machine learning bias arising from misuse of public data

Efrat Shimron; Jonathan I Tamir; Ke Wang; Michael Lustig

doi:10.1073/pnas.2117203119

Implicit data crimes: Machine learning bias arising from misuse of public data

Proc Natl Acad Sci U S A. 2022 Mar 29;119(13):e2117203119. doi: 10.1073/pnas.2117203119. Epub 2022 Mar 21.

Authors

Efrat Shimron¹, Jonathan I Tamir^{2

3

4}, Ke Wang¹, Michael Lustig¹

Affiliations

¹ Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, CA 94720.
² Department of Electrical and Computer Engineering, The University of Texas at Austin, Austin, TX 78712.
³ Department of Diagnostic Medicine, Dell Medical School, The University of Texas at Austin, Austin, TX 78712.
⁴ Oden Institute for Computational Engineering and Sciences, The University of Texas at Austin, Austin, TX 78712.

Abstract

SignificancePublic databases are an important resource for machine learning research, but their growing availability sometimes leads to "off-label" usage, where data published for one task are used for another. This work reveals that such off-label usage could lead to biased, overly optimistic results of machine-learning algorithms. The underlying cause is that public data are processed with hidden processing pipelines that alter the data features. Here we study three well-known algorithms developed for image reconstruction from magnetic resonance imaging measurements and show they could produce biased results with up to 48% artificial improvement when applied to public databases. We relate to the publication of such results as implicit "data crimes" to raise community awareness of this growing big data problem.

Keywords: MRI; bias; big data; data crimes; inverse problem.

MeSH terms

Algorithms*
Bias
Crime
Image Processing, Computer-Assisted
Machine Learning*

Abstract

MeSH terms

Grants and funding