Post-randomization for controlling identification risk in releasing microdata from general surveys

Cheng Zhang; Tapan K Nayak

doi:10.1080/02664763.2020.1732310

Post-randomization for controlling identification risk in releasing microdata from general surveys

J Appl Stat. 2020 Feb 26;48(3):455-470. doi: 10.1080/02664763.2020.1732310. eCollection 2021.

Authors

Cheng Zhang¹, Tapan K Nayak^{2

3}

Affiliations

¹ MedStar Cardiovascular Research Network, Washington, DC, USA.
² Center for Statistical Research and Methodology, U.S. Census Bureau, Washington, DC, USA.
³ Department of Statistics, George Washington University, Washington, DC, USA.

Abstract

Before releasing survey data, statistical agencies usually perturb the original data to keep each survey unit's information confidential. One significant concern in releasing survey microdata is identity disclosure, which occurs when an intruder correctly identifies the records of a survey unit by matching the values of some key (or pseudo-identifying) variables. We examine a recently developed post-randomization method for a strict control of identification risks in releasing survey microdata. While that procedure well preserves the observed frequencies and hence statistical estimates in case of simple random sampling, we show that in general surveys, it may induce considerable bias in commonly used survey-weighted estimators. We propose a modified procedure that better preserves weighted estimates. The procedure is illustrated and empirically assessed with an application to a publicly available US Census Bureau data set.

Keywords: 62D05; Identity disclosure; data partitioning; key variable; post-randomization block; survey weighted estimator; total variation distance.

This work was authored as part of the Contributor's official duties as an Employee of the United States Government and is therefore a work of the United States Government. In accordance with 17 U.S.C. 105, no copyright protection is available for such works under U.S. Law.