Using systematic data categorisation to quantify the types of data collected in clinical trials: the DataCat project

Evelyn Crowley; Shaun Treweek; Katie Banister; Suzanne Breeman; Lynda Constable; Seonaidh Cotton; Anne Duncan; Adel El Feky; Heidi Gardner; Kirsteen Goodman; Doris Lanz; Alison McDonald; Emma Ogburn; Kath Starr; Natasha Stevens; Marie Valente; Gordon Fernie

doi:10.1186/s13063-020-04388-x

Using systematic data categorisation to quantify the types of data collected in clinical trials: the DataCat project

Trials. 2020 Jun 16;21(1):535. doi: 10.1186/s13063-020-04388-x.

Authors

Evelyn Crowley¹, Shaun Treweek², Katie Banister³, Suzanne Breeman⁴, Lynda Constable⁴, Seonaidh Cotton⁴, Anne Duncan⁴, Adel El Feky³, Heidi Gardner³, Kirsteen Goodman⁵, Doris Lanz⁶, Alison McDonald⁴, Emma Ogburn⁷, Kath Starr⁴, Natasha Stevens⁸, Marie Valente⁹, Gordon Fernie⁴

Affiliations

¹ Health Research Board Clinical Research Facility, University of Cork, Cork, Ireland.
² Health Services Research Unit, University of Aberdeen, Aberdeen, UK. streweek@mac.com.
³ Health Services Research Unit, University of Aberdeen, Aberdeen, UK.
⁴ Centre for Healthcare Randomised Trials, Health Services Research Unit, University of Aberdeen, Aberdeen, UK.
⁵ Nursing, Midwifery and Allied Health Professions (NMAHP) Research Unit, Glasgow Caledonian University, Glasgow, UK.
⁶ Institute of Population Health Sciences, Queen Mary University of London, London, UK.
⁷ Primary Care Clinical Trials Unit, University of Oxford, Oxford, UK.
⁸ Pragmatic Clinical Trials Unit, Queen Mary University of London, London, UK.
⁹ Birmingham Clinical Trials Unit, University of Birmingham, Birmingham, UK.

Abstract

Background: Data collection consumes a large proportion of clinical trial resources. Each data item requires time and effort for collection, processing and quality control procedures. In general, more data equals a heavier burden for trial staff and participants. It is also likely to increase costs. Knowing the types of data being collected, and in what proportion, will be helpful to ensure that limited trial resources and participant goodwill are used wisely.

Aim: The aim of this study is to categorise the types of data collected across a broad range of trials and assess what proportion of collected data each category represents.

Methods: We developed a standard operating procedure to categorise data into primary outcome, secondary outcome and 15 other categories. We categorised all variables collected on trial data collection forms from 18, mainly publicly funded, randomised superiority trials, including trials of an investigational medicinal product and complex interventions. Categorisation was done independently in pairs: one person having in-depth knowledge of the trial, the other independent of the trial. Disagreement was resolved through reference to the trial protocol and discussion, with the project team being consulted if necessary.

Key results: Primary outcome data accounted for 5.0% (median)/11.2% (mean) of all data items collected. Secondary outcomes accounted for 39.9% (median)/42.5% (mean) of all data items. Non-outcome data such as participant identifiers and demographic data represented 32.4% (median)/36.5% (mean) of all data items collected.

Conclusion: A small proportion of the data collected in our sample of 18 trials was related to the primary outcome. Secondary outcomes accounted for eight times the volume of data as the primary outcome. A substantial amount of data collection is not related to trial outcomes. Trialists should work to make sure that the data they collect are only those essential to support the health and treatment decisions of those whom the trial is designed to inform.

MeSH terms

Clinical Trials as Topic / statistics & numerical data*
Data Collection / classification*
Data Collection / standards*
Data Interpretation, Statistical
Humans

Grants and funding

HSRU1/CSO_/Chief Scientist Office/United Kingdom