Fragmentstein-facilitating data reuse for cell-free DNA fragment analysis

Bioinformatics. 2024 Jan 2;40(1):btae017. doi: 10.1093/bioinformatics/btae017.

Abstract

Summary: Method development for the analysis of cell-free DNA (cfDNA) sequencing data is impeded by limited data sharing due to the strict control of sensitive genomic data. An existing solution for facilitating data sharing removes nucleotide-level information from raw cfDNA sequencing data, keeping alignment coordinates only. This simplified format can be publicly shared and would, theoretically, suffice for common functional analyses of cfDNA data. However, current bioinformatics software requires nucleotide-level information and cannot process the simplified format. We present Fragmentstein, a command-line tool for converting non-sensitive cfDNA-fragmentation data into alignment mapping (BAM) files. Fragmentstein complements fragment coordinates with sequence information from a reference genome to reconstruct BAM files. We demonstrate the utility of Fragmentstein by showing the feasibility of copy number variant (CNV), nucleosome occupancy, and fragment length analyses from non-sensitive fragmentation data.

Availability and implementation: Implemented in bash, Fragmentstein is available at https://github.com/uzh-dqbm-cmi/fragmentstein, licensed under GNU GPLv3.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Cell-Free Nucleic Acids*
  • Genome
  • Genomics
  • High-Throughput Nucleotide Sequencing / methods
  • Nucleotides
  • Sequence Analysis, DNA / methods
  • Software*

Substances

  • Cell-Free Nucleic Acids
  • Nucleotides

Grants and funding