A semi-automated workflow for biodiversity data retrieval, cleaning, and quality control

Biodivers Data J. 2014 Dec 11:(2):e4221. doi: 10.3897/BDJ.2.e4221. eCollection 2014.

Abstract

The compilation and cleaning of data needed for analyses and prediction of species distributions is a time consuming process requiring a solid understanding of data formats and service APIs provided by biodiversity informatics infrastructures. We designed and implemented a Taverna-based Data Refinement Workflow which integrates taxonomic data retrieval, data cleaning, and data selection into a consistent, standards-based, and effective system hiding the complexity of underlying service infrastructures. The workflow can be freely used both locally and through a web-portal which does not require additional software installations by users.

Keywords: biodiversity informatics; data cleaning; e-Science; service oriented architecture; web services; workflows.