Fish Segmentation in Sonar Images by Mask R-CNN on Feature Maps of Conditional Random Fields

Sensors (Basel). 2021 Nov 17;21(22):7625. doi: 10.3390/s21227625.

Abstract

Imaging sonar systems are widely used for monitoring fish behavior in turbid or low ambient light waters. For analyzing fish behavior in sonar images, fish segmentation is often required. In this paper, Mask R-CNN is adopted for segmenting fish in sonar images. Sonar images acquired from different shallow waters can be quite different in the contrast between fish and the background. That difference can make Mask R-CNN trained on examples collected from one fish farm ineffective to fish segmentation for the other fish farms. In this paper, a preprocessing convolutional neural network (PreCNN) is proposed to provide "standardized" feature maps for Mask R-CNN and to ease applying Mask R-CNN trained for one fish farm to the others. PreCNN aims at decoupling learning of fish instances from learning of fish-cultured environments. PreCNN is a semantic segmentation network and integrated with conditional random fields. PreCNN can utilize successive sonar images and can be trained by semi-supervised learning to make use of unlabeled information. Experimental results have shown that Mask R-CNN on the output of PreCNN is more accurate than Mask R-CNN directly on sonar images. Applying Mask R-CNN plus PreCNN trained for one fish farm to new fish farms is also more effective.

Keywords: conditional random fields; fish segmentation; mask R-CNN; sonar images.

MeSH terms

  • Image Processing, Computer-Assisted*
  • Neural Networks, Computer*
  • Sound
  • Specimen Handling
  • Supervised Machine Learning