Volumetric memory network for interactive medical image segmentation

Tianfei Zhou; Liulei Li; Gustav Bredell; Jianwu Li; Jan Unkelbach; Ender Konukoglu

doi:10.1016/j.media.2022.102599

Volumetric memory network for interactive medical image segmentation

Med Image Anal. 2023 Jan:83:102599. doi: 10.1016/j.media.2022.102599. Epub 2022 Sep 6.

Authors

Tianfei Zhou¹, Liulei Li², Gustav Bredell³, Jianwu Li², Jan Unkelbach⁴, Ender Konukoglu³

Affiliations

¹ Computer Vision Laboratory, ETH Zurich, Switzerland. Electronic address: tianfei.zhou@vision.ee.ethz.ch.
² School of Computer Science and Technology, Beijing Institute of Technology, China.
³ Computer Vision Laboratory, ETH Zurich, Switzerland.
⁴ Department of Radiation Oncology, University Hospital of Zurich, Zurich, Switzerland.

PMID: 36327652
DOI: 10.1016/j.media.2022.102599

Abstract

Despite recent progress of automatic medical image segmentation techniques, fully automatic results usually fail to meet clinically acceptable accuracy, thus typically require further refinement. To this end, we propose a novel Volumetric Memory Network, dubbed as VMN, to enable segmentation of 3D medical images in an interactive manner. Provided by user hints on an arbitrary slice, a 2D interaction network is firstly employed to produce an initial 2D segmentation for the chosen slice. Then, the VMN propagates the initial segmentation mask bidirectionally to all slices of the entire volume. Subsequent refinement based on additional user guidance on other slices can be incorporated in the same manner. To facilitate smooth human-in-the-loop segmentation, a quality assessment module is introduced to suggest the next slice for interaction based on the segmentation quality of each slice produced in the previous round. Our VMN demonstrates two distinctive features: First, the memory-augmented network design offers our model the ability to quickly encode past segmentation information, which will be retrieved later for the segmentation of other slices; Second, the quality assessment module enables the model to directly estimate the quality of each segmentation prediction, which allows for an active learning paradigm where users preferentially label the lowest-quality slice for multi-round refinement. The proposed network leads to a robust interactive segmentation engine, which can generalize well to various types of user annotations (e.g., scribble, bounding box, extreme clicking). Extensive experiments have been conducted on three public medical image segmentation datasets (i.e., MSD, KiTS₁₉, CVC-ClinicDB), and the results clearly confirm the superiority of our approach in comparison with state-of-the-art segmentation models. The code is made publicly available at https://github.com/0liliulei/Mem3D.

Keywords: Attention; Deep learning; Interactive image segmentation; Memory-augmented network; fully convolutional network.