Text Recognition Model Based on Multi-Scale Fusion CRNN

Le Zou; Zhihuang He; Kai Wang; Zhize Wu; Yifan Wang; Guanhong Zhang; Xiaofeng Wang

doi:10.3390/s23167034

Text Recognition Model Based on Multi-Scale Fusion CRNN

Sensors (Basel). 2023 Aug 8;23(16):7034. doi: 10.3390/s23167034.

Authors

Le Zou¹, Zhihuang He¹, Kai Wang¹, Zhize Wu¹, Yifan Wang¹, Guanhong Zhang¹, Xiaofeng Wang¹

Affiliation

¹ School of Artificial Intelligence and Big Data, Hefei University, Hefei 230601, China.

Abstract

Scene text recognition is a crucial area of research in computer vision. However, current mainstream scene text recognition models suffer from incomplete feature extraction due to the small downsampling scale used to extract features and obtain more features. This limitation hampers their ability to extract complete features of each character in the image, resulting in lower accuracy in the text recognition process. To address this issue, a novel text recognition model based on multi-scale fusion and the convolutional recurrent neural network (CRNN) has been proposed in this paper. The proposed model has a convolutional layer, a feature fusion layer, a recurrent layer, and a transcription layer. The convolutional layer uses two scales of feature extraction, which enables it to derive two distinct outputs for the input text image. The feature fusion layer fuses the different scales of features and forms a new feature. The recurrent layer learns contextual features from the input sequence of features. The transcription layer outputs the final result. The proposed model not only expands the recognition field but also learns more image features at different scales; thus, it extracts a more complete set of features and achieving better recognition of text. The results of experiments are then presented to demonstrate that the proposed model outperforms the CRNN model on text datasets, such as Street View Text, IIIT-5K, ICDAR2003, and ICDAR2013 scenes, in terms of text recognition accuracy.

Keywords: feature fusion; multi-scale; text recognition.

Abstract

Grants and funding