Multilingual character recognition dataset for Moroccan official documents

Ali Benaissa; Abdelkhalak Bahri; Ahmad El Allaoui

doi:10.1016/j.dib.2023.109953

Multilingual character recognition dataset for Moroccan official documents

Data Brief. 2023 Dec 13:52:109953. doi: 10.1016/j.dib.2023.109953. eCollection 2024 Feb.

Authors

Ali Benaissa^{1

2}, Abdelkhalak Bahri¹, Ahmad El Allaoui³

Affiliations

¹ Data Science and Competitive Intelligence Team (DSCI), ENSAH, Abdelmalek Essaadi University (UAE), Tetouan, Morocco.
² Finance and Governance of Organizations team, Governance and Performance of Organizations laboratory, The National School of Management, Abdelmalek Essaadi University, Tangier, Morocco.
³ Decisional Computing and Systems Modelling Team, Engineering Sciences and Techniques Laboratory, Faculty of Sciences and Techniques Errachidia, Moulay Ismail University of Meknes, Morocco.

Abstract

This article focuses on the construction of a dataset for multilingual character recognition in Moroccan official documents. The dataset covers languages such as Arabic, French, and Tamazight and are built programmatically to ensure data diversity. It consists of sub-datasets such as Uppercase alphabet (26 classes), Lowercase alphabet (26 classes), Digits (9 classes), Arabic (28 classes), Tifinagh letters (33 classes), Symbols (14 classes), and French special characters (16 classes). The dataset construction process involves collecting representative fonts and generating multiple character images using a Python script, presenting a comprehensive variety essential for robust recognition models. Moreover, this dataset contributes to the digitization of these diverse official documents and archival papers, essential for preserving cultural heritage and enabling advanced text recognition technologies. The need for this work arises from the advancements in character recognition techniques and the significance of large-scale annotated datasets. The proposed dataset contributes to the development of robust character recognition models for practical applications.

Keywords: Character recognition; Documents digitization; Moroccan characters images; Moroccan documents; OCR dataset; Printed characters.