0% Complete
Home
/
13th International Conference on Computer and Knowledge Engineering
SUT: a new multi-purpose synthetic dataset for Farsi document image analysis
Authors :
Elham Shabaninia
1
Fatemeh sadat Eslami
2
Ali Afkari Fahandari
3
Hossein Nezamabadi-pour
4
1- Department of Applied mathematics, Faculty of Sciences and Modern Technologies, Graduate University of Advanced Technology, Kerman, Iran
2- Department of Computer Engineering, Sirjan University of Technology, Sirjan, Iran
3- Department of Electrical Engineering, Shahid Bahonar University of Kerman, Kerman, Iran
4- Department of Electrical Engineering, Shahid Bahonar University of Kerman, Kerman, Iran
Keywords :
Document Image Analysis،Farsi database،Document Classification،Optical Character Recognition
Abstract :
This paper introduces a new large-scale dataset for Farsi document images, named SUT, which aims to tackle the challenges associated with obtaining diverse and substantial ground-truth data for supervised models in document image analysis (DIA) tasks, like document image classification, text detection and recognition, and information retrieval. The dataset comprises 62,453 images that have been categorized into 21 distinct classes, including identity documents featuring synthetically generated personal information superimposed on various backgrounds. The dataset also includes corresponding files with labeling information for the images. The ground-truth data is organized in CSV files containing image file paths and associated information about the embedded data. To demonstrate the efficacy of the SUT dataset in DIA tasks, it was utilized for document classification (achieving an accuracy of 86% using a convolutional neural network) and OCR (achieving a CER of 0.083 and 0.072 using Tesseract and EasyOCR engines, respectively). The SUT dataset serves as an esteemed asset for scholars engaged in the development and assessment of supervised models in Farsi document image analysis.
Papers List
List of archived papers
Chaotic multi-population ABC algorithm based on memory and levy flight for solving dynamic job shop scheduling problems
Mohammad Ali Zarif - Javad Hamidzadeh
An interactive user groups recommender system based on reinforcement learning
Hediyeh Naderi Allaf - Mohsen Kahani
Lossless Watermarking in Encrypted Triangular Mesh Models Based on Optimized Vertex Estimation and Error Histogram Shifting
Alireza Ghaemi - Habibollah Danyali - Kamran Kazemi - Zahra Qodrati - Amirhossein Ghaemi - Seyedeh Masoumeh Taji
Ramp Progressive Secret Image Sharing using Ensemble of Simple Methods
Atieh Mokhtari - Mohammad Taheri
Evaluating the Impact of Traveling on COVID-19 Prevalence and Predicting the New Confirmed Cases According to the Travel Rate Using Machine Learning: A Case Study in Iran
Anita Ghandehari - Soheil Shirvani - Hadi Moradi
Paddy Plant Stress Identification Using Few-Shot Learning Framework
Ervin Gubin Moung - Pavindrah Naidu a/l Narayanasamy Naiidu - Maisarah Mohd Sufian - Valentino Liaw - Ali Farzamnia - Lorita Angeline
Adversarial Robustness Evaluation with Separation Index
Bahareh Kaviani Baghbaderani - Afsaneh Hasanebrahimi - Ahmad Kalhor - Reshad Hosseini
A Federated Learning-Based Hybrid Deep Learning Framework for Enhanced Human Activity Recognition
Jamileh Azmoudeh - Sajjad Arghaee - Parisa Valizadeh - Samaneh Dandani - Iman Havangi - Mohammad Hossein Yaghmaee
Segmentation of Coronary Artery Stenosis in X-ray Angiography using Mamba Models
Fatemeh Fouladi - Ali Rostami - Hedieh Sajedi
Explainable Error Detection Method for Structured Data using HoloDetect framework
Abolfazl Mohajeri Khorasani - Sahar Ghassabi - Behshid Behkamal - Mostafa Milani
more
Samin Hamayesh - Version 41.3.1