0% Complete
Home
/
14th International Conference on Computer and Knowledge Engineering
A Comprehensive Dataset of Real-scene Images for Text Detection and Recognition in Persian
Authors :
Iman Souzanchi
1
Ramin Rahimi
2
Mohammad Ali Majidi Anvari
3
Atefeh Baniasadi
4
Ashkan Sadeghi
5
Mohammad Reza Mohammadi
6
1- PART AI Research Center
2- PART AI Research Center
3- PART AI Research Center
4- PART AI Research Center
5- PART AI Research Center
6- School of Computer Engineering, Iran University of Science and Technology
Keywords :
Persian scene text dataset،Scene text recognition،Deep learning
Abstract :
Extracting text from scene images is a widely utilized field owing to the abundance of information available in scene images and their potential utilization in computer vision applications such as self-driving cars, text translation, information extraction from invoices, shopfronts, license plate retrieval, etc. Nonetheless, this field presents challenges because of the varying fonts, styles, sizes, and other characteristics of the text. Despite the existence of numerous studies on scene text recognition for languages such as English that employ deep learning models, a major barrier to implementing these models in Persian is the lack of an appropriate and sufficient dataset both in terms of quantity and quality. This paper aims to introduce a comprehensive collection of Persian scene images obtained from diverse sources, including newspapers, magazines, books, business cards, road signs, advertising billboards, shopfronts, invoices, and scanned documents. This dataset comprises over 250k of annotated text lines from 5000 images, including various lengths, fonts, and sizes that have been prepared under different conditions, including varying brightness and viewing angles. Additionally, more than 2,500,000 images of meaningful sentences have been synthesized since the annotation of real data is so expensive. In order to assess the efficacy of our dataset, a scene text recognition model was trained from existing models, and a word-accuracy of 83.9% was achieved on challenging test images.
Papers List
List of archived papers
Hybrid Vision Transformer for Detection of Dentigerous Cysts in Dental Radiography Images
Reza Tavasoli - Arya VarastehNezhad - Hamed Farbeh
A Robust Network for Embedded Traffic Sign Recognation.
Omid Nejati Manzari - Shahriar Baradaran Shokouhi
DPRNN-FORMER: AN EFFICIENT WAY TO DEAL WITH BLIND SOURCE SEPARATION
Ramin Ghorbani - Sajad Haghzad Klidbary
Probabilistic Short-Term Load Forecasting Using GBDT-Based Sister Forecasts and Ensemble Methods
Hossein Shahinzadeh - Hamed Nafisi - Amirafshin Zamani - Saiedeh Mehrabani-Najafabadi - Arezou Mahmoudi - Farshad Ebrahimi
A Deep CNN Model Based Ensemble Approach for Semantic and Instance Segmentation of Indoor Environment
Sajad Rezaei - Jafar Tanha - Zahra Jafari - SeyedEhsan Roshan - Mohammad-Amin Memar Kochebagh
Uncertainty-Aware Deep Ensembles for Confident Customer Churn Prediction with Rejection Option
Fatemeh Moradi - Mehran Tarif - Mohammadhossein Homaei
A Systematic Embedded Software Design Flow for Robotic Applications
Navid Mahdian - Seyed-Hosein Attarzadeh-Niaki - Armin Salimi-Badr
Efficient T-Count Fault-tolerant Quantum Clifford+T Multiplexer
Negin Mashayekhi - Shekoofeh Moghimi - Mohammad Reza Reshadinezhad
A New Time Series Approach in Churn Prediction with Discriminatory Intervals
Hedieh Ahmadi - Seyed Mohammad Hossein Hasheminejad
Realism in Action: Anomaly-Aware Diagnosis of Brain Tumors from Medical Images Using YOLOv8 and DeiT
Seyed Mohammad Hossein Hashemi - Leila Safari - Mohsen Hooshmand - Amirhossein Dadashzadeh Taromi
more
Samin Hamayesh - Version 44.5.0