0% Complete
Home
/
14th International Conference on Computer and Knowledge Engineering
AvashoG2P: A multi-module G2P Converter for Persian
Authors :
Ali Moghadaszadeh
1
Fatemeh Pasban
2
Mohsen Mahmoudzadeh
3
Maryam Vatanparast
4
Amirmohammad Salehoof
5
1- Part AI Research Center
2- Part AI Research Center
3- Ferdowsi University of Mashhad
4- Part AI Research Center
5- Part AI Research Center
Keywords :
TTS،G2P
Abstract :
The conversion of graphemes to phonemes (G2P) is a fundamental task in text-to-speech (TTS) and automatic speech recognition (ASR) systems. Over the years, G2P systems have evolved from rule-based and statistical methods to advanced neural network-based approaches. Despite these advancements, G2P conversion for Persian remains challenging due to the complex relationship between spelling and pronunciation and the scarcity of high-quality datasets. This paper introduces the AvashoG2P, a multi-module novel solution for Persian G2P conversion. The AvashoG2P system leverages a sequence-to-sequence (seq2seq) model with a GRU-based recurrent unit and an attention mechanism. This model is trained on both diacritized and non-diacritized words, enhancing its understanding of phonemes and their relationships. The system achieves a Word Error Rate (WER) of 15\% and a Phoneme Error Rate (PER) of 5\%, demonstrating its effectiveness. One of the critical components of AvashoG2P is its homograph disambiguation module, which utilizes a single model for all homographs, addressing a significant challenge in Persian text processing. Our method leverages a classification approach for homograph disambiguation, which assigns a phoneme label to the entire input window. Our system achieves high accuracy while optimizing for latency and memory consumption. We achieve significant improvements in accuracy and F1 scores using transformer-based models and machine learning classifiers. Our results highlight the superior performance of the XLMRoberta model among transformer models, with an F1 Weighted score of 94.7, and the SVC model among machine learning classifiers, with an F1 Weighted score of 89.96. Additionally, we present the AvashoG2P-Benchmark, a comprehensive test dataset designed to facilitate future research and benchmarking in Persian G2P tasks (available at: https://huggingface.co/datasets/PartAI/AvashoG2P-Benchmark).
Papers List
List of archived papers
Exploring 3D Transfer Learning CNN Models for Alzheimer’s Disease Diagnosis from MRI Images
Fatemehsadat Ghanadi Ladani - Hamidreza Baradaran Kashani
Designing a High Perfomance and High Profit P2P Energy Trading System Using a Consortium Blockchain Network
Poonia Taheri Makhsoos - Behnam Bahrak - Fattaneh Taghiyareh
An Efficient Planning Method for Autonomous Navigation of a Wheeled-Robot based on Deep Reinforcement Learning
Ali Salimi Sadr - Mahdi Shahbazi Khojasteh - Hamed Malek - Armin Salimi-Badr
Adaptive Multi-Scale Attentional Network for Semantic Segmentation of Remote Sensing Images
Melika Zare - Sattar Hashemi
A Federated Learning-Based Hybrid Deep Learning Framework for Enhanced Human Activity Recognition
Jamileh Azmoudeh - Sajjad Arghaee - Parisa Valizadeh - Samaneh Dandani - Iman Havangi - Mohammad Hossein Yaghmaee
FaaScaler: An Automatic Vertical and Horizontal Scaler for Serverless Computing Environments
Zahra Rezaei - Saeid Abrishami - Seid Nima Moeintaghavi
SCDS: A Secure Clustering Protocol Using Dempster-Shafer Theory for VANET in Smart City
Hoda Mosadegh - Nazbanoo Farzaneh
Fine-tuned Generative Adversarial Network-based Model for Medical Image Super-Resolution
Alireza Aghelan - Modjtaba Rouhani
I-ACS: An Improved Ant Colony System to Solve the Time-Dependent Orienteering Problem
Zahra Bakhshandeh - Morteza Keshtkaran
Adaptive-A-GCRNN: Enhancing Real-time Multi-band Spectrum Prediction through Attention-based Spatial-Temporal Modeling
Seyed majid Hosseini - Seyedeh Mozhgan Rahmatinia - Seyed Amin Hosseini Seno - Hadi Sadoghi yazdi
more
Samin Hamayesh - Version 41.7.6