Please wait ...
0% Complete
Home
/
13th International Conference on Computer and Knowledge Engineering
AgeNet-AT: An End-to-End Model for Robust Joint Speaker Age Estimation and Gender Recognition Based on Attention Mechanism and Titanet
Authors :
Mahsa Zamani Tarashandeh
1
Amirhossein Torkanloo
2
Mohammad Hossein Moattar
3
1- Department of Electrical Engineering, Faculty of Engineering Ferdowsi University of Mashhad
2- Department of Computer Engineering Ferdowsi University of Mashhad Mashhad, Iran
3- ,Department of Computer Engineering Mashhad Branch, Islamic Azad University, Mashhad, IRAN
Keywords :
Age estimation،Gender classification،Multi-task learning،Attention mechanism،Titanet
Abstract :
Speaker age estimation has become popular in recent years due to its potential applications in various fields, including forensics and human-computer interaction. However, noise and utterance length robustness is a key factor in the performance of the approaches. In this work, a robust age estimation and gender recognition model named AgeNet-AT is proposed based on an attention mechanism and Titanet model. The proposed approach applies Titanet as the embedding extractor, and attention mechanism to create an end-to-end architecture for age estimation. Since Titanet is a model designed to distinguish different speaker identities, it is hypothesized that some of its extracted features may contain properties that can differentiate speakers’ age and gender. Therefore, Titanet is chosen as the embedding approach in this study. Additionally, an attention layer is used to focus on the most valuable features for age estimation. Furthermore, an auxiliary task of gender classification is added to the model in order to improve the estimation performance. The experiments are conducted on TIMIT dataset for different evaluation conditions, such as various utterance lengths and noise levels. The experimental results indicate the robustness of the AgeNet-AT model. The model has outperformed the state-of-the-art age estimation results on TIMIT dataset with Root Mean Square Error (RMSE) of 5.92 and 6.85 and Mean Absolute Error (MAE) of 4.30 and 4.73 for male and female speakers, respectively.
Papers List
List of archived papers
Early detection of Parkinson’s disease using Convolutional Neural Networks on SPECT images
Reyhaneh Dehghan - Marjan Naderan - Seyyed Enayatallah Alavi
Blind image quality assessment based on Multi-resolution Local Structures
Seyed Majid Khorashadizadeh - Mehdi Sadeghi Bakhi - Fatemeh Seifishahpar - AliMohammad Latif
Lightweight Local Transformer for COVID-19 Detection Using Chest CT Scans
Hojat Asgarian Dehkordi - Hossein Kashiani - Amir Abbas Hamidi Imani - Shahriar Baradaran Shokouhi
Emotion Recognition In Persian Speech Using Deep Neural Networks
Ali Yazdani - Hossein Simchi - Yasser Shekofteh
Damage Detection After the Earthquake Using Sentinel-1 and 2 Images and Machine Learning Algorithms (Case Study: Sarpol-e Zahab Earthquake)
Niloofar Alizadeh - Behnam Asghari Beirami - Mehdi Mokhtarzade
Adaptive Ensemble Learning for Software Defect Prediction: A Dynamic Weighted Hybrid Model Using SVM, DT, and ANFIS-PSO
Mohsen EsfandyariDoulabi - Amin Esfandiyari Doulabi - Javad Khaligh
Atlas-based segmentation of cardiac chambers in systolic and diastolic phases of echocardiographic images
Elham Fathipour - Mahdi Saadatmand
ParsHomo: A T5-Powered Approach to High-Precision Persian Homograph Disambiguation
Hasan Jalali - Taha Mohaddesi
Joint mobility-aware offloading and UAV position optimization in Blockchain-enabled 5G
Zeinab Rabbani - Zeinab Movahedi
SingAll: Scalable Control Flow Checking for Multi-Process Embedded Systems
Mehdi Amininasab - Ahmad Patooghy - Mahdi Fazeli
more
Samin Hamayesh - Version 44.9.3