Please wait ...
0% Complete
Home
/
13th International Conference on Computer and Knowledge Engineering
Distilling Knowledge from CNN-Transformer Models for Enhanced Human Action Recognition
Authors :
Hamid Ahmadabadi
1
Omid Nejati Manzari
2
Ahmad Ayatollahi
3
1- School of Electrical Engineering, Iran University of Science and Technology, Tehran, Iran
2- School of Electrical Engineering, Iran University of Science and Technology, Tehran, Iran
3- School of Electrical Engineering, Iran University of Science and Technology, Tehran, Iran
Keywords :
Action Recognition،Still Image،Deep Learning،Knowledge Distillation
Abstract :
This paper presents a study on improving human action recognition through the utilization of knowledge distillation, and the combination of CNN and ViT models. The research aims to enhance the performance and efficiency of smaller student models by transferring knowledge from larger teacher models. The proposed method employs a Transformer vision network as the student model, while a convolutional network serves as the teacher model. The teacher model extracts local image features, whereas the student model focuses on global features using an attention mechanism. The Vision Transformer (ViT) architecture is introduced as a robust framework for capturing global dependencies in images. Additionally, advanced variants of ViT, namely PVT, Convit, MVIT, Swin Transformer, and Twins, are discussed, highlighting their contributions to computer vision tasks. The ConvNeXt model is introduced as a teacher model, known for its efficiency and effectiveness in computer vision. The paper presents performance results for human action recognition on the Stanford 40 dataset, comparing the accuracy and mAP of student models trained with and without knowledge distillation. The findings illustrate that the suggested approach significantly improves the accuracy and mAP when compared to training networks under regular settings. These findings emphasize the potential of combining local and global features in action recognition tasks.
Papers List
List of archived papers
Object Detection on Detecting Skin Lesion using Dab-DETR
Sheida Shadman - Amirreza Rouhbakhshmeghrazi - Shayan Nalbandian - Bo Li - Shaghayegh Shadman - Malik Muhammad Owais Siddique
Financial Market Prediction Using Deep Neural Networks with Hardware Acceleration
Dara Rahmati - Mohammad Hadi Foroughi - Ali Bagherzadeh - Mehdi Foroughi - Saeid Gorgin
A New Hypercube Variant: Pruned Shuffle Connected Cube
Reza Latifi - Mahmoud Naghibzadeh
Hybrid Vision Transformer for Detection of Dentigerous Cysts in Dental Radiography Images
Reza Tavasoli - Arya VarastehNezhad - Hamed Farbeh
A New Inter-layer Similarity metric for link prediction in multilayer networks
Alireza Abdollahpouri - Samira Rafiee
MultiPath ViT OCR: A Lightweight Visual Transformer-based License Plate Optical Character Recognition
Alireza Azadbakht - Saeed Reza Kheradpisheh - Hadi Farahani
Identifying novel disease genes based on protein complexes and biological features
Mahshad Hashemi - Eghbal Mansoori
A Review on Machine Learning Methods for Workload Prediction in Cloud Computing
Mohammad Yekta - Hadi Shahriar Shahhoseini
Lossless Watermarking in Encrypted Triangular Mesh Models Based on Optimized Vertex Estimation and Error Histogram Shifting
Alireza Ghaemi - Habibollah Danyali - Kamran Kazemi - Zahra Qodrati - Amirhossein Ghaemi - Seyedeh Masoumeh Taji
FedFog: A Serverless and Privacy-Aware Federated Learning Simulator for Edge–Fog Networks
Seyed Vahid Hashemi Nik - Seyed Mohammad Mahdi Asaadi - Somayeh Sobati-M
more
Samin Hamayesh - Version 44.9.3