0% Complete
Home
/
14th International Conference on Computer and Knowledge Engineering
TriMAE: Fashion visual search with Triplet Masked Auto Encoder Vision Transformer
Authors :
Lachin Zamani
1
Reza Azmi
2
1- Department of Computer Engineering, Faculty of Engineering, Alzahra University, Tehran, Iran
2- Department of Computer Engineering, Faculty of Engineering, Alzahra University, Tehran, Iran
Keywords :
Visual Search،Triplet Network،Masked Auto Encoders Vision Transformer
Abstract :
Visual search is a technology that identifies images similar to a provided query image and presents results ranked by similarity. In the realm of apparel, this innovative tool revolutionizes shopping by enabling users to effortlessly find desired items based on visual preference. Visual search remains a challenging problem despite its potential to significantly enhance user experience. The existence of differences in minute details, the presence of multiple garments in a single image, discrepancies between user-taken and catalog images, and the inherent flexibility of clothing are among the challenges associated with this issue. By selecting robust features and improving the learning of similarity and dissimilarity between images, superior results can be obtained. Consequently, a method has been proposed to yield enhanced outcomes. Convolutional Neural Networks and Vision Transformers are commonly used as the backbone of triplet neural networks for visual search tasks. These networks are designed to better learn the similarities and differences between images. In this research, we employ a combination of triplet neural networks and a masked auto-encoder vision transformer model. A triplet loss function is used during network training to learn the similarity between images. We evaluate our method on the DeepFashion In-shop dataset, which comprises different categories of clothing images. Through extensive experiments on this benchmark, our model achieves an impressive Recall@1 of 93.2% for visual search.
Papers List
List of archived papers
Automating Theory of Mind Assessment with a LLaMA-3-Powered Chatbot: Enhancing Faux Pas Detection in Autism
Avisa Fallah - Ali Keramati - Mohammad Ali Nazari - Fatemeh Sadat Mirfazeli
Object Detection on Detecting Skin Lesion using Dab-DETR
Sheida Shadman - Amirreza Rouhbakhshmeghrazi - Shayan Nalbandian - Bo Li - Shaghayegh Shadman - Malik Muhammad Owais Siddique
IranITJobs2021: a Dataset for Analyzing Iranian Online IT Job Advertisements Collected Using a New Crowdsourcing Process
Fakhroddin Noorbehbahani - Nikta Akbarpour - Mohammad Reza Saeidi
Optimization of quantum secret sharing communication using corresponding bits
Mahsa Khorrampanah - Mohammad Bolokian - Monireh Houshmand
Towards Efficient Video Object Detection on Embedded Devices
Mohammad Hajizadeh - Adel Rahmani - Mohammad Sabokrou
Attention Transfer in Self-Regulated Networks for Recognizing Human Actions from Still Images
Masoumeh Chapariniya - Sara Vesali Barazande - Seyed Sajad Ashrafi - Shahriar B.Shokouhi
Early detection of Parkinson’s disease using Convolutional Neural Networks on SPECT images
Reyhaneh Dehghan - Marjan Naderan - Seyyed Enayatallah Alavi
Atlas-based segmentation of cardiac chambers in systolic and diastolic phases of echocardiographic images
Elham Fathipour - Mahdi Saadatmand
Enhanced Duplicate Bug Report Detection in Anonymized Environments: A Parallelized Multi-Task Learning Framework
Alireza Shorafa - Abolfazl Zarghani
A Chaotic Crow Search Algorithm for Overlapping Clustering
Mostafa Sabzekar - Seyed Vahid Mousavainejad
more
Samin Hamayesh - Version 44.5.0