0% Complete
Home
/
11th International Conference on Computer and Knowledge Engineering
A Language-Independent Approach to Classification of Textual File Fragments: Case Study of Persian, English, and Chinese Languages
Authors :
Fatemeh Mansouri Hanis
1
Hamidreza Khoshvaghti
2
Mehdi Teimouri
3
Hadi Veisi
4
1- University of Tehran
2- University of Tehran
3- University of Tehran
4- University of Tehran
Keywords :
Classification, File fragments, language-dependent file type identification, textual files, context language, file format.
Abstract :
With the advent of communications systems in recent decades, the transmission of electronic files on computer networks has dramatically increased. In this situation, identifying the type of files is important in many applications such as digital forensics and file carving. The state-of-the-art methods for identifying the file type of a file fragment are based on the content of the fragments. To the best of the authors' knowledge, there is no study addressing the effect of context language in identifying the file type of textual file fragments. In this paper, we have considered a machine learning approch for the classification among five types of common text file formats: PDF, DOC, DOCX, RTF, and TXT. Also, we have examined the effect of context language on the classification of the file fragments. Two scenarios are considered. In the first one, the language for both training and testing phases are the same, that the best results are achieved; the accuracies of the test for Persian, English, and Chinese languages are 85.6%, 76.4%, 86.1%, respectively. In the second scenario, the languages of training and testing sets are not the same, in which the training is done using one language and the evaluation is performed on the two other languages. In this case, the average accuracy values for Persian, English, and Chinese languages are 60.0%, 58.5%, and 71.4%, respectively. The evaluations of the second scenario show that the language-independent machine learning approach is robust in the identification of DOC, DOCX, and RTF formats.
Papers List
List of archived papers
An Evolutionary Approach with Surrogate Models for Feature Selection in Intrusion Detection Systems
Sadeq Moradi - Hadi Shahriar Shahhoseini
Depression Diagnosis Using Optimization of Nonlinear EEG Features Based on Parametric Learning Tactics
Ali Asadi Zeidabadi - Melika Changizi - Mahdi Zolfagharzadeh Kermani - Sara Bargi Barkouk
A Stacking Ensemble Framework for Ransomware Detection on the Bitcoin Blockchain Using Transaction Graph Analytics
Mohammad Mobin Teymourpour - Parsa Hedayatnia - Mohammad Allahbakhsh - Haleh Amintoosi
IranITJobs2021: a Dataset for Analyzing Iranian Online IT Job Advertisements Collected Using a New Crowdsourcing Process
Fakhroddin Noorbehbahani - Nikta Akbarpour - Mohammad Reza Saeidi
DFIG-WECS Renewable Integration to the Grid and Stability Improvement through Optimal Damping Controller Design
Theophilus Ebuka Odoh - Aliyu Sabo - Hossien Shahinzadeh - Noor Izzri Abdul Wahab - Farshad Ebrahimi
Adversarial Robustness Evaluation with Separation Index
Bahareh Kaviani Baghbaderani - Afsaneh Hasanebrahimi - Ahmad Kalhor - Reshad Hosseini
Practical Implementation of Real-Time Waste Detection and Recycling based on Deep Learning for Delta Parallel Robot
Hasan Jalali - Shaya Garjani - Ahmad Kalhor - Mehdi Tale Masouleh - Parisa Yousefi
Cross-project Defect Prediction with An Enhanced Transfer Boosting Algorithm
Nazgol Nikravesh - Mohammad Reza Keyvanpour
Predicting the Recovery Rate of COVID-19 Using a Novel Hybrid Method
Fatemeh Ahouz - Ebrahim Sayahi
Fatty Liver Level Recognition Using Particle Swarm Optimization (PSO) Image Segmentation and Analysis
Seyed Muhammad Hossein Mousavi - Vyacheslav Lyashenko - Atiye Ilanloo - S. Younes Mirinezhad
more
Samin Hamayesh - Version 44.5.0