Please wait ...
0% Complete
Home
/
11th International Conference on Computer and Knowledge Engineering
A Language-Independent Approach to Classification of Textual File Fragments: Case Study of Persian, English, and Chinese Languages
Authors :
Fatemeh Mansouri Hanis
1
Hamidreza Khoshvaghti
2
Mehdi Teimouri
3
Hadi Veisi
4
1- University of Tehran
2- University of Tehran
3- University of Tehran
4- University of Tehran
Keywords :
Classification, File fragments, language-dependent file type identification, textual files, context language, file format.
Abstract :
With the advent of communications systems in recent decades, the transmission of electronic files on computer networks has dramatically increased. In this situation, identifying the type of files is important in many applications such as digital forensics and file carving. The state-of-the-art methods for identifying the file type of a file fragment are based on the content of the fragments. To the best of the authors' knowledge, there is no study addressing the effect of context language in identifying the file type of textual file fragments. In this paper, we have considered a machine learning approch for the classification among five types of common text file formats: PDF, DOC, DOCX, RTF, and TXT. Also, we have examined the effect of context language on the classification of the file fragments. Two scenarios are considered. In the first one, the language for both training and testing phases are the same, that the best results are achieved; the accuracies of the test for Persian, English, and Chinese languages are 85.6%, 76.4%, 86.1%, respectively. In the second scenario, the languages of training and testing sets are not the same, in which the training is done using one language and the evaluation is performed on the two other languages. In this case, the average accuracy values for Persian, English, and Chinese languages are 60.0%, 58.5%, and 71.4%, respectively. The evaluations of the second scenario show that the language-independent machine learning approach is robust in the identification of DOC, DOCX, and RTF formats.
Papers List
List of archived papers
Persis: A Persian Font Recognition Pipeline Using Convolutional Neural Networks
Mehrdad Mohammadian - Neda Maleki - Tobias Olsson - Fredrik Ahlgren
A Novel Method For Fake News Detection Based on Propagation Tree
Mansour Davoudi - Mohammad Reza Moosavi - Mohammad Hadi Sadreddini
DRL-based Decision-Making for Autonomous Vehicle Collision Avoidance
Hoda Gholamrezaee - Seyedreza Taghizadeh - Ali Honarjoo
Adaptive Pronunciation Scoring: Aligning Automated Assessments with Human Expert Evaluations
Omid Aghdaei - Mohammad Sadegh Safari - Mohammad Hassan Rasoolizadeh - Abedeh Mirzaee
Analyzing the Impact of COVID-19 on Economy from the Perspective of User’s Reviews
Fatemeh Salmani - Hamed Vahdat-Nejad - Hamideh Hajiabadi
ExaASC: A General Target-Based Stance Detection Corpus in Arabic Language
Mohammad Mehdi Jaziriyan - Ahmad Akbari - Hamed Karbasi
Enhancing Lighter Neural Network Performance with Layer-wise Knowledge Distillation and Selective Pixel Attention
Siavash Zaravashan - Sajjad Torabi - Hesam Zaravashan
Adaptive-A-GCRNN: Enhancing Real-time Multi-band Spectrum Prediction through Attention-based Spatial-Temporal Modeling
Seyed majid Hosseini - Seyedeh Mozhgan Rahmatinia - Seyed Amin Hosseini Seno - Hadi Sadoghi yazdi
DTranIDS: A Two-Tiered Intrusion Detection System for RPL-based IoT Networks based on Decision Tree and Transformer Models
Mohammad Fazeli - Mohsen Raji - Mohammad Mahdi Fazeli
A Comparative Analysis of Clinical Note Categories for Mortality Prediction in ICU Patients
Maryam Karrabi - Mohsen Kahani - Mina Afzali - Nadieh Armin
more
Samin Hamayesh - Version 44.9.3