0% Complete
Home
/
15th International Conference on Computer and Knowledge Engineering
Persian Legal Text Simplification Leveraging Transformer-Based Models
Authors :
Mohammadreza Joneidi Jafari
1
Saedeh Tahery
2
Amirhossein Nikoofard
3
1- Department of Electronics, Faculty of Electrical Engineering
2- Department of Artificial Intelligence, Faculty of Computer Engineering
3- Department of Systems and Control, Faculty of Electrical Engineering
Keywords :
Text Simplification،Persian Legal Documents،ChatGPT،Large Language Models،Transformer-Based Models
Abstract :
Legal documents often use complex and domain-specific language, which limits their accessibility to the general public. Despite growing interest in text simplification within natural language processing, the legal domain in Persian remains largely unexplored due to the lack of annotated corpora and domain-adapted models. This study presents a practical approach to Persian legal text simplification by leveraging synthetic supervision alongside fine-tuned transformer-based models. This paper creates a new labeled dataset by generating simplified versions of legal rulings using ChatGPT, and validate a subset of the outputs through expert review to ensure data quality. To address resource constraints common in real-world applications, we fine-tune lightweight encoder-decoder models, enabling efficient deployment without requiring large-scale annotation or extensive inference infrastructure. Our results show that a compact model such as ParsT5 outperforms zero-shot large language models like PersianLLaMA. The model is also enhanced with an existing attention extension, which enables efficient processing of long inputs without truncation. As a result, this work introduces the first benchmark for Persian legal text simplification, demonstrating that well-adapted, efficient models can achieve high performance in low-resource and domain-specific scenarios. By taking this solid first step, it paved the way for future research and development in natural language processing for Persian legal texts. The released dataset and code are publicly available at https://github.com/mrjoneidi/Simplification-Legal-Texts.
Papers List
List of archived papers
Parallel Local Feature Selection For High-dimensional Data
Zhaleh Manbari - Chiman Salavati - Fardin AkhlaghianTab - Barzan Saeedpoor - Himan Delbina - Mahmud Abdulla Mohammad
Intracranial Hemorrhage Classification using CBAM Attention Module and Convolutional Neural Networks
Parnian Rahimi - Marjan Naderan - Amir Jamshidnezhad - Shahram Rafie
African Vultures Optimization Algorithm for Optimal Damping Controllers Design in the Electrical Power Grid System
Aliyu Sabo - Theophilus Ebuka Odoh - Samuel Habu - Hossein Shahinzadeh - Farshad Ebrahimi
Reliability Evaluation of 4:2 Compressors Based on Hammock Networks
Farshad Safaei - Mohammad mahdi Emadi Kouchak - Sara Talebpour
An intelligent linguistic error detection approach to automated diagnosis of Dyslexia disorder in Persian speaking children
Fatemeh Asghari - Mahsa Khorasani - Mohsen Kahani - Seyed Amir Amin Yazdi - Mahdi Arkhodi Ghalenoei
MIPS-Core Application Specific Instruction-Set Processor for IDEA Cryptography − Comparison between Single-Cycle and Multi-Cycle Architectures
Ahmad Ahmadi - Reza Faghih Mirzaee
Leveraging the Power of Object Detection Models in Identifying Litter for a Significant Reduction in Environmental Pollution
Lim Zhen Xian - Ervin Gubin Moung - Jason Teo Tze Wi - Nordin Saad - Farashazillah Yahya - Tiong Lin Rui - Ali Farzamnia
Leveraging a structure-based and learning-based predictor using various feature groups in bioinformatics (case study: protein-peptide region residue-level interaction)
Shima Shafiee - Abdolhossein Fathi
Low-Cost and Hardware Efficient Implementation of Pooling Layers for Stochastic CNN Accelerators
Mobin Vaziri - Hadi Jahanirad
Virtual machine consolidation using SLA-aware genetic algorithm placement for data centers with non-stationary workloads
Hossein Monshizadeh Naeen
more
Samin Hamayesh - Version 44.5.0