أطروحات الدكتوراه (Doctoral Dissertations)
Permanent URI for this community
Browse
Browsing أطروحات الدكتوراه (Doctoral Dissertations) by Author "Eman Abed Al-Kareem Mousa Naser"
Now showing 1 - 1 of 1
Results Per Page
Sort Options
- ItemTowards Automated Arabic Synonyms Extraction(Al-Quds University, 2025-01-11) Eman Abed Al-Kareem Mousa Naser; ايمان عبد الكريم موسى نصرSynonyms extraction gains special attention as synonyms are essential in improving Natural Language Processing (NLP) application performance. The Lexical Substitution (LS) task is utilized for Synonym extraction, which generates a set of equivalent substitutions (i.e., synonyms) to the target word or phrase in a sentence that saves the sentence's meaning. This task can enhance writing, language understanding, and NLP models and address ambiguity. Recently, LS has attracted much attention in many languages. Despite the richness of Arabic vocabulary, limited research has been performed on the LS task due to the lack of annotated data. To bridge this gap, we present the first Arabic LS benchmark dataset, AraLexSubD for benchmarking LS pipelines. AraLexSubD is manually built by eight native Arabic speakers and linguists (six linguist annotators, a doctor, and an economist) who annotate the 630 sentences. AraLexSubD covers three domains: general, finance, and medical. It encompasses 2476 substitution candidates ranked according to their semantic relatedness. We also present an Arabic LS pipeline, AraLexSubPro, which offers different techniques for generating, selecting, and ranking substitutions. To make a thorough comparison, AraLexSubPro uses four different methods as baselines to generate substitute candidates for the target words: a synonym dictionary-based approach using Arabic Word Net (AWN), a pre-trained language model-based approach (AraBERT), AraBERT dropout (partial masking), and a hybrid approach between AraBERT and AWN. The results showed that the hybrid approach achieved the best results compared to the other approaches. The generated substitutions are filtered and then ranked based on six high-quality features to compare thoroughly: word similarity, word frequency, BERT prediction order (BERT probability), BERT-based language model (Loss), BERT similarity, and the BERTscore. The substitutions are then reranked based on our AraLexSubPro ranker. Additionally, an error analysis of the experiment is reported. To evaluate the AraLexSubPro pipeline, we use our first benchmark dataset for the Arabic LS task AraLexSubD dataset, which can automatically evaluate the Arabic LS systems. To our knowledge, this is the first study on Arabic lexical substitution. The results were encouraging and fundamental for Arabic LS research. To speed up research on this field, we have put the AraLexSubD data on GitHub at the following link: https://github.com/karajah2024/Arabic-Lexical-Substitution.git