Modeling Adverse Events Detection as Time Series Using LSTM and Labeled Data Fusion

Date
2026-05-27
Authors
Rasha Zaki Mustafa Assaf
رشا زكي مصطفى عساف
Journal Title
Journal ISSN
Volume Title
Publisher
Al-Quds Univeersity
Abstract
Adverse events (AEs) , characterized as undesirable and unintentional outcomes of medical treatments, pose significant challenges in healthcare. Accurate prediction of these events is crucial for enhancing patient safety, optimizing resource allocation, and improving overall health- care outcomes. Traditional methods, such as statistical analyses and early machine learning techniques, often fail to capture complex, nonlinear relationships of medical data. This dissertation explores the application of advanced machine learning models, particularly Long Short- Term Memory (LSTM ) and transformers, to predict adverse events and outcome using Food and Drug Administration (FDA ) Adverse Event Reporting System ( FAERS), a database containing information on drug safety, and Medical Entities in Digital Records Annotated ( Med- DRA) datasets. The dissertation focuses on integrating these comprehensive datasets, followed by data cleaning and pre-processing to ensure data accuracy and reliability. Latent Dirichlet Al- location( LDA) and clustering, a statistical method for topic modeling was used to address the complexity of classifying over 50,000 adverse event categories, reducing the number of classes to 10. The study further investigates the efficacy of sequence-to-sequence models, such as the Transformer, in predicting adverse events. Results indicate that Transformer-based models out- perform traditional machine learning algorithms in predicting adverse events, demonstrating improved accuracy and robustness. The sequence-to-sequence models achieved an accuracy of 92.81%, with a precision of 93.22%, a recall of 91.89%, and an F1 score of 92.55%. The integration of self-attention mechanisms enhances these models by allowing them to capture complex patterns and relationships within the data—insights that traditional approaches often overlook s. We implemented a Graph Neural Network ( GNN) and evaluated precision and re- call across multiple categories, yielding low overall F1 scores of 0.36 for term type classification and 0.42 for FAERS reaction classification.
تُعد الأحداث السلبية (AEs)، التي تُعرف بأنها نتائج غير مرغوب فيها وغير مقصودة للعلاجات الطبية، من التحديات الكبيرة في قطاع الرعاية الصحية. إذ يُعد التنبؤ الدقيق بهذه الأحداث أمرًا بالغ الأهمية لتعزيز سلامة المرضى، وتحسين تخصيص الموارد، والارتقاء بنتائج الرعاية الصحية بشكل عام. إلا أن الطرق التقليدية، كتحليل البيانات الإحصائي والتقنيات المبكرة للتعلم الآلي، غالبًا ما تفشل في التقاط العلاقات المعقدة وغير الخطية في البيانات الطبية. يستعرض هذا البحث تطبيق نماذج التعلم الآلي المتقدمة، وخاصة نماذج الذاكرة طويلة وقصيرة المدى (LSTM) ونماذج المحولات (Transformers)، في التنبؤ بالأحداث السلبية والنتائج الطبية باستخدام قاعدة بيانات نظام الإبلاغ عن الأحداث السلبية لإدارة الغذاء والدواء الأمريكية (FAERS)، والتي تحتوي على معلومات تتعلق بسلامة الأدوية، إلى جانب بيانات الكيانات الطبية المرمزة (MedDRA). يركّز البحث على دمج هذه المصادر البيانية الشاملة، تليها عمليات تنظيف ومعالجة للبيانات لضمان الدقة والموثوقية. وللتعامل مع تعقيد تصنيف أكثر من 50,000 فئة من فئات الأحداث السلبية، تم استخدام طريقة النمذجة الموضوعية (Latent Dirichlet Allocation - LDA) وتقنيات التجميع (clustering) لتقليل عدد الفئات إلى 10 فئات رئيسية. كما تناولت الدراسة فعالية نماذج التسلسل إلى التسلسل (sequence-to-sequence)، مثل المحولات، في التنبؤ بالأحداث السلبية. وقد أظهرت النتائج أن النماذج المعتمدة على المحولات تتفوق على خوارزميات التعلم الآلي التقليدية من حيث الدقة والموثوقية، حيث حققت هذه النماذج دقة بلغت 92.81%، ودقة إيجابية (precision) بنسبة 93.22%، واسترجاع (recall) بنسبة 91.89%، ودرجة F1 بلغت 92.55%. كما أن دمج آلية الانتباه الذاتي (self-attention) ساعد هذه النماذج في التقاط الأنماط والعلاقات المعقدة في البيانات—وهي علاقات غالبًا ما تغفلها الأساليب التقليدية. إضافة إلى ذلك، تم تطبيق شبكة العصبونات البيانية (Graph Neural Network - GNN)، وتقييم مقاييس الدقة والاسترجاع عبر عدة فئات، لكنها حققت درجات F1 منخفضة عمومًا، بلغت 0.36 لتصنيف نوع المصطلح، و0.42 لتصنيف التفاعلات في قاعدة بيانات FAERS.
Description
Keywords
Citation