Computer Science and Engineering

Permanent URI for this collectionhttps://hdl.handle.net/10323/11886

Browse

Recent Submissions

Now showing 1 - 20 of 25
  • Item type: Item ,
    Advancing Software Testing: Mutation Testing in Actor Concurrency and Empirical Insights into Machine Learning Test Practices
    (2026-01-01) Moradi Moghadam, Mohsen; Bagherzadeh, Mehdi; Ming, Hua; Sen, Amartya; Srauy, Sam
    Software testing, especially in large-scale projects, is becoming increasingly important in real-world and mission-critical software. This requires the tests to be highly effective at finding bugs that occur in the real world. However, developers find designing such effective tests challenging. My research aims to tackle this challenge by advancing software testing in two key domains: actor concurrency and machine learning that has been widely used across web, mobile, and desktop applications and adopted at scale by companies like Google and Facebook. This dissertation presents µAkka, a framework for mutation testing of Akka actor concurrency using real actor bugs. The research analyzes 186 real Akka bugs, designs 32 mutation operators, and implements them in an Eclipse plugin. µAkka generates 11.7k mutants of 10 GitHub applications and runs 7.9k tests. The evaluation compares results to PIT with 26.2k mutants. µAkka mutants are higher quality, cover more bugs, and tests are less effective in detecting them. Additionally, this dissertation includes a large-scale study on a set of real-world ML test cases in GitHub to understand their topics, popularities, difficulties, correlations, and historical changes. The study curates a set of 2,525 ML projects and their 136,463 test cases from GitHub; uses topic modeling to group these test cases into ML test topics; groups similar topics into an ML test topic hierarchy; discusses these ML topics using sample test cases; analyze the popularity and difficulty of these topics; and study the correlation and historical trends over the past 10 years. Together, these efforts address two critical domains: actor concurrency and machine learning. In actor concurrency, this work introduce a mutation testing framework designed to measure and strengthen test suites. In machine learning, we provide empirical insights into testing practices and gaps through a large-scale analysis. Both contributions advance understanding of how software is tested in areas of growing importance.
  • Item type: Item ,
    Leveraging Large Language Models for Adaptive and Secure UAV Control
    (2026-01-01) Dharmalingam, Balakrishnan; Liu, Anyi; Ganesan, Subra; Wardat, Mohammad; Deng, Xiaodong
    Unmanned Aerial Vehicles (UAV) are increasingly integral to mission-critical operations, spanning public safety, decentralized logistics, and autonomous surveillance. However, as UAV operations become highly autonomous within contested environments, their reliance on unencrypted and unauthenticated communication standards exposes a severe cyber-physical attack surface. Kinetic cyber-attacks targeting these vulnerabilities—such as sensor spoofing and Man-in-the-Middle (MITM) intrusions—can yield catastrophic physical consequences. While current research proposes securing UAV operations via deterministic heuristics or conventional deep learning models (e.g., Long Short Term Memory (LSTM), Autoencoders), these methods struggle to adapt to dynamic zero-day threats and operate as opaque “black boxes,” lacking the semantic explainability required for safety-critical aviation. Large Language Models (LLM) demonstrate unprecedented capabilities in contextual reasoning, zero-shot anomaly detection, and semantic explainability. Yet, their deployment in autonomous aviation is severely bottlenecked by the strict Size, Weight, and Power (SWaP) constraints of UAV microcontrollers, necessitating highly optimized, domain-specific training paradigms. To resolve this dichotomy, this dissertation introduces Aero-LLM, a novel, distributed cyber-physical security framework that harnesses the cognitive reasoning of LLMs to secure UAV operations. The proposed architecture is comprised of four primary subsystems: (1) A Data Collector that synthesizes the Heterogeneous Generative Dataset for UASes (HGDAVE) utilizing Digital Twins (Software-In-The-Loop (SITL)/Hardware-In-The-Loop (HITL)), firmware fuzzing, and AI-generated MITM attacks orchestrated by a generative adversary (Net-GPT); (2) A Fine-Tuner that leverages the Zero Redundancy Optimizer (ZeRO), Parameter-Efficient-Fine-Tuning (PEFT), and Reinforment Learing from Human Feedback (RLHF) to compress and align models for the aviation domain; (3) An Online Controller deploying specialized models across an Edge-Fog-Cloud distributed architecture for real-time threat mitigation; and (4) An Offline Processor that continuously calibrates thresholds to harden the system against emerging attack vectors. This research contributes a unified cyber-physical data collection pipeline, a resource-efficient generative threat model, and a scalable LLM deployment strategy for embedded flight systems. Experimental evaluations demonstrate the profound efficacy of the proposed architectures. The Net-GPT offensive module achieved an average payload synthesis accuracy of 95.30%, successfully mimicking protocol-compliant network traffic. Defensively, the Aero-LLM anomaly detection framework achieved an accuracy of 92.60%, with a precision of 92.70%, a recall of 92.06%, and an F1-score of 90.55%. While the generative inference introduces a latency overhead of approximately 850ms compared to lightweight neural networks, the framework provides unparalleled semantic explainability and cyber-resilience, establishing a scalable foundation for secure, intelligent autonomous UAV operations.
  • Item type: Item ,
    A Trust-Centric Framework for Secure, User-Centric AI Systems with Actionable and Personalized Recourse
    (2026-01-01) Oyeniyi, Oluwafeyisayo O; Sen, Amartya; Xu, Lanyu; Raj, Sunny; Caushaj, Eralda
    Artificial Intelligence (AI) applications increasingly mediate high-stakes decisions in domains such as finance, healthcare, and social welfare. Yet their deployment often exposes users and institutions to risks arising from bias, opacity, and technical vulnerabilities. High predictive accuracy alone fails to ensure trust, particularly in contexts where user engagement, recourse, and accountability are critical. These challenges motivate a principled, layered approach to operationalizing trust in AI. This dissertation presents a trust-centric AI framework that formalizes trust as a multi-layer architectural objective. Trust is structured across three interdependent layers: technical trust, ensuring model integrity, security, and resilience against adversarial attacks; cognitive and practical trust, realized through causally grounded, feasible, cost-aware, and interpretable counterfactual explanations; and user-aligned trust, achieved by leveraging preference-aware, lexicographic optimization that personalizes actionable recourse. By aligning these layers through secure deployment mechanisms, causal verification, and multi-objective optimization, the framework shows how failures in security, explainability, or personalization can be systematically mitigated within a coherent trust-centric design. The framework is evaluated within tabular and structured decision-making environments, providing a testbed for secure model deployment, causally grounded explanations, and preference-aware actionable recourse. These contexts illustrate how layered trust can reconcile predictive accuracy, interpretability, and user agency under uncertainty and high personal stakes. By embedding verifiable, actionable, and user-aligned trust into AI-driven systems, this work supports responsible deployment in high-stakes domains and provides technical guidance for policymakers, organizations, and system architects seeking robust, transparent, and user-centered AI.
  • Item type: Item ,
    Graded Quantum Codes
    (2026-01-01) Shaska, Tanush; Ma, Tianle; Liu, Anyi; Qiang, Yao
    This thesis introduces Quantum Weighted Algebraic Geometry Codes (QWAGs), a novel class of quantum error-correcting codes derived from weighted superelliptic curves over finite fields. By extending classical algebraic geometry codes, the weighted framework incorporates graded rings and orbifold corrections, enhancing parameter flexibility and self-orthogonality for Calderbank–Shor–Steane (CSS) constructions. We develop divisor theory, Riemann–Roch spaces, and duality in quasi-smooth weighted settings, proving Euclidean duality via canonical divisors and residue pairings. These foundations enable QWAC families with parameters shaped by graded geometry. A homological perspective through evaluation chain complexes elucidates CSS conditions, leading to a refined quantum Singleton bound incorporating orbifold terms. Complementing the theory, we present a Python-based computational framework automating curve construction, point enumeration, Riemann–Roch computations, and CSS verification. This work unifies weighted algebraic geometry, coding theory, and quantum stabilizers, advancing tools for quantum error correction.
  • Item type: Item ,
    Advancing Software Testing: Mutation Testing in Actor Concurrency and Empirical Insights into Machine Learning Test Practices
    (2026-01-01) Moradi Moghadam, Mohsen; Bagherzadeh, Mehdi; Ming, Hua; Sen, Amartya; Srauy, Sam
    Software testing, especially in large-scale projects, is becoming increasingly important in real-world and mission-critical software. This requires the tests to be highly effective at finding bugs that occur in the real world. However, developers find designing such effective tests challenging. My research aims to tackle this challenge by advancing software testing in two key domains: actor concurrency and machine learning that has been widely used across web, mobile, and desktop applications and adopted at scale by companies like Google and Facebook. This dissertation presents µAkka, a framework for mutation testing of Akka actor concurrency using real actor bugs. The research analyzes 186 real Akka bugs, designs 32 mutation operators, and implements them in an Eclipse plugin. µAkka generates 11.7k mutants of 10 GitHub applications and runs 7.9k tests. The evaluation compares results to PIT with 26.2k mutants. µAkka mutants are higher quality, cover more bugs, and tests are less effective in detecting them. Additionally, this dissertation includes a large-scale study on a set of real-world ML test cases in GitHub to understand their topics, popularities, difficulties, correlations, and historical changes. The study curates a set of 2,525 ML projects and their 136,463 test cases from GitHub; uses topic modeling to group these test cases into ML test topics; groups similar topics into an ML test topic hierarchy; discusses these ML topics using sample test cases; analyze the popularity and difficulty of these topics; and study the correlation and historical trends over the past 10 years. Together, these efforts address two critical domains: actor concurrency and machine learning. In actor concurrency, this work introduce a mutation testing framework designed to measure and strengthen test suites. In machine learning, we provide empirical insights into testing practices and gaps through a large-scale analysis. Both contributions advance understanding of how software is tested in areas of growing importance.
  • Item type: Item ,
    Topology-Aware Correction of Retinal Vessel Graphs Using Graph Neural Networks
    (2026-01-01) Kssayrawi, Fatima Mohamad Omar; Qu, Guangzhi; Ma, Tianle; Zhao, Kaiqi
    Vessel graphs are commonly used to study the structure of retinal vessels, but graphs obtained from vessel segmentation often contain topology errors, including broken vessels and spurious connections. These errors arise from limitations in pixel level processing and negatively impact graph-based analysis.Our work addresses this problem by formulating topology correction as a graph-level learning task. We represent connectivity errors as missing or incorrect edges and correct them through a candidate edge prediction task. A perturbation framework is introduced to simulate realistic structural errors, enabling supervised learning across various scenarios. A graph neural network is used to encode both the geometry and the graph structure of the vessels, enabling the prediction of valid vessel connections. Experimental results show that this approach improves the structural consistency of vessel graphs and enhances the accuracy of connectivity modeling.
  • Item type: Item ,
    From Perceived Trust to Provable Security: A Multi-Tiered Framework for the Design, Assessment, and Verification of Security in Resource-Constrained IOT Ecosystems
    (2026-01-01) Wakam Younang, Victorine Clotilde; Sen, Amartya; Fu, Huirong; Raj, Sunny; Drignei, Dorin
    The widespread adoption of Internet of Things (IoT) devices and Autonomous Vehicles (AVs) has created a significant assurance gap between what users expect and what is technically feasible in resource-limited environments. As these systems transition from prototypes to critical infrastructure, it becomes essential not only to make them efficient, but also to ensure they are demonstrably secure and capable of earning public trust. This dissertation introduces a multi-layered research framework that tackles this gap across four stages of the technology lifecycle: socio-technical, operational, security-evaluative, and formal-foundational.The first phase situates the problem by examining the disconnect between governmental policies and end-user priorities related to AV safety and privacy. After identifying that public trust is frequently weakened by opaque automated decision-making, the work moves in the second phase, which focuses on security, especially assessing security risk in IoT networks, by leveraging complex numbers Bayesian Attack Graphs to model and quantify unpredictable attacker behavior. Once the risk can be assessed, the most common step is to mitigate the risk. The work focuses on assessing how realistically security models can be deployed on edge devices, especially Intrusion Detection Systems (IDS) models, on vehicles CAN buses, introducing new feasibility metrics, that account for the IDS models accuracy, as well as the CAN bus resources restrictions. The final phase directly addresses the core of the security assurance gap by connecting imperative software implementations to rigorous mathematical proofs. It introduces the Tiered Isomorphic Alignment (TIA) framework, which helps automate the translation of Python programs into Isabelle/HOL formal specifications, and introduces CryptoGraph, a neuro-symbolic system that employs Graph Attention Networks to analyze the security of non-standard lightweight cryptographic ciphers. Taken together, these contributions establish a scalable, end-to-end methodology for designing the next generation of IoT ecosystems, one that is grounded in mathematical correctness and automated security assessment, capable of maintaining and enhancing public trust.
  • Item type: Item ,
    Deep Learning Methods for Predicting Alzheimer’s Disease Progression: Transformer and Graph Neural Network Approaches
    (2026-01-01) Moghaddami, Mahdi; Siadat, Mohammad-Reza; Ma, Tianle; Li, Jia
    Alzheimer’s disease (AD) is a progressive neurodegenerative disorder affecting millions globally. Early detection of disease progression is critical for timely intervention. This thesis presents two complementary deep learning approaches for predicting AD progression using longitudinal data. The first approach employs a Transformer-based model that leverages visit history features, including cognitive test scores and imaging features. The second approach uses a history-aware graph neural network (HA-GNN) that operates on functional connectivity derived from resting-state functional MRI data. Both approaches address key challenges in longitudinal prediction, including irregular visit spacing, missing data, and class imbalance. Results demonstrate that Transformer-based models achieve superior performance on multi-class diagnosis prediction (82.4 accuracy), particularly for identifying disease converters. The HA-GNN model achieves 82.9 accuracy in binary conversion prediction, offering potential for early intervention. These findings emphasize the importance of model selection in predicting cognitive decline, with implications for clinical decision-making and patient outcomes.
  • Item type: Item ,
    A Cloud-Based Kalman Filter Approach for Noise Reduction in Diagnostic Trouble Codes (DTCS)
    (2026-01-01) Taleb, Yamen; Zohdy, Mohamed; Louis, Steven; Kaur, Amanpreet; Alwerfali, Daw
    This dissertation presents a cloud-based Kalman filter framework designed to reduce noise in diagnostic trouble code (DTC) signals and improve First Notice of Loss (FNOL) accuracy within connected vehicle systems. Modern vehicles generate extensive CAN-based diagnostic and sensor data, yet these signals are often influenced by noise, transient disturbances, and low-severity impacts that can trigger false DTC activations and unnecessary FNOL alerts. These inaccuracies create operational inefficiencies for automative OEMs and contribute to increased claim volumes, higher warranty expenses, and inconsistent assessments for insurance partners.To address these challenges, this research develops and evaluates a Kalman filter-based noise reduction model executed in the cloud. Using CAN signal files collected from real vehicles testing, the model estimations the true underlying diagnostic state by filtering out random fluctuations, road disturbances, and non-critical impact signatures. The framework also incorporates a smart contract mechanism that securely timestamps validated FNOL events, ensuring trusted, tamper-resistant data exchange between OEM and insurance companies. Validation is conduction using CAN logs, simulated DTC fault scenarios, and synthetic impact profiles. Results demonstrate that the proposed model significantly reduce false-positive FNOL detection and accurately differentiates true collision events from low-impact disturbances. The integration of smart contracts provides an automated and verifiable method for sharing incident data with insurers, enabling faster claim triage, reduced fraud risk, and more precise repair cost estimation. This work offers a scalable, cloud driven methodology that enhances vehicle diagnostic reliability, improves FNOL accuracy and strengthens collaboration between OEMs and insurance companies. The framework lays the foundation for a more secure, transparent, and data-driven ecosystem in modern mobility.
  • Item type: Item ,
    Towards Physics-Based Industrial Malware Detection with Performance and Integrity Analysis Enhanced Via Machine and Deep Learning
    (2025-01-01) Al-Tarazi, Hussam; Rrushi, Julian; Cholis, Ilias; Barrak, Amine; Ming, Hua
    The security of power system protection infrastructure poses emerging challenges, particularly due to the implications of threats that exploit the physical behavior of power transformers. While prior research has made significant progress in detecting malware within industrial control systems (ICS), there remains a pressing need for more robust and adaptive methodologies capable of addressing evolving cyber threats and advanced intrusive techniques.This study emphasizes the critical role of protection relay algorithms specifically differential and harmonic restraint algorithms in safeguarding power transformers from electrical faults. These algorithms, however, are susceptible to malicious interference; they can be disabled, altered, or suppressed by malware, thereby compromising the reliability and safety of the system. We begin by presenting an overview of a typical electrical power substation architecture and the core functionalities of protection relay algorithms. Furthermore, we explore preliminary insights into malware strategies that involve physics-based data manipulation, potentially leading to transformer maloperation, large-scale blackouts, and permanent equipment damage. To address these vulnerabilities, we propose the design of a hybrid machine learning and deep learning framework aimed at detecting protection anomalies through analysis of computational resource metrics, including memory usage and execution time, associated with relay algorithm execution. The model was developed using Python and evaluated through simulation of differential and harmonic protection logic. Synthetic datasets were generated using a generative adversarial network (GAN) implemented on a virtualized environment to facilitate comprehensive model training and testing
  • Item type: Item ,
    From Core to Apex: Towards a Multi-Facet Optimization of Software Systems
    (2025-01-01) Ghammam, Anwar; Barrak, Amine; Barrak, Amine; Kessentini, Marouane; Raj, Sunny; Bagherzadeh, Mehdi; Srauy, Sam; Ming, Hua
    As software evolves and deviates from its original design, it becomes prone to increased complexity, decreased performance, and higher maintenance costs. Software optimization has thus become a critical research area in software engineering, driven by modern systems' growing complexity and scale. Refactoring - the process of restructuring code without altering its external behavior - has been extensively studied to improve software performance and maintainability. Consequently, refactoring recommendation tools have emerged to assist developers with the process of refactoring, but these tools often fail to align with industry practical needs, resulting in a gap between theoretical recommendations by researchers and practitioners' needs in real-world applications. Beyond source code, software optimization spans broader lifecycle processes, including software resource management and build systems. Our research addresses these challenges holistically, advancing software optimiation from core resource allocation to apex-level development practices. First, this dissertation presents a multi-objective approach to software management, focusing on resource-aware containerization. By optimizing workload balancing in cloud environments and extending these techniques to edge environments such as Software-Defined Vehicles (SDVs), this research ensures efficient resource utilization and scalability. Using NSGA-III, our approach demonstrates superior performance in managing constraints and optimizing conflicting objectives, with practical applications validated through industrial partnership. Second, our research proposes significant contribution to source code refactoring by addressing gaps between industry practices and existing refactoring recommendation tools. Current tools often overlook key criteria crucial to developers' decisions. By addressing the gap, we provide a broader perspective for refactoring tool developers, enabling them to enhance their tools and make them more efficient and practical for industry practitioners, thereby enhancing code maintainability and developer productivity. Finally, this research delves into the underexplored domain of build systems refactoring. Our analysis was conducted on 725 examined build-file-related commits, where we identified 24 build-related refactorings, divided into 6 main categories. We also investigate whether these refactorings are used to tackle technical debts in build systems. Building on these insights, we introduce BuildRefMiner to detect and analyze refactorings within Buildfiles using LLMs. This work redefines software optimization as an integrated process, combining refactoring insights, resource-aware strategies, and lifecycle improvements to ensure robust, efficient, and sustainable software systems
  • Item type: Item ,
    Domain-Specific LLMS-Based Refactoring Techniques for Software Vulnerabilities
    (2025-01-01) Ananbeh, Obieda; Kim, Dae-Kyoo; Lu, Lunjin; Li, Li; Ming, Hua
    Software vulnerabilities pose a significant threat to the security of systems. While much of the existing research focuses on detecting vulnerabilities, the potential of refactoring for vulnerability mitigation has not been explored much. To bridge this gap, this work introduces a set of refactoring techniques aimed at mitigating various types of vulnerabilities. The process begins by categorizing vulnerabilities based on the weakness types defined in the CWE system, concentrating on eight categories that pose significant risks. For each CWE within each category, five samples are generated using ChatGPT, resulting in a dataset of 405 samples. These samples are rigorously analyzed manually for their validity. Based on the dataset, the characteristics of each category are identified, the core problem within the category is defined, and a specific refactoring solution is developed to address it. These techniques are evaluated using the Snyk tool on twenty-one active open-source projects. The results demonstrate an 89% reduction in vulnerabilities after applying the refactoring techniques, providing insights on enhancing software security through refactoring-based strategies. Building upon this foundation, this work introduces VulnFixAI, an automated tool designed to detect and repair various types of vulnerabilities in Java code by integrating refactoring techniques with fine-tuned domain-specific large language models (DSLLMs). Specifically, VulnFixAI implements three targeted refactoring algorithms, Whitelist Validation Refactoring (WVR), Output Safety Refactoring (OSR), and TrustChain Verification Refactoring (TCVR), which were selected for their demonstrated effectiveneess against prevalent CWE vulnerabilities identified in the initial study. A dataset of 10,000 vulnerabile Java snippets was collected, extracted from 907 open-source, and fine-tuned Llama 3.2 (3B parameters) model to enhance vulnerability detection and automated repair. VulnFixAI was evaluated on 20 real-world open-source projects listed in the GitHub Advisory Database. The results demonstrate an 89% overall effectiveness in detecting and repairing vulnerabilities, representing a 51% improvement over non-fine-tuned Llama 3.2 (3B) and an average 27% improvement over other LLMs, including ChatGPT4, Claude 3.5 Sonnet, and Gemini 2.0 Flash. These findings underscore how VulnFixAI's hybrid approach, integrating fine-tuned DSLLMs and refactoring algorithms, provides an efficient, scalable, and highly effective solution for enhancing software security. This work introduces eight novel refactoring techniques specifically designed to address the root causes of vulnerabilities and presents VulnFix AI, an automated tool that integrates these techniques with DSLLMs to detect and repair vulnerabilities with high precision. Evaluated on 20 real-world projects, VulnFixAI achieved an overall effectiveness rate of 89&, significantly outperforming leading LLMs such as ChatGPT-4 74%, Claude 3.5 Sonnet 70%, Gemini 2.0 Flash 67%, and non-fine-tuned Llama 3.2 59%. These results demonstrate VulnFixAI's superior ability to identify and repair vulnerabilities through domain-specific tuning and structured refactoring
  • Item type: Item ,
    User-Centric Secure Service Provisioning with Sustainability in Resource Constrained Platforms Using Predictive Machine Learning Models
    (2025-01-01) Alao, Damilola Adedamola; Sen, Amartya; Raj, Sunny; Xu, Lanyu; Caushaj, Eralda
    This dissertation addresses the challenges of resource constraints and security vulnerabilities inherent in the increasing deployment of IoT devices and services. It embarks on a journey to optimize IoT service provisioning by recognizing the dynamic nature of user preferences and the need for robust security measures compatible with the limited computational, energy, and storage capacities of IoT devices. This foundational understanding leads to the initial contribution: a novel IoT simulation framework. This framework uniquely integrates lightweight cryptography - NtruEncrypt, within a user-centric service provisioning model, simulated using CupCarbon. The data collected from this framework validates the efficacy of lightweight security in enhancing Quality of Service (QoS) and also highlights the need for sophisticated resource management. Insights from the initial simulation inform subsequent works, driving the research towards comprehensive IoT optimization. Recognizing the urgency for improved resource management, the dissertation presents a two-phase IoT resource management optimization framework. This solution leverages predictive analytics of user service requests combined with a service recommendation methodology to dynamically reallocate IoT resources, thereby minimizing service delays and optimizing device efficiency. Focusing on sustained device operation, particularly in the context of emerging energy harvesting solutions, the work further develops an enhanced energy management framework. This framework dynamically optimizes the power distribution of harvested radio Frequency (RF) energy using a multi-objective optimization algorithm (NSGA-III) and a custom energy allocation algorithm, ensuring continuous service availability despite intermittent energy supplies. The final contribution is a secure energy scheduling framework for Energy Harvesting (EH)-enabled Sensor Cloud platforms. Directly countering energy depletion attacks by implementing dynamic, user-centric secure admission control, projecting each request’s net energy impact to block malicious requests before they can compromise the energy integrity of IoT devices. These contributions form a coherent and robust solution for dynamic, user-centric, and secure IoT service provisioning, resource management, and energy sustainability.
  • Item type: Item ,
    Detecting and refactoring technical debt for software containers
    (2024-01-01) Ksontini, Emna; Ming, Hua; Kessentini, Marouane; Wilson, Steven; Bagherzadeh, Mehdi; Srauy, Sam
    In today’s fast-paced software development environment, containerization has emerged as a cornerstone of modern infrastructure, enabling consistent and scalable deployments across varied platforms. Docker, as the leading containerization platform, plays a critical role in this landscape, but the management of Dockerfiles - scripts that automate the creation of container images - presents significant challenges. These challenges include the accumulation of technical debt, inefficient image sizes, prolonged build durations, and the emergence of anti-patterns, all of which can undermine the efficiency, maintainability, and quality of Docker-based projects. This dissertation presents and advances a suite of techniques and tools to enhance the quality of Docker projects through carefully targeted refactoring strategies, anti-pattern detection, and the mitigation of technical debt. It begins with an empirical study of open-source Docker projects, identifying 38 Docker-specific refactoring techniques and nine distinct categories of technical debt. These findings illuminate the unique challenges inherent in Dockerfile management, which differ fundamentally from those in traditional software development due to Docker's Infrastructure as Code (IaC) nature. Building on these insights, the research introduces DRMiner, a tool designed to detect and analyze refactorings within Dockerfiles. Utilizing an Enhanced Abstact Syntax Tree (E-AST) approach, DRMiner addresses the specific complexities of Docker artifacts, automating the identification of refactoring opportunities and significantly reducing the manual effort required for Dockerfile maintenance. The dissertation also proposes a novel method for the specification and detection of Docker-specific anti-patterns, expanding the scope beyond traditional code smells to address broader design flaws. This method defines five new anti-patterns and develops a metric-based framework for their automated detection. Finally, this research delves into automating Dockerfile refactoring through large language models (LLMs). The findings reveal that LLM-driven refactoring reduces image sizes and builds durations and enhances maintainability and understandability, outperforming manual refactoring methods. Together, these contributions provide a robust, empirically grounded framework for enhancing the quality and sustainability of Docker projects, offering valuable tools and insights for both practitioners and researchers in containerization
  • Item type: Item ,
    Optimizing performance of distributed memory workloads for high performance computing (hpc) and data center networks
    (2024-01-01) Newaz, Md Nahid; Qu, Guangzhi; Lu, Lunjin; Mollah, Md Atiqul; Cholis, Ilias
    Today’s high-performance computing (HPC) and large-scale data center systems are constructed by interconnecting thousands of computing nodes, storage devices, hardware accelerators, and I/O devices. Due to the rapid growth in the scale of these systems, along with their complex interconnect topologies and the need to support diverse applications, the interconnection network has become one of the major bottlenecks for performance, efficiency, and resource utilization. My research focuses on developing novel, efficient, and scalable routing methods to reduce latency and optimized scheduling methods to improve throughput and system utilization. I validated these proposed approaches through various experimental studies using state-of-the-art benchmarks, as well as real scientific and synthetic applications. The goal is to facilitate the future adoption of these approaches in next-generation HPC and data center network systems
  • Item type: Item ,
    Correlation of Cyber and Physical Dynamics in Power Grid Relays for Malware Detection
    (2025-01-01) Alali, Dafer Rzok; Rrushi, Julian; Zohdy, Mohamed A; Siadat, Mohammad R; Ming, Hua; Li, Li
    Exploits and malicious operations modules make up the malware application that targets the electrical power system. The exploits resemble those found in conventional malware for general-purpose computing. These malicious programs infiltrate an industrial computer, i.e. relay and then release functional components. In order to take over computational functions of the relay, malware run physics-aware modules that target physical equipment. An example is fabrication of fictitious status power data that indicates a power transformer is functioning normally, but in reality, the attacks are causing a malfunction in the transformers harmonic protection algorithm. This research explores the relationship between harmonic activities in power transformers and their impact on system behavior in substation computers. We privilege mimic a set of harmonic malwares and a power transformer. In this study, we contribute in multiple ways: 1) we use these emulations to examine a power transformer's cyberattack surface; 2) using these insights, we develop a number of attacks that harmonic malware could employ against a power transformer; 3) we use Python to simulate these cyberattacks in order to monitor and measure their harmful effects on a power transformer empirically. Also, we used Hypersim simulation provided by OPAR-R, to compare it with our Python simulation; 4) we use the ProcMon tool to dynamically observe system activities; and 5) we use machine learning models to forecast and assess the impact of cyberattacks on computer systems. The results demonstrated that the machine learning model achieved high accuracy in detecting malware-induced anomalies, with the Python simulation yielding a 91% accuracy and the Hypersim simulation achieving 86% accuracy in predicting cyber-physical disturbances for the type of activity "Results". Confusion matrix analyses revealed a strong correlation between harmonic distortions and malware activity, validating the effectiveness of frequency-domain analysis in anomaly detection. Furthermore, the model successfully differentiated between normal harmonic fluctuations and malicious injections, reducing false positives while maintaining high detection precision. This study highlights the critical role of machine learning in enhancing cybersecurity resilience in power grids. By integrating real-time anomaly detection and dynamic system activity monitoring, the proposed framework offers a robust defense mechanism against evolving cyber threats targeting power transformer operations. The findings provide a foundation for future research in developing AI-driven cybersecurity solutions tailored to industrial control systems
  • Item type: Item ,
    Detecting and Classifying Malware in Electrical Power Grids Via Cyberdeception
    (2024-01-01) Omar, Tallal Mohamed; Zohdy, Mohamed A; Edwards, William; Caushaj, Eralda; Sutton, Sara
    Artificial intelligence (AI) has become an essential instrument for enterprises aiming to protect their digital assets within a progressively aggressive cyber landscape. As the dependence on digital technologies increases among companies and individuals, the risks associated with cyberattacks are also advancing in terms of complexity and magnitude. AI and the proliferation of technology has led to a significant concern over security, mostly due to the escalating prevalence of malware on industrial computers. This has resulted in potential physical harm to computer systems and the individuals involved. Malware is a collection of malicious programming code that aims to inflict harm against computer systems, programs, or online apps. These applications lack the ability to differentiate between legitimate system calls and those that are intended to cause harm. Therefore, it is imperative to ensure that computer systems and online applications are constructed in a manner that enables the identification and differentiation of malicious activities from legitimate application activities. The utilization of AI in the realm of cybersecurity is revolutionizing the domain of digital protection. There are various techniques that can be used to identify malicious activity, leveraging innovative concepts such as AI, machine learning, and deep learning. The present study presents a proposal for utilizing AI approaches to identify and mitigate malware activity in computer memory, with the aim of safeguarding against unauthorized access to and manipulation of physical data within the system. This research aims to combine the traditional K-means algorithm with other methods and functionalities to perform data aggregation tasks on a physical dataset. The primary objective is to identify anomalies in the dataset using clustering techniques. These anomalies will serve as triggers for creating a replica of the main process as a decoy thread. The decoy thread will be equipped with decoy sensors and actuators. The analysis will be conducted on the decoy thread rather than the main process, allowing for intrusive observation. The same host environment will be provided to memory-resident malware, enabling it to continue operating within the main operating system process. The analysis process involves utilizing a replicated instance of malware that resides within a deceptive thread.
  • Item type: Item ,
    Resilient Suppliers Selection System Using Machine Learning Algorithms and Risk Assessment Methodology
    (2022-01-01) Albadrani, Abdullah Meteb; Zohdy, Mohamed A.; Edwards, William; Ruegg, Erica
    Demand, supply, pricing, and lead time are all unpredictable in the manufacturingindustry, and the manufacturer must function in this environment. Because of the enormous amount of data available and the introduction of new technologies, such as the internet of things (IoT), machine learning (ML), and Blockchain, administrators and government officials are better able to deal with uncertainty by applying intelligent decision-making principles to their situations. All supply chains must make use of new technology and analyze previous data to forecast and improve the success of future operations. At the moment, we rely on the supply chain and its facilities for healthcare when we receive our vaccine, for our food when we go grocery shopping, and for transportation when we drive our cars. Supplier selection is exposed to the three most significant factors: quality, delivery, and performance history, which are all evaluated separately. The use of data analytic capabilities in the selection of robust supplier portfolios has not been thoroughly investigated. Manufacturers typically have three to four resilient suppliers for the same item, but occasionally one or two of them will fail, causing a ripple effect throughout the entire supply chain. This is a frequent problem that the supply chain must deal with on a daily basis. Supply chain resilience, on the other hand, ignores or is incompatible with the risk profiles of suppliers’ performance.
  • Item type: Item ,
    Application of Several Machine Learning Algorithms for Multiple Stage Inference Data
    (2024-01-01) Amen, Khalid A; Zohdy, Mohamed A; Rrushi, Julian; McDonald, Gary; Mahmoud, Mohammed
    Historically, machine learning techniques have been dependent on utilizing data from two distinct phases to predict and identify particular occurrences. The outcomes of these studies may exhibit either validity or inaccuracy, represented by binary values of one or zero. An alternative term for this is a prognostication of one of two potential results. Several issues are present in this approach, which have the potential to yield inaccurate outcomes. The issues encompassed in this context consist of data imbalance, overfitting, and error propagation. This study aims to employ and use a multiple stage outcome approach to enhance accuracy and optimize the performance of outcomes. In this step of our research, we will be implementing the Multiclass Classification One-vs.-All methodology to analyze the data collected from various stages of the experiment's conclusion. In the subsequent phase, it is necessary to engage in the utilization or investigation of a diverse range of potential supervised models, which are trained through the application of machine learning algorithms. Subsequently, the determination of the model that exhibits a superior level of accuracy will be made by designating it as the victor. In our study, we employ and evaluate five distinct machine learning algorithms, namely Support Vector Machines (SVM), Logistic Regression (LR), Random Forest (RF), Gradient Tree Boosting (GTB), and Extremely Randomized Trees (ERF). These algorithms are used within our machine learning framework to analyze multi-stage data and ascertain the technique that exhibits the highest accuracy in predicting outcome stages. This multi-stage conclusion would effectively narrow down the problem or difficulties at hand, reduce the potential for errors, and enhance the ability to accurately predict and diagnose medical diseases or cyber security threats. A Python-based model was developed to execute the proposed methodology. The utilized notion employs a binary format, which has been substantiated by empirical evidence and offers two potential outcomes. Upon the completion of our research, it was determined that the Logistic Regression and Support Vector Machine algorithms exhibited better performance compared to the other algorithms when a multiple stage outcome was employed. The results were assessed in terms of accuracy, precision, recall, and the F measure
  • Item type: Item ,
    Knowledge Net: An Automated Clinical Knowledge Graph Generation Framework for Evidence Based Medicine
    (2023-01-01) Alam, Fakhare; Malik, Khalid Mahmood; Siadat, Mohammad-Reza; Ma, Tianle; Homayouni, Ramin; Corliss, David
    To practice the evidence-based medicine, clinicians are interested to find the most suitable research for the clinical decision making. The use of knowledge graphs (KGs) and Neuro-Symbolic methods to integrate and analyze complex and heterogeneous healthcare data is critical to enable evidence-based treatment in clinical decision support systems (CDSS). Healthcare generates a vast amount of data, including electronic health records (EHRs), medical images, genetic information, research papers, and clinical guidelines. Neuro-symbolic AI can leverage its neural network component to process unstructured data, while using symbolic reasoning to interpret the data and make logical inferences. It also enables a deeper understanding of patient data, leading to more accurate diagnoses, personalized treatment plans, and improved patient outcomes. By incorporating symbolic reasoning, Neuro-Symbolic AI systems can provide explanations for their outputs, making them more transparent and interpretable. To enable Neuro-Symbolic AI in healthcare, large-scale KGs play a pivotal role as it can integrate heterogeneous and big healthcare data including medical ontologies, clinical guidelines, drug databases, patient records, and research literature.The existing KG construction frameworks are not fully automated and predominantly carried out using manual or semi-automated approach, requiring substantial effort and expertise. The challenges encompass identifying knowledge sources, disambiguating concepts in context, enriching semantics, determining relationships, and conducting inferential reasoning. Automating the extraction of coherent knowledge and constructing KGs from diverse data forms remains a longstanding goal in AI research. Also, the current frameworks for constructing KGs fail to generate KGs that provide relevant information for evidence-based practitioners. This is because the organization of constructed subgraphs is neither topic-specific nor evidence-based PICO (Participants/Problem P, Intervention-I, Comparison C, Outcome O) query-friendly. These KGs, built through manual or semi-automated processes, are incapable of adapting to new domains and incorporating the constantly changing information into their knowledge base. Consequently, they gradually lose relevance over time and miss out on important evidence. Thus, ignoring temporal information and failing to incorporate dynamic nature of entities and relations can lead to erroneous information extraction and suboptimal decision-making. This dissertation proposes fully automated knowledge graph curation framework to curate information and create KG of different clinical domains by employing concept extraction, semantic enrichment, optimized clustering using Neuro-Symbolic approach, and state of art Recurrent Neural Networks (RNNs) with BioBERT based encoded representation to categorize PICO elements and predict relationships between concepts using huge corpus of publicly available literature on COVID-19 and cerebral aneurysms. The evaluation shows that the proposed framework achieves significant improvement over baseline models and has 93 , and 82 accuracy on aneurysm and COVID data set respectively for PICO classification. The Neuro-Symbolic clustering approach outperforms traditional baseline models by 43 and achieves average precision of 88 across all identified clusters. Also, the relationship extraction module has an accuracy of 96 with precision and recall being 92 , and 90 respectively. The incorporation of domain-specific and language models has proven to enhance the performance of machine learning models, particularly in the context of Neuro-Symbolic clustering, PICO classification, and relation extraction. The integration of deep learning and symbolic reasoning techniques has demonstrated significant improvements in clustering performance, especially in biomedical research domains. The utilization of the BioBERT embedded layer and LSTM model has notably boosted the accuracy of PICO classification tasks by 11 for both the COVID-19 dataset and cerebral aneurysm dataset. Furthermore, when BioBERT is combined with Bi-LSTM and CNN, the performance of the RE model also experiences substantial enhancements. Future work will focus on parallelizing the data processing pipeline to enhance the efficiency and scalability of the knowledge graph framework, while also developing an interactive user interface for visualization. Additionally, efforts will be dedicated to extending the frameworks application across diverse domains such as the food supply chain, dietary recommendations, agriculture, and fisheries, addressing unique challenges and expanding its impact. This expansion aims to advance multiple industries and leverage the potential benefits of the approach in various domains.