A Survey on Malware Detection with Graph Representation Learning

Обзор методов обнаружения вредоносного программного обеспечения с использованием обучения представлений на графах
Nour El Madhoun, Tristan Bilot, Khaldoun Al Agha, Anis Zouaoui
2024-05-21

control flow graphsfunction call graphsgraph neural networksgraph representation learningmalware detection
Malware detection has become a major concern due to the increasing number and complexity of malware. Traditional detection methods based on signatures and heuristics are used for malware detection, but unfortunately, they suffer from poor generalization to unknown attacks and can be easily circumvented using obfuscation techniques. In recent years, Machine Learning (ML) and notably Deep Learning (DL) achieved impressive results in malware detection by learning useful representations from data and have become a solution preferred over traditional methods. Recently, the application of Graph Representation Learning (GRL) techniques on graph-structured data has demonstrated impressive capabilities in malware detection. This success benefits notably from the robust structure of graphs, which are challenging for attackers to alter, and their intrinsic explainability capabilities. In this survey, we provide an in-depth literature review to summarize and unify existing works under the common approaches and architectures. We notably demonstrate that Graph Neural Networks (GNNs) reach competitive results in learning robust embeddings from malware represented as expressive graph structures such as Function Call Graphs (FCGs) and Control Flow Graphs (CFGs). This study also discusses the robustness of GRL-based methods to adversarial attacks, contrasts their effectiveness with other ML/DL approaches, and outlines future research for practical deployment.
1
Graph Neural Networks achieve competitive performance in learning robust malware embeddings from expressive Function Call Graphs and Control Flow Graphs.
2
Graph Representation Learning has shown strong potential for malware detection by exploiting robust graph structures that are difficult for attackers to modify.
3
Graph-based malware detection offers intrinsic explainability and is reviewed alongside robustness to adversarial attacks and comparisons with other ML/DL methods.
4
The survey identifies future research directions needed to support practical deployment of Graph Representation Learning for malware detection.
5
Traditional signature- and heuristic-based malware detection methods generalize poorly to unknown attacks and are vulnerable to obfuscation.

Malware represented as expressive graph structures, particularly Function Call Graphs (FCGs) and Control Flow Graphs (CFGs)

Graph Representation Learning (GRL) approaches for malware detection, including robust embedding performance, adversarial robustness, explainability, and comparative effectiveness against other ML/DL methods

Publication Details
Publication Date
2024-05-21
Journal
Publisher
ISSN
Cited by
81
Access Type
Author Information
Authors
Nour El Madhoun
Tristan Bilot
Khaldoun Al Agha
Anis Zouaoui
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%