Can Open Large Language Models Catch Vulnerabilities?

Способны ли открытые большие языковые модели выявлять уязвимости?
DeepSeek-AI
2025-01-01

Big-Vul datasetCommon Weakness EnumerationLarge language modelsVulnerability classificationVulnerability detection
As Large Language Models (LLMs) become increasingly integrated into secure software development workflows, a critical question remains unanswered: can these models not only detect insecure code but also reliably classify vulnerabilities according to standardized taxonomies? In this work, we conduct a systematic evaluation of three state-of-the-art LLMs - Llama3, Codestral, and Deepseek R1 - using a carefully filtered subset of the Big-Vul dataset annotated with eight representative Common Weakness Enumeration categories. Adopting a closed-world classification setup, we assess each model’s performance in both identifying the presence of vulnerabilities and mapping them to the correct CWE label. Our findings reveal a sharp contrast between high detection rates and markedly poor classification accuracy, with frequent overgeneralization and misclassification. Moreover, we analyze model-specific biases and common failure modes, shedding light on the limitations of current LLMs in performing fine-grained security reasoning.These insights are especially relevant in educational contexts, where LLMs are being adopted as learning aids despite their limitations. A nuanced understanding of their behaviour is essential to prevent the propagation of misconceptions among students. Our results expose key challenges that must be addressed before LLMs can be reliably deployed in security-sensitive environments.
1
A systematic evaluation tested Llama3, Codestral, and Deepseek R1 on vulnerability detection and CWE classification using a filtered Big-Vul subset.
2
Frequent overgeneralization and misclassification demonstrate that current LLMs struggle with fine-grained vulnerability reasoning.
3
The analysis identifies model-specific biases and recurring failure modes that limit reliable security-sensitive deployment.
4
The models achieved high rates of detecting whether code was vulnerable but markedly poor accuracy when assigning the correct CWE category.
5
These limitations are particularly important in education, where incorrect LLM security guidance could propagate misconceptions among students.

Three open Large Language Models (Llama3, Codestral, Deepseek R1) evaluated on a filtered Big-Vul dataset annotated with eight CWE categories

The models’ ability and limitations in detecting vulnerabilities and classifying them into standardized CWE categories, including biases and failure modes

Publication Details
Publication Date
2025-01-01
Journal
Publisher
ISSN
Cited by
532
Access Type
Author Information
Authors
DeepSeek-AI
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%