R+R: Security Vulnerability Dataset Quality Is Critical

Anurag Swarnim Yadav
Joseph Wilson

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Large Language Models (LLMs) are of great interest in vulnerability detection and repair. The effectiveness of these models hinges on the quality of the datasets used for both training and evaluation. Our investigation reveals that a number of studies featured in prominent software engineering conferences have employed datasets that are plagued by high duplication rates, questionable label accuracy, and incomplete samples. Using these datasets for experimentation will yield incorrect results that are significantly different from actual expected behavior. For example, the state-of-the-art VulRepair Model, which is reported to have 44% accuracy, on average yielded 9% accuracy when test-set duplicates were removed from its training set and 13% accuracy when training-set duplicates were removed from its test set. In an effort to tackle these data quality concerns, we have retrained models from several papers without duplicates and conducted an accuracy assessment of labels for the top ten most hazardous Common Weakness Enumerations (CWEs). Our findings indicate that 56% of the samples had incorrect labels and 44% comprised incomplete samples—only 31% were both accurate and complete. Finally, we employ transfer learning using a large deduplicated bug-fix corpus to show that these models can exhibit better performance if given larger amounts of high-quality pre-training data, leading us to conclude that while previous studies have over-estimated performance due to poor dataset quality, this does not demonstrate that better performance is not possible.

Version published to 10.32388/bt8ocz
Mar 26, 2025

Multi-Sallm: A Multilingual Security Assessment of Generated Code

This article has 5 authors:
1. Mohammed Latif Siddiq
2. Noshin Ulfat
3. Nishat Raihan
4. Joanna C. S. Santos
5. Marcos Zampieri
This article has no evaluationsLatest version Dec 16, 2025
Explainable and Adversarial Robust Deep Learning for Malware Campaigns Forensic Attribution

This article has 4 authors:
1. Idowu Olugbenga ADEWUMI
2. Wumi AJAYI
3. Tolulope OLUFEMI
4. Ayoade Oluwafisayo BABATOPE
This article has no evaluationsLatest version Jan 7, 2026
SMELL AWARE BUG CLASSIFICATION

This article has 1 author:
1. Khyber Zaland
This article has no evaluationsLatest version Jan 19, 2026

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Multi-Sallm: A Multilingual Security Assessment of Generated Code

Explainable and Adversarial Robust Deep Learning for Malware Campaigns Forensic Attribution

SMELL AWARE BUG CLASSIFICATION