ASEET Talk

PatchDB After Five Years: Experiences and Lessons Learned in Building a Dataset for AI-Enabled Software Vulnerability Research and Education


Abstract


Large-scale, high-quality datasets play a critical role in advancing AI-enabled software vulnerability research and education. In this talk, I revisit PatchDB, a sizable software vulnerability dataset that integrates NVD, wild, and synthetic security vulnerabilities and patches, as well as reflect on five years of experience in its development, use, maintenance, and evolution.

I will discuss key design decisions behind PatchDB, including strategies for scalable data collection, automated patch identification, and the use of synthesis techniques to address data scarcity and imbalance. Beyond the technical pipeline, I will highlight practical challenges in ensuring data quality, diversity, and reproducibility, as well as lessons learned in maintaining and extending the dataset over time. In particular, this talk will explore how PatchDB has been leveraged not only to advance AI-based vulnerability detection research, but also as a pedagogical resource in software engineering and cybersecurity education. I will share insights on integrating real-world security datasets into coursework and supporting hands-on learning.

Speaker


Xinda Wang's avatar
Professor Xinda Wang USA

Department of Computer Science

University of Texas at Dallas


Xinda Wang is an Assistant Professor of Computer Science at the University of Texas at Dallas. Her research focuses on software and systems security, with an emphasis on applying AI to vulnerability detection, software patch management, and program analysis. She is also interested in identifying attacks and developing defense techniques for AI-based security systems. Her work has been published in premier security and software engineering venues (e.g., IEEE S&P, USENIX Security, ACSAC, DSN, ICSME, AsiaCCS, TDSC, TIFS) and recognized by 2026 NSF CAREER Award. She led the development of PatchDB, a widely used dataset for software vulnerability and security patch analysis, which has been downloaded by over 1,000 external researchers, selected by Google DeepMind to evaluate the Gemini 1.5 LLM, and chosen by the NSF as 1 of only 10 datasets integrated into the NAIRR Pilot to advance AI literacy, education, and innovation.