started · updated
Austrian Thesis Evaluates Datasets for Cybersecurity Small Language Models
A master thesis by Cristóbal Ricardo Jesus Veas Chávez, supervised by Jasmin Wachter at the University of Klagenfurt (May 2026), investigates how to train small language models (SLMs) for autonomous penetration testing. The work uses a collection of 585 real‑world penetration‑testing interaction logs (about 32.5 million tokens) from the CAI autonomous penetration‑testing framework.
Five dataset construction strategies are compared: a raw corpus baseline, three clustering‑based methods (feature‑vector clustering, semantic‑embedding clustering, and combined representations), and a statistical curation approach that selects logs structurally similar to successful Capture‑the‑Flag (CTF) task completions. Each dataset fine‑tunes the Qwen3‑8B model under identical conditions. Evaluation on the CyBench benchmark (33 CTF challenges) shows the statistical curation method solving five tasks, outperforming the raw corpus baseline, which solves two. The results suggest that carefully curated, high‑quality training data based on task success yields better performance than larger, unsupervised datasets for cybersecurity SLMs.
Entities
CAI autonomous penetration testing framework · Cristóbal Ricardo Jesus Veas Chávez · CyBench benchmark · Jasmin Wachter · Qwen3-8B