Toward Dynamic and Risk-Aware Evaluation of Cybersecurity LLMs: A Survey and the RIRAG Framework
DOI:
https://doi.org/10.65879/3070-5789.2026.02.05Keywords:
Large Language Models, Cybersecurity Benchmarking, Retrieval-Augmented Generation, Dynamic Benchmarking, Adversarial Robustness, Operational Risk Assessment, Human-AI Teaming, Cyber Threat IntelligenceAbstract
The rapid adoption of large language models (LLMs) in cybersecurity has created a growing need for evaluation methods that reflect operational risk rather than isolated language capability. Existing cybersecurity benchmarks assess useful dimensions such as factual knowledge, vulnerability analysis, secure coding, penetration testing, and threat intelligence reasoning, but many remain limited by static datasets, weak diagnostic granularity, limited adversarial testing, and insufficient attention to human-AI decision-making. This survey analyzes recent LLM cybersecurity benchmarks through three evaluation paradigms: knowledge-oriented, task-oriented, and holistic evaluation. From this analysis, we identify five recurring gaps between benchmark performance and real-world cybersecurity risk: knowledge-reasoning mismatch, capability-risk separation, limited failure attribution, staticity and contamination, and adversarial fragility. To address these gaps, we introduce the Reflective and Iterative Retrieval-Augmented Generation (RIRAG) framework, a dynamic and risk-aware evaluation architecture for cybersecurity LLMs. RIRAG combines continuously updated cybersecurity knowledge, retrieval- and generation-specific metrics, diagnostic logging, independent evaluation, adversarial red teaming, operational risk scoring, and human-AI teaming assessment. The framework is specified through design principles, formal components, implementation guidance, and illustrative case studies. Rather than presenting RIRAG as a validated production system, this article defines a falsifiable empirical validation protocol for future study. The central contribution is a survey-grounded reference architecture for evaluating cybersecurity LLMs as evolving, adversarially exposed, and human-interactive systems.
References
[1] Deason L, Bali A, Bejean C, Bolocan D, Crnkovich J, Croitoru I, et al. CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning. arXiv preprint arXiv:2509.20166; 2025.
[2] Microsoft Security. Microsoft Copilot for Security. 2024. Available from: https://www.microsoft.com/enus/security/business/aimachine-learning/microsoft-copilot-security.
[3] Google Cloud. Supercharge security with AI. 2024. Available from: https://cloud.google.com/security/ai.
[4] Wang D, Zhou G, Li X, Bai Y, Chen L, Qin T, et al. The Digital Cybersecurity Expert: How Far Have We Come? arXiv preprint arXiv:2504.11783; 2025.
[5] Alam MT, Bhusal D, Nguyen L, Rastogi N. CTIBench: A benchmark for evaluating LLMs in cyber threat intelligence. arXiv preprint arXiv:2406.07599; 2024.
[6] Perez F, Ribeiro I. PromptInject: A framework for quantitative analysis of robustness to adversarial prompt attacks. arXiv preprint arXiv:2211.09527; 2022.
[7] Zou A, Phute Z, Bubeck S, Ribeiro MT. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043; 2023.
[8] Liu Y, Jia Y, Jia J, Song D, Gong NZ. DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks. In: Proceedings of the 2025 IEEE Symposium on Security and Privacy (SP); 2025.
[9] Zou W, Geng R, Wang B, Jia J. PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models. In: Proceedings of the 34th USENIX Security Symposium (USENIX Security 25); 2025.
[10] Srikanth S, Hasanuzzaman M, Meem FT. Evaluating the usability of LLMs in threat intelligence enrichment. arXiv preprint arXiv:2409.15072; 2024.
[11] Liu Z, Shi J, Buford J. CyberBench: A multi-task benchmark for evaluating large language models in cybersecurity. In: Proceedings of the AAAI Conference on Artificial Intelligence; 2024.
[12] Ji H, Yang J, Chai L, Wei C, Yang L, Duan Y, et al. SevenLLM: Benchmarking, eliciting, and enhancing abilities of large language models in cyber threat intelligence. arXiv preprint arXiv:2405.03446; 2024.
[13] He M, et al. Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks. arXiv preprint arXiv:2505.12786; 2025.
[14] Fang R, Bindu R, Gupta A, Kang D. LLM Agents can Autonomously Exploit One-day Vulnerabilities. arXiv preprint arXiv:2404.08144; 2024.
[15] Fang R, Bindu R, Gupta A, Kang D. Teams of LLM Agents can Exploit Zero-Day Vulnerabilities. In: Proc. 18th Conf. European Chapter of the ACL (EACL); 2026.
[16] Deng G, Liu Y, Mayoral-Vilches V, Liu P, Li Y, Xu Y, et al. PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing. arXiv preprint arXiv:2308.06782; 2024.
[17] Wang A, Singh A, Michael J, Hill F, Levy O, Bowman SR. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In: Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP; 2018. p. 353–355. doi:10.18653/v1/W18-5446.
[18] Hendrycks D, Burns C, Basart S, Zou A, Mazeika M, Song D, et al. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300; 2020.
[19] Liang P, Bommasani R, Lee T, Tsipras D, Soylu D, Yasunaga M, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110; 2022.
[20] Ullah S, Han M, Pujar S, Pearce H, Coskun A, Stringhini G. LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks. arXiv preprint arXiv:2312.12575; 2024.
[21] Lukošiūtė K, Swanda A. LLM Cyber Evaluations Don’t Capture Real-World Risk. arXiv preprint arXiv:2502.00072; 2025.
[22] Alrashedy K, Aljasser A. Can LLMs patch security issues? arXiv preprint arXiv:2312.00024; 2023.
[23] Yu A, et al. Retrieval Augmented Generation Evaluation in the Era of Large Language Models: A Comprehensive Survey. arXiv preprint arXiv:2504.14891; 2025.
[24] Es S, James J, Espinosa-Anke L, Schockaert S. RAGAS: Automated Evaluation of Retrieval Augmented Generation. In: Proc. 18th Conf. European Chapter of the ACL: System Demonstrations; 2024. p. 150–158. doi:10.18653/v1/2024.eacl-demo.16.
[25] Chao P, Debenedetti E, Robey A, Andriushchenko M, Croce F, Sehwag V, et al. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models. In: Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track; 2024.
[26] Nasr M, Carlini N, et al. The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections. arXiv preprint arXiv:2510.09023; 2025.
[27] Cho S, et al. Towards More Robust Retrieval-Augmented Generation: Evaluating RAG Under Adversarial Poisoning Attacks. arXiv preprint arXiv:2412.16708; 2024.
[28] ElZemity A, Arief B, Li S. CyberLLMInstruct: A Pseudo-malicious Dataset Revealing Safety-performance Trade-offs in Cyber Security LLM Fine-tuning. arXiv preprint arXiv:2503.09334; 2025.
[29] Li G, Li Y, Wang G, Yang H, Yu Y. SecEval: A comprehensive benchmark for evaluating cybersecurity knowledge of foundation models. GitHub; 2023. Available from: https://github.com/XuanwuAI/SecEval.
[30] Tihanyi N, Ferrag MA, Jain R, Bisztray T, Debbah M. CyberMetric: A benchmark dataset based on retrieval-augmented generation for evaluating LLMs in cybersecurity knowledge. arXiv preprint arXiv:2402.07688; 2024.
[31] Jing P, et al. SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity. arXiv preprint arXiv:2412.20787; 2024.
[32] Yu Z, Zeng J, Chen S, Xu W, Xu D, Liu X, et al. CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity. arXiv preprint arXiv:2411.16239; 2024.
[33] Liu Y, Gao L, Yang M, Xie Y, Chen P, Zhang X, et al. VulDetectBench: Evaluating the deep capability of vulnerability detection with large language models. arXiv preprint arXiv:2406.07595; 2024.
[34] Bhatt M, Chennabasappa S, Nikolaidis C, Wan S, Evtimov I, Gabi D, et al. Purple Llama CyberSecEval: A secure coding benchmark for language models. arXiv preprint arXiv:2312.04724; 2023.
[35] Shao M, et al. NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security. arXiv preprint arXiv:2406.05590; 2024.
[36] Cheng Y, Liu Y, Li C, Song D, Gao P. CTIArena: Benchmarking LLM Knowledge and Reasoning Across Heterogeneous Cyber Threat Intelligence. arXiv preprint arXiv:2510.11974; 2025.
[37] Bhusal D, Alam MT, Nguyen L, Mahara A, Lightcap Z, Frazier R, et al. SECURE: Benchmarking Generative Large Language Models for Cybersecurity Advisory. arXiv preprint arXiv:2405.20441; 2024.
[38] NetSPI. Balancing Security and Usability of Large Language Models: An LLM Benchmarking Framework. 2024.
[39] Zhang AK, et al. Cybench: A framework for evaluating cybersecurity capabilities and risks of language models. arXiv preprint arXiv:2408.08926; 2024.
[40] Tan RS, et al. LLMs in the SOC: An Empirical Study of Human-AI Collaboration in Security Operations Centres. arXiv preprint arXiv:2508.18947; 2025.
[41] Cronbach LJ, Meehl PE. Construct Validity in Psychological Tests. Psychological Bulletin. 1955;52(4):281–302. doi:10.1037/h0040957.
[42] Messick S. Validity of Psychological Assessment: Validation of Inferences from Persons’ Responses and Performances as Scientific Inquiry into Score Meaning. American Psychologist. 1995;50(9):741–749. doi:10.1037/0003-066X.50.9.741.
[43] Campbell DT. Assessing the Impact of Planned Social Change. Evaluation and Program Planning. 1979;2(1):67–90. doi:10.1016/0149-7189(79)90048-X.
[44] Saltzer JH, Schroeder MD. The Protection of Information in Computer Systems. Proceedings of the IEEE. 1975;63(9):1278–1308. doi:10.1109/PROC.1975.9939.
[45] Joint Task Force Transformation Initiative. Guide for Conducting Risk Assessments. National Institute of Standards and Technology; 2012. Report No.: NIST SP 800-30, Revision 1. doi:10.6028/NIST.SP.800-30r1.
[46] Bansal G, Wu T, Zhou J, Fok R, et al. Does the Whole Exceed Its Parts? The Effect of AI Explanations on Complementary Team Performance. In: Proc. CHI Conference on Human Factors in Computing Systems; 2021. doi:10.1145/3411764.3445717.
[47] Buçinca Z, Malaya MB, Gajos KZ. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making. Proc. ACM on Human-Computer Interaction (CSCW). 2021;5. doi:10.1145/3449287.
[48] Samvelyan M, Raparthy SC, Lupu A, Hambro E, Markosyan AH, Bhatt M, et al. Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts. In: Advances in Neural Information Processing Systems (NeurIPS); 2024.
[49] Cheng Y, Bajaber O, Tsegai SA, Song D, Gao P. CTINexus: Automatic Cyber Threat Intelligence Knowledge Graph Construction Using Large Language Models. In: Proceedings of the 2025 IEEE European Symposium on Security and Privacy (EuroS&P); 2025.
[50] Hu Y, Zou F, Han J, Sun X, Wang Y. LLM-TIKG: Threat intelligence knowledge graph construction utilizing large language model. Computers & Security. 2024;145:103999. doi:10.1016/j.cose.2024.103999.
[51] Aguilera-Martínez F, Berzal F. LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures. arXiv preprint arXiv:2505.01177; 2025.
[52] Gong C, Li Z, Li X. Information Security Based on LLM Approaches: A Review. arXiv preprint arXiv:2507.18215; 2025.
[53] Verma A, et al. Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations. arXiv preprint arXiv:2410.09097; 2024.
[54] Deng G, et al. Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents. arXiv preprint arXiv:2602.02164; 2026.
[55] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. In: Advances in neural information processing systems; 2017. p. 5998–6008.
[56] Devlin J, Chang M, Lee K, Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In: Proc. NAACL-HLT; 2019. doi:10.18653/v1/N19-1423.
[57] Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language Models are Unsupervised Multitask Learners. OpenAI; 2019.
[58] Brown TB, Mann B, Ryder N, Subbiah M, et al. Language Models are Few-Shot Learners. In: Advances in Neural Information Processing Systems (NeurIPS); 2020.
[59] Raffel C, Shazeer N, Roberts A, Lee K, et al. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. Journal of Machine Learning Research. 2020;21.
[60] Lewis M, Liu Y, Goyal N, Ghazvininejad M, et al. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In: Proc. ACL; 2020. doi:10.18653/v1/2020.acl-main.703.
[61] Ouyang L, Wu J, Jiang X, Almeida D, et al. Training Language Models to Follow Instructions with Human Feedback. In: Advances in Neural Information Processing Systems (NeurIPS); 2022.
[62] Wei J, Bosma M, Zhao VY, Guu K, et al. Finetuned Language Models are Zero-Shot Learners. In: Proc. ICLR; 2022.
[63] Sanh V, Webson A, Raffel C, Bach SH, et al. Multitask Prompted Training Enables Zero-Shot Task Generalization. In: Proc. ICLR; 2022.
[64] Chung HW, Hou L, Longpre S, Zoph B, et al. Scaling Instruction-Finetuned Language Models. Journal of Machine Learning Research. 2024;25.
[65] Christiano PF, Leike J, Brown T, Martic M, Legg S, Amodei D. Deep Reinforcement Learning from Human Preferences. In: Advances in Neural Information Processing Systems (NeurIPS); 2017.
[66] Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal Policy Optimization Algorithms. arXiv preprint arXiv:1707.06347; 2017.
[67] Rafailov R, Sharma A, Mitchell E, Ermon S, Manning CD, Finn C. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. In: Advances in Neural Information Processing Systems (NeurIPS); 2023.
[68] Bai Y, Kadavath S, Kundu S, et al. Constitutional AI: Harmlessness from AI Feedback. arXiv preprint arXiv:2212.08073; 2022.
[69] Touvron H, Lavril T, Izacard G, et al. Llama: Open and Efficient Foundation Language Models. arXiv preprint arXiv:2302.13971; 2023.
[70] Touvron H, Martin L, Stone K, et al. Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv preprint arXiv:2307.09288; 2023.
[71] Chowdhery A, Narang S, Devlin J, et al. PaLM: Scaling Language Modeling with Pathways. Journal of Machine Learning Research. 2023;24.
[72] Hu EJ, Shen Y, Wallis P, Allen-Zhu Z, et al. LoRA: Low-Rank Adaptation of Large Language Models. In: Proc. ICLR; 2022.
[73] Dettmers T, Pagnoni A, Holtzman A, Zettlemoyer L. QLoRA: Efficient Finetuning of Quantized LLMs. In: Advances in Neural Information Processing Systems (NeurIPS); 2023.
[74] Kaplan J, McCandlish S, Henighan T, Brown TB, et al. Scaling Laws for Neural Language Models. arXiv preprint arXiv:2001.08361; 2020.
[75] Hoffmann J, Borgeaud S, Mensch A, et al. Training Compute-Optimal Large Language Models. In: Advances in Neural Information Processing Systems (NeurIPS); 2022.
[76] Wei J, Tay Y, Bommasani R, Raffel C, et al. Emergent Abilities of Large Language Models. Transactions on Machine Learning Research. 2022.
[77] Bubeck S, Chandrasekaran V, Eldan R, et al. Sparks of Artificial General Intelligence: Early Experiments with GPT-4. arXiv preprint arXiv:2303.12712; 2023.
[78] Bommasani R, Hudson DA, Adeli E, Altman R, et al. On the Opportunities and Risks of Foundation Models. arXiv preprint arXiv:2108.07258; 2021.
[79] OpenAI. GPT-4 Technical Report. arXiv preprint arXiv:2303.08774; 2023.
[80] Wang A, Pruksachatkun Y, Nangia N, Singh A, et al. SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems. In: Advances in Neural Information Processing Systems (NeurIPS); 2019.
[81] Srivastava A, Rastogi A, Rao A, et al. Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models. Transactions on Machine Learning Research. 2023.
[82] Zhong W, Cui R, Guo Y, Liang Y, et al. AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models. arXiv preprint arXiv:2304.06364; 2023.
[83] Wang Y, Ma X, Zhang G, Ni Y, et al. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. Advances in Neural Information Processing Systems (NeurIPS). 2024.
[84] Wei J, Wang X, Schuurmans D, Bosma M, et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In: Advances in Neural Information Processing Systems (NeurIPS); 2022.
[85] Cobbe K, Kosaraju V, Bavarian M, et al. Training Verifiers to Solve Math Word Problems. arXiv preprint arXiv:2110.14168; 2021.
[86] Hendrycks D, Burns C, Kadavath S, Arora A, et al. Measuring Mathematical Problem Solving with the MATH Dataset. In: Proceedings of the NeurIPS Datasets and Benchmarks Track; 2021.
[87] Chen M, Tworek J, Jun H, Yuan Q, et al. Evaluating Large Language Models Trained on Code. arXiv preprint arXiv:2107.03374; 2021.
[88] Austin J, Odena A, Nye M, Bosma M, et al. Program Synthesis with Large Language Models. arXiv preprint arXiv:2108.07732; 2021.
[89] Lin S, Hilton J, Evans O. TruthfulQA: Measuring How Models Mimic Human Falsehoods. In: Proc. ACL; 2022. doi:10.18653/v1/2022.acl-long.229.
[90] Zellers R, Holtzman A, Bisk Y, Farhadi A, Choi Y. HellaSwag: Can a Machine Really Finish Your Sentence? In: Proc. ACL; 2019. doi:10.18653/v1/P19-1472.
[91] Guha N, Nyarko J, Ho DE, Ré C, et al. LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models. Advances in Neural Information Processing Systems (NeurIPS). 2023.
[92] Jin Q, Dhingra B, Liu Z, Cohen WW, Lu X. PubMedQA: A Dataset for Biomedical Research Question Answering. Proc. EMNLP-IJCNLP. 2019. doi:10.18653/v1/D19-1259.
[93] Zheng L, Chiang W, Sheng Y, Zhuang S, et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In: Advances in Neural Information Processing Systems (NeurIPS); 2023.
[94] Min S, Krishna K, Lyu X, Lewis M, et al. FActScore: Fine-Grained Atomic Evaluation of Factual Precision in Long Form Text Generation. In: Proc. EMNLP; 2023.
[95] Manakul P, Liusie A, Gales MJ. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. In: Proc. EMNLP; 2023.
[96] Lewis P, Perez E, Piktus A, Petroni F, et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In: Advances in Neural Information Processing Systems (NeurIPS); 2020.
[97] Guu K, Lee K, Tung Z, Pasupat P, Chang M. REALM: Retrieval-Augmented Language Model Pre-Training. In: Proc. ICML; 2020.
[98] Borgeaud S, Mensch A, Hoffmann J, Cai T, et al. Improving Language Models by Retrieving from Trillions of Tokens. In: Proc. ICML; 2022.
[99] Izacard G, Lewis P, Lomeli M, Hosseini L, et al. Atlas: Few-shot Learning with Retrieval Augmented Language Models. Journal of Machine Learning Research. 2023;24.
[100] Karpukhin V, Oguz B, Min S, Lewis P, et al. Dense Passage Retrieval for Open-Domain Question Answering. In: Proc. EMNLP; 2020. doi:10.18653/v1/2020.emnlp-main.550.
[101] Khattab O, Zaharia M. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In: Proc. ACM SIGIR; 2020. doi:10.1145/3397271.3401075.
[102] Johnson J, Douze M, Jégou H. Billion-Scale Similarity Search with GPUs. IEEE Transactions on Big Data. 2021;7(3). doi:10.1109/TBDATA.2019.2921572.
[103] Malkov YA, Yashunin DA. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2020;42(4). doi:10.1109/TPAMI.2018.2889473.
[104] Asai A, Wu Z, Wang Y, Sil A, Hajishirzi H. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. In: Proc. ICLR; 2024.
[105] Shi W, Min S, Yasunaga M, Seo M, et al. REPLUG: Retrieval-Augmented Black-Box Language Models. In: Proc. NAACL; 2024.
[106] Gao Y, Xiong Y, Gao X, Jia K, et al. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv preprint arXiv:2312.10997; 2023.
[107] Zhong Z, Huang Z, Wettig A, Chen D. Poisoning Retrieval Corpora by Injecting Adversarial Passages. Proc. EMNLP. 2023.
[108] Chaudhari H, Severi G, Abascal J, Jagielski M, et al. Phantom: General Trigger Attacks on Retrieval Augmented Language Generation. arXiv preprint arXiv:2405.20485; 2024.
[109] Xiang C, Wu T, Zhong Z, Wagner D, Chen D, Mittal P. Certifiably Robust RAG Against Retrieval Corruption. arXiv preprint arXiv:2405.15556; 2024.
[110] Yao S, Zhao J, Yu D, Du N, et al. ReAct: Synergizing Reasoning and Acting in Language Models. In: Proc. ICLR; 2023.
[111] Schick T, Dwivedi-Yu J, Dessì R, Raileanu R, et al. Toolformer: Language Models Can Teach Themselves to Use Tools. In: Advances in Neural Information Processing Systems (NeurIPS); 2023.
[112] Shinn N, Cassano F, Gopinath A, Narasimhan K, Yao S. Reflexion: Language Agents with Verbal Reinforcement Learning. In: Advances in Neural Information Processing Systems (NeurIPS); 2023.
[113] Park JS, O’Brien J, Cai CJ, Morris MR, et al. Generative Agents: Interactive Simulacra of Human Behavior. In: Proc. ACM UIST; 2023. doi:10.1145/3586183.3606763.
[114] Qin Y, Liang S, Ye Y, Zhu K, et al. ToolLLM: Facilitating Large Language Models to Master 16000+ Real-World APIs. In: Proc. ICLR; 2024.
[115] Wu Q, Bansal G, Zhang J, Wu Y, et al. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. arXiv preprint arXiv:2308.08155; 2023.
[116] Hong S, Zhuge M, Chen J, Zheng X, et al. MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework. In: Proc. ICLR; 2024.
[117] Xi Z, Chen W, Guo X, He W, et al. The Rise and Potential of Large Language Model Based Agents: A Survey. arXiv preprint arXiv:2309.07864; 2023.
[118] Liu X, Yu H, Zhang H, Xu Y, et al. AgentBench: Evaluating LLMs as Agents. In: Proc. ICLR; 2024.
[119] Fang R, Bindu R, Gupta A, Zhan Q, Kang D. LLM Agents can Autonomously Hack Websites. arXiv preprint arXiv:2402.06664; 2024.
[120] Xu J, Stokes JW, McDonald G, Bai X, et al. AutoAttacker: A Large Language Model Guided System to Implement Automatic Cyber-attacks. arXiv preprint arXiv:2403.01038; 2024.
[121] Happe A, Cito J. Getting Pwn’d by AI: Penetration Testing with Large Language Models. Proc. 31st ACM Joint European Software Engineering Conference and Symposium (ESEC/FSE). 2023. doi:10.1145/3611643.3613083.
[122] Zhan Q, Liang Z, Ying Z, Kang D. InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents. Findings of ACL. 2024.
[123] Debenedetti E, Zhang J, Balunovic M, Beurer-Kellner L, et al. AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. In: Proceedings of the NeurIPS Datasets and Benchmarks Track; 2024.
[124] Deng G, et al. A Survey on Agentic Security: Applications, Threats and Defenses. arXiv preprint arXiv:2510.06445; 2025.
[125] Wei A, Haghtalab N, Steinhardt J. Jailbroken: How Does LLM Safety Training Fail? In: Advances in Neural Information Processing Systems (NeurIPS); 2023.
[126] Chao P, Robey A, Dobriban E, Hassani H, Pappas GJ, Wong E. Jailbreaking Black Box Large Language Models in Twenty Queries. In: Proceedings of the IEEE Conference on Secure and Trustworthy Machine Learning (SaTML); 2025.
[127] Mehrotra A, Zampetakis M, Kassianik P, Nelson B, et al. Tree of Attacks: Jailbreaking Black-Box LLMs Automatically. In: Advances in Neural Information Processing Systems (NeurIPS); 2024.
[128] Liu X, Xu N, Chen M, Xiao C. AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. In: Proc. ICLR; 2024.
[129] Shen X, Chen Z, Backes M, Shen Y, Zhang Y. “Do Anything Now”: Characterizing and Evaluating In-the-Wild Jailbreak Prompts on Large Language Models. In: Proc. ACM CCS; 2024.
[130] Andriushchenko M, Croce F, Flammarion N. Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks. Proc. ICLR. 2025.
[131] Zeng Y, Lin H, Zhang J, Yang D, Jia R, Shi W. How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs. Proc. ACL. 2024.
[132] Ren Q, Li H, Liu D, Xie Z, et al. Derail Yourself: Multi-turn LLM Jailbreak Attack through Self-discovered Clues. arXiv preprint arXiv:2410.10700; 2024.
[133] Schulhoff S, Pinto J, Khan A, Bouchard L, et al. Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking Competition. Proc. EMNLP. 2023.
[134] Greshake K, Abdelnabi S, Mishra S, Endres C, Holz T, Fritz M. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In: Proc. 16th ACM Workshop on Artificial Intelligence and Security (AISec); 2023.
[135] Abdelnabi S, Greshake K, Mishra S, et al. More than You’ve Asked For: A Comprehensive Analysis of Novel Prompt Injection Threats to Application-Integrated Large Language Models. arXiv preprint arXiv:2302.12173; 2023.
[136] Liu Y, Jia Y, Geng R, Jia J, Gong NZ. Formalizing and Benchmarking Prompt Injection Attacks and Defenses. In: Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24); 2024.
[137] Yi J, Xie Y, Zhu B, Hines K, et al. Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models. arXiv preprint arXiv:2312.14197; 2024.
[138] Carlini N, Tramer F, Wallace E, Jagielski M, et al. Extracting Training Data from Large Language Models. In: Proceedings of the 30th USENIX Security Symposium (USENIX Security 21); 2021.
[139] Carlini N, Ippolito D, Jagielski M, Lee K, Tramer F, Zhang C. Quantifying Memorization Across Neural Language Models. In: Proc. ICLR; 2023.
[140] Wallace E, Feng S, Kandpal N, Gardner M, Singh S. Universal Adversarial Triggers for Attacking and Analyzing NLP. In: Proc. EMNLP-IJCNLP; 2019. doi:10.18653/v1/D19-1221.
[141] Ganguli D, Lovitt L, Kernion J, Askell A, et al. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv preprint arXiv:2209.07858; 2022.
[142] Perez E, Huang S, Song F, Cai T, et al. Red Teaming Language Models with Language Models. In: Proc. EMNLP; 2022. doi: 10.18653/v1/2022.emnlp-main.225.
[143] Yu J, Lin X, Yu Z, Xing X. GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts. arXiv preprint arXiv:2309.10253; 2023.
[144] Lin L, Mu H, Zhai Z, Wang M, et al. Large Language Model Red Teaming: A Survey of Attacks, Defenses, and Evaluation. arXiv preprint arXiv:2407.14937; 2024.
[145] Mazeika M, Phan L, Yin X, Zou A, et al. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal. Proc. ICML. 2024.
[146] Xu Z, Liu Y, Deng G, Li Y, Picek S. A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models. Findings of ACL. 2024.
[147] Jain N, Schwarzschild A, Wen Y, Somepalli G, et al. Baseline Defenses for Adversarial Attacks Against Aligned Language Models. arXiv preprint arXiv:2309.00614; 2023.
[148] Robey A, Wong E, Hassani H, Pappas GJ. SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks. arXiv preprint arXiv:2310.03684; 2023.
[149] Kumar A, Agarwal C, Srinivas S, Feizi S, Lakkaraju H. Certifying LLM Safety Against Adversarial Prompting. arXiv preprint arXiv:2309.02705; 2023.
[150] Inan H, Upasani K, Chi J, Rungta R, et al. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. arXiv preprint arXiv:2312.06674; 2023.
[151] Weidinger L, Mellor J, Rauh M, Griffin C, et al. Ethical and Social Risks of Harm from Language Models. arXiv preprint arXiv:2112.04359; 2021.
[152] Ji Z, Lee N, Frieske R, Yu T, et al. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys. 2023;55(12). doi:10.1145/3571730.
[153] Huang L, Yu W, Ma W, Zhong W, et al. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Transactions on Information Systems. 2025;43(2).
[154] Sommer R, Paxson V. Outside the Closed World: On Using Machine Learning for Network Intrusion Detection. In: Proceedings of the IEEE Symposium on Security and Privacy; 2010. doi:10.1109/SP.2010.25.
[155] Arp D, Quiring E, Pendlebury F, Warnecke A, et al. Dos and Don’ts of Machine Learning in Computer Security. In: Proceedings of the 31st USENIX Security Symposium (USENIX Security 22); 2022.
[156] Apruzzese G, Laskov P, Oca EM, Mallouli W, et al. The Role of Machine Learning in Cybersecurity. In: Proceedings of the Digital Threats: Research and Practice; 2023.
[157] Tavallaee M, Bagheri E, Lu W, Ghorbani AA. A Detailed Analysis of the KDD CUP 99 Data Set. In: Proceedings of the IEEE Symposium on Computational Intelligence for Security and Defense Applications (CISDA); 2009. doi:10.1109/CISDA.2009.5356528.
[158] Moustafa N, Slay J. UNSW-NB15: A Comprehensive Data Set for Network Intrusion Detection Systems. In: Proceedings of the Military Communications and Information Systems Conference (MilCIS); 2015. doi:10.1109/MilCIS.2015.7348942.
[159] Sharafaldin I, Lashkari AH, Ghorbani AA. Toward Generating a New Intrusion Detection Dataset and Intrusion Traffic Characterization. In: Proc. ICISSP; 2018. doi:10.5220/0006639801080116.
[160] Anderson HS, Roth P. EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models. arXiv preprint arXiv:1804.04637; 2018.
[161] Raff E, Barker J, Sylvester J, Brandon R, Catanzaro B, Nicholas CK. Malware Detection by Eating a Whole EXE. In: Proceedings of the AAAI Workshop on Artificial Intelligence for Cyber Security; 2018.
[162] Downing E, Mirsky Y, Park K, Lee W. DeepReflect: Discovering Malicious Functionality through Binary Reconstruction. In: Proceedings of the 30th USENIX Security Symposium (USENIX Security 21); 2021.
[163] Buczak AL, Guven E. A Survey of Data Mining and Machine Learning Methods for Cyber Security Intrusion Detection. IEEE Communications Surveys & Tutorials. 2016;18(2). doi:10.1109/COMST.2015.2494502.
[164] Ucci D, Aniello L, Baldoni R. Survey of Machine Learning Techniques for Malware Analysis. Computers & Security. 2019;81. doi: 10.1016/j.cose.2018.11.001.
[165] Quiring E, Klein D, Arp D, Johns M, Rieck K. Adversarial Preprocessing: Understanding and Preventing Image-Scaling Attacks in Machine Learning. In: Proceedings of the 29th USENIX Security Symposium (USENIX Security 20); 2020.
[166] Sahingoz OK, Buber E, Demir O, Diri B. Machine Learning Based Phishing Detection from URLs. Expert Systems with Applications. 2019;117. doi: 10.1016/j.eswa.2018.09.029.
[167] Li Z, Zou D, Xu S, Ou X, Jin H, et al. VulDeePecker: A Deep Learning-Based System for Vulnerability Detection. In: Proc. Network and Distributed System Security Symposium (NDSS); 2018.
[168] Zhou Y, Liu S, Siow J, Du X, Liu Y. Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks. In: Advances in Neural Information Processing Systems (NeurIPS); 2019.
[169] Chakraborty S, Krishna R, Ding Y, Ray B. Deep Learning Based Vulnerability Detection: Are We There Yet? IEEE Transactions on Software Engineering. 2022;48(9). doi:10.1109/TSE.2021.3078750.
[170] Fu M, Tantithamthavorn C. LineVul: A Transformer-Based Line-Level Vulnerability Prediction. In: Proc. 19th International Conference on Mining Software Repositories (MSR); 2022. doi:10.1145/3524842.3528452.
[171] Fan J, Li Y, Wang S, Wang G, Chen X. BigVul: A large-scale, fine-grained vulnerability dataset. arXiv preprint arXiv:2010.08772; 2020.
[172] Feng Z, Guo D, Tang D, Duan N, et al. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In: Findings of EMNLP; 2020. doi: 10.18653/v1/2020.findings-emnlp.139.
[173] Guo D, Ren S, Lu S, Feng Z, et al. GraphCodeBERT: Pre-training Code Representations with Data Flow. In: Proc. ICLR; 2021.
[174] Wang Y, Wang W, Joty S, Hoi SC. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In: Proc. EMNLP; 2021. doi: 10.18653/v1/2021.emnlp-main.685.
[175] Nijkamp E, Pang B, Hayashi H, Tu L, et al. CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. In: Proc. ICLR; 2023.
[176] Li R, Allal LB, Zi Y, Muennighoff N, et al. StarCoder: May the Source Be with You! Transactions on Machine Learning Research. 2023.
[177] Rozière B, Gehring J, Gloeckle F, Sootla S, et al. Code Llama: Open Foundation Models for Code. arXiv preprint arXiv:2308.12950; 2023.
[178] Pearce H, Ahmad B, Tan B, Dolan-Gavitt B, Karri R. Asleep at the Keyboard? Assessing the Security of GitHub Copilot’s Code Contributions. In: Proceedings of the IEEE Symposium on Security and Privacy; 2022.
[179] Pearce H, Tan B, Ahmad B, Karri R, Dolan-Gavitt B. Examining Zero-Shot Vulnerability Repair with Large Language Models. In: Proceedings of the IEEE Symposium on Security and Privacy; 2023.
[180] Sandoval G, Pearce H, Nys T, Karri R, Garg S, Dolan-Gavitt B. Lost at C: A User Study on the Security Implications of Large Language Model Code Assistants. In: Proceedings of the 32nd USENIX Security Symposium (USENIX Security 23); 2023.
[181] Tihanyi N, Bisztray T, Jain R, Ferrag MA, et al. The FormAI Dataset: Generative AI in Software Security through the Lens of Formal Verification. Proc. 19th International Conference on Predictive Models and Data Analytics in Software Engineering. 2023. doi:10.1145/3617555.3617874.
[182] Khare A, Dutta S, Li Z, Solko-Breslin A, et al. Can ChatGPT Pass the Vulnerability Detection Test? An Empirical Study. arXiv preprint arXiv:2311.16169; 2023.
[183] Husari G, Al-Shaer E, Ahmed M, Chu B, Niu X. TTPDrill: Automatic and Accurate Extraction of Threat Actions from Unstructured Text of CTI Sources. In: Proc. 33rd Annual Computer Security Applications Conference (ACSAC); 2017.
[184] Liao X, Yuan K, Wang X, Li Z, Xing L, Beyah R. Acing the IOC Game: Toward Automatic Discovery and Analysis of Open-Source Cyber Threat Intelligence. In: Proc. ACM CCS; 2016. doi:10.1145/2976749.2978315.
[185] Li Z, Zeng J, Chen Y, Liang Z. AttacKG: Constructing Technique Knowledge Graph from Cyber Threat Intelligence Reports. In: Proc. European Symposium on Research in Computer Security (ESORICS); 2022.
[186] Gao P, Shao F, Liu X, Xiao X, et al. Enabling Efficient Cyber Threat Hunting with Cyber Threat Intelligence. In: Proc. IEEE 37th International Conference on Data Engineering (ICDE); 2021.
[187] Milajerdi SM, Gjomemo R, Eshete B, Sekar R, Venkatakrishnan V. HOLMES: Real-Time APT Detection through Correlation of Suspicious Information Flows. In: Proceedings of the IEEE Symposium on Security and Privacy; 2019.
[188] Milajerdi SM, Eshete B, Gjomemo R, Venkatakrishnan V. POIROT: Aligning Attack Behavior with Kernel Audit Records for Cyber Threat Hunting. In: Proc. ACM CCS; 2019.
[189] Strom BE, Applebaum A, Miller DP, Nickels KC, et al. MITRE ATT&CK: Design and Philosophy. The MITRE Corporation; 2018.
[190] Lai V, Chen C, Liao QV, Smith-Renner A, Tan C. Towards a Science of Human-AI Decision Making: A Survey of Empirical Studies. arXiv preprint arXiv:2112.11471; 2021.
[191] Alahmadi BA, Axon L, Martinovic I. 99% False Positives: A Qualitative Study of SOC Analysts’ Perspectives on Security Alarms. In: Proceedings of the 31st USENIX Security Symposium (USENIX Security 22); 2022.
[192] Hassan WU, Guo S, Li D, Chen Z, et al. NoDoze: Combatting Threat Alert Fatigue with Automated Provenance Triage. In: Proc. Network and Distributed System Security Symposium (NDSS); 2019.
[193] Vielberth M, Bohm F, Fichtinger I, Pernul G. Security Operations Center: A Systematic Study and Open Challenges. IEEE Access. 2020;8. doi:10.1109/ACCESS.2020.3045514.
[194] Yao Y, Duan J, Xu K, Cai Y, Sun Z, Zhang Y. A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the Ugly. High-Confidence Computing. 2024;4(2).
[195] Motlagh FN, Hajizadeh M, Majd M, Najafi P, et al. Large Language Models in Cybersecurity: State-of-the-Art. arXiv preprint arXiv:2402.00891; 2024.
[196] Xu H, Wang S, Li N, Wang K, et al. Large Language Models for Cyber Security: A Systematic Literature Review. arXiv preprint arXiv:2405.04760; 2024.
[197] Hassanin M, Moustafa N. A Comprehensive Overview of Large Language Models (LLMs) for Cyber Defences: Opportunities and Directions. arXiv preprint arXiv:2405.14487; 2024.
[198] Ferrag MA, Alwahedi F, Battah A, Cherif B, et al. Generative AI and Large Language Models for Cyber Security: All Insights You Need. arXiv preprint arXiv:2405.12750; 2024.
[199] Zhang J, Bu H, Wen H, Liu Y, et al. When LLMs Meet Cybersecurity: A Systematic Literature Review. Cybersecurity. 2025;8.
[200] Das BC, Amini MH, Wu Y. A Survey on Large Language Model (LLM) Security: Threats, Vulnerabilities, and Defenses. ACM Computing Surveys. 2025.
[201] Happe A, Cito J. Benchmarking Practices in LLM-driven Offensive Security: Testbeds, Metrics, and Experiment Design. arXiv preprint arXiv:2504.10112; 2025.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 RAVI PRASAD, Feroz Ahmed, Md Shohel Rana, Charan Gudla, Sujan Kumar Reddy Challa, Aditya Garg (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.