A novel detection mechanism for sensor spoofing attacks using spatio-temporal consistency analysis.
Oct 14, 2024

An LLM-empowered automated penetration testing framework that leverages domain knowledge inherent in LLMs, achieving 228.6% task completion improvement over baseline GPT models.
Aug 14, 2024
A comprehensive taxonomy and effective detection methods for glitch tokens in Large Language Models.
Jul 15, 2024
A comprehensive guide to jailbreaking ChatGPT via prompt engineering techniques.
Apr 20, 2024

A comprehensive framework for automated jailbreaking of Large Language Model chatbots, featuring novel attack methodologies and systematic analysis of defense mechanisms.
Feb 26, 2024
MASTERKEY is a research framework for automated jailbreak attack generation and defense evaluation for LLM chatbots. The framework supports systematic analysis of jailbreak strategies across commercial chatbot systems and was published at NDSS 2024.
Feb 26, 2024
A comprehensive analysis of jailbreak attack and defense techniques for Large Language Models.
Feb 20, 2024
Novel attack framework exploiting RAG mechanisms to jailbreak LLMs through retrieval database poisoning. Distinguished Paper Award winner.
Feb 1, 2024
PANDORA studies jailbreak attacks against retrieval-augmented generation systems through retrieval database poisoning. The work received the Distinguished Paper Award at AISCC 2024 and highlights a practical attack surface in RAG-enhanced LLM deployments.
Feb 1, 2024
A comprehensive study of prompt injection attacks against LLM-integrated applications.
Jun 9, 2023