Scheduled Maintenance Notice
Please note that Researcher Profiles will be undergoing scheduled maintenance on Wednesday 7th Oct, from 8:00am to 9:00am. During this time, the Researcher Profiles system will be unavailable. We apologise for any inconvenience and appreciate your understanding.
Select Publications
Preprints
, 2026, How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study, http://dx.doi.org/10.48550/arxiv.2604.03070
, 2026, PUFFERDOS: Efficient and Effective Attack String Generation for Regular Expression Denial of Service Vulnerabilities, http://dx.doi.org/10.48550/arxiv.2606.19654
, 2026, Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment, http://dx.doi.org/10.48550/arxiv.2405.13068
, 2026, Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems, http://dx.doi.org/10.48550/arxiv.2604.03081
, 2026, STEAMROLLER: A Multi-Agent System for Inclusive Automatic Speech Recognition for People who Stutter, http://dx.doi.org/10.48550/arxiv.2601.10223
, 2025, Prompt Injection attack against LLM-integrated Applications, http://dx.doi.org/10.48550/arxiv.2306.05499
, 2025, Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-based Test Oracles, http://dx.doi.org/10.48550/arxiv.2504.12312
, 2025, SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation, http://dx.doi.org/10.48550/arxiv.2508.06194
, 2025, NeuroDeX: Unlocking Diverse Support in Decompiling Deep Neural Network Executables, http://dx.doi.org/10.48550/arxiv.2509.06402
, 2025, TombRaider: Entering the Vault of History to Jailbreak Large Language Models, http://dx.doi.org/10.48550/arxiv.2501.18628
, 2025, Source Code Summarization in the Era of Large Language Models, http://dx.doi.org/10.48550/arxiv.2407.07959
, 2025, Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models, http://dx.doi.org/10.48550/arxiv.2508.12566
, 2025, "Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas, http://dx.doi.org/10.48550/arxiv.2508.07284
, 2025, CrossPL: Evaluating Large Language Models on Cross Programming Language Code Generation, http://dx.doi.org/10.48550/arxiv.2507.19904
, 2025, MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation, http://dx.doi.org/10.48550/arxiv.2503.10497
, 2025, A Rusty Link in the AI Supply Chain: Detecting Evil Configurations in Model Repositories, http://dx.doi.org/10.48550/arxiv.2505.01067
, 2025, ai.txt: A Domain-Specific Language for Guiding AI Interactions with the Internet, http://dx.doi.org/10.48550/arxiv.2505.07834
, 2025, Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation, http://dx.doi.org/10.48550/arxiv.2505.01065
, 2025, Demystifying RCE Vulnerabilities in LLM-Integrated Apps, http://dx.doi.org/10.48550/arxiv.2309.02926
, 2025, IllusionCAPTCHA: A CAPTCHA based on Visual Illusion, http://dx.doi.org/10.48550/arxiv.2502.05461
, 2025, Detecting LLM Fact-conflicting Hallucinations Enhanced by Temporal-logic-based Reasoning, http://dx.doi.org/10.48550/arxiv.2502.13416
, 2024, MiniScope: Automated UI Exploration and Privacy Inconsistency Detection of MiniApps via Two-phase Iterative Hybrid Analysis, http://dx.doi.org/10.48550/arxiv.2401.03218
, 2024, TransferFuzz: Fuzzing with Historical Trace for Verifying Propagated Vulnerability Code, http://dx.doi.org/10.48550/arxiv.2411.18347
, 2024, A Survey of Protocol Fuzzing, http://dx.doi.org/10.48550/arxiv.2401.01568
, 2024, Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language Models, http://dx.doi.org/10.48550/arxiv.2405.00648
, 2024, Image-Based Geolocation Using Large Vision-Language Models, http://dx.doi.org/10.48550/arxiv.2408.09474
, 2024, How Effective Are They? Exploring Large Language Model Based Fuzz Driver Generation, http://dx.doi.org/10.48550/arxiv.2307.12469
, 2024, Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models, http://dx.doi.org/10.48550/arxiv.2407.13796
, 2024, PentestGPT: An LLM-empowered Automatic Penetration Testing Tool, http://dx.doi.org/10.48550/arxiv.2308.06782
, 2024, A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models, http://dx.doi.org/10.48550/arXiv.2402.13457
, 2024, Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective Detection, http://dx.doi.org/10.48550/arxiv.2404.09894
, 2024, Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study, http://dx.doi.org/10.48550/arxiv.2305.13860
, 2024, MeTMaP: Metamorphic Testing for Detecting False Vector Matching Problems in LLM Augmented Generation, http://dx.doi.org/10.48550/arxiv.2402.14480
, 2024, Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning, http://dx.doi.org/10.48550/arxiv.2402.08416
, 2024, Digger: Detecting Copyright Content Mis-usage in Large Language Model Training, http://dx.doi.org/10.48550/arxiv.2401.00676
, 2023, MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots, http://dx.doi.org/10.48550/arxiv.2307.08715
, 2023, ASTER: Automatic Speech Recognition System Accessibility Testing for Stutterers, http://dx.doi.org/10.48550/arxiv.2308.15742
, 2023, ACETest: Automated Constraint Extraction for Testing Deep Learning Operators, http://dx.doi.org/10.48550/arxiv.2305.17914
, 2023, Automatic Code Summarization via ChatGPT: How Far Are We?, http://dx.doi.org/10.48550/arxiv.2305.12865
, 2022, Morest: Model-based RESTful API Testing with Execution Feedback, http://dx.doi.org/10.48550/arxiv.2204.12148
, 2020, MUZZ: Thread-aware Grey-box Fuzzing for Effective Bug Hunting in Multithreaded Programs, http://dx.doi.org/10.48550/arxiv.2007.15943
, 2020, An Empirical Study on Benchmarks of Artificial Software Vulnerabilities, http://dx.doi.org/10.48550/arxiv.2003.09561
, 2020, LEOPARD: Identifying Vulnerable Code for Vulnerability Assessment through Program Metrics, http://dx.doi.org/10.48550/arxiv.1901.11479
, GeneRAG: Enhancing Large Language Models with Gene-Related Task by Retrieval-Augmented Generation, http://dx.doi.org/10.1101/2024.06.24.600176