Scheduled Maintenance Notice

Please note that Researcher Profiles will be undergoing scheduled maintenance on Wednesday 7th Oct, from 8:00am to 9:00am. During this time, the Researcher Profiles system will be unavailable. We apologise for any inconvenience and appreciate your understanding.

Select Publications

Preprints

Chen Z; Zhang Y; Liu Y; Deng G; Li Y; Zhang Y; Ning J; Zhang LY; Ma L; Li Z, 2026, How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study, http://dx.doi.org/10.48550/arxiv.2604.03070

Xu S; Ding Z; Cheng X; Li Y; Sun N; Turnbull B; Kan S; Ma S, 2026, PUFFERDOS: Efficient and Effective Attack String Generation for Regular Expression Denial of Service Vulnerabilities, http://dx.doi.org/10.48550/arxiv.2606.19654

Li Y; Liu Y; Li Y; Shi L; Deng G; Chen S; Wang K, 2026, Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment, http://dx.doi.org/10.48550/arxiv.2405.13068

Qu Y; Liu Y; Geng T; Deng G; Li Y; Zhang LY; Zhang Y; Ma L, 2026, Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems, http://dx.doi.org/10.48550/arxiv.2604.03081

Xu Z; Liu Y; Li Y; Shi L; Wang K; Zhao Y, 2026, STEAMROLLER: A Multi-Agent System for Inclusive Automatic Speech Recognition for People who Stutter, http://dx.doi.org/10.48550/arxiv.2601.10223

Liu Y; Deng G; Li Y; Wang K; Wang Z; Wang X; Zhang T; Liu Y; Wang H; Zheng Y; Zhang LY; Liu Y, 2025, Prompt Injection attack against LLM-integrated Applications, http://dx.doi.org/10.48550/arxiv.2306.05499

Xu Z; Ding J; Lou Y; Zhang K; Gong D; Li Y, 2025, Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-based Test Oracles, http://dx.doi.org/10.48550/arxiv.2504.12312

Jiang L; Li Y; Zhang X; Ding Y; Pan L, 2025, SceneJailEval: A Scenario-Adaptive Multi-Dimensional Framework for Jailbreak Evaluation, http://dx.doi.org/10.48550/arxiv.2508.06194

Li Y; Meng G; Sun M; Wang Y; Sun K; Chang H; Li Y, 2025, NeuroDeX: Unlocking Diverse Support in Decompiling Deep Neural Network Executables, http://dx.doi.org/10.48550/arxiv.2509.06402

Ding J; Zhang J; Liu Y; Ding Z; Deng G; Li Y, 2025, TombRaider: Entering the Vault of History to Jailbreak Large Language Models, http://dx.doi.org/10.48550/arxiv.2501.18628

Sun W; Miao Y; Li Y; Zhang H; Fang C; Liu Y; Deng G; Liu Y; Chen Z, 2025, Source Code Summarization in the Era of Large Language Models, http://dx.doi.org/10.48550/arxiv.2407.07959

Song W; Zhong H; Ding Z; Xue J; Li Y, 2025, Help or Hurdle? Rethinking Model Context Protocol-Augmented Large Language Models, http://dx.doi.org/10.48550/arxiv.2508.12566

Ding J; Jiang P; Xu Z; Ding Z; Zhu Y; Jiang J; Li Y, 2025, "Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas, http://dx.doi.org/10.48550/arxiv.2508.07284

Xiong Z; Wang D; Li Y; An X; Wang W, 2025, CrossPL: Evaluating Large Language Models on Cross Programming Language Code Generation, http://dx.doi.org/10.48550/arxiv.2507.19904

Xuan W; Yang R; Qi H; Zeng Q; Xiao Y; Feng A; Liu D; Xing Y; Wang J; Gao F; Lu J; Jiang Y; Li H; Li X; Yu K; Dong R; Gu S; Li Y; Xie X; Juefei-Xu F; Khomh F; Yoshie O; Chen Q; Teodoro D; Liu N; Goebel R; Ma L; Marrese-Taylor E; Lu S; Iwasawa Y; Matsuo Y; Li I, 2025, MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation, http://dx.doi.org/10.48550/arxiv.2503.10497

Ding Z; Fu Q; Ding J; Deng G; Liu Y; Li Y, 2025, A Rusty Link in the AI Supply Chain: Detecting Evil Configurations in Model Repositories, http://dx.doi.org/10.48550/arxiv.2505.01067

Li Y; Song W; Zhu B; Gong D; Liu Y; Deng G; Chen C; Ma L; Sun J; Walsh T; Xue J, 2025, ai.txt: A Domain-Specific Language for Guiding AI Interactions with the Internet, http://dx.doi.org/10.48550/arxiv.2505.07834

Jin D; Fu Q; Li Y, 2025, Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation, http://dx.doi.org/10.48550/arxiv.2505.01065

Liu T; Deng Z; Meng G; Li Y; Chen K, 2025, Demystifying RCE Vulnerabilities in LLM-Integrated Apps, http://dx.doi.org/10.48550/arxiv.2309.02926

Ding Z; Deng G; Liu Y; Ding J; Chen J; Sui Y; Li Y, 2025, IllusionCAPTCHA: A CAPTCHA based on Visual Illusion, http://dx.doi.org/10.48550/arxiv.2502.05461

Li N; Song Y; Wang K; Li Y; Shi L; Liu Y; Wang H, 2025, Detecting LLM Fact-conflicting Hallucinations Enhanced by Temporal-logic-based Reasoning, http://dx.doi.org/10.48550/arxiv.2502.13416

Wang S; Li Y; Wang K; Liu Y; Li H; Liu Y; Wang H, 2024, MiniScope: Automated UI Exploration and Privacy Inconsistency Detection of MiniApps via Two-phase Iterative Hybrid Analysis, http://dx.doi.org/10.48550/arxiv.2401.03218

Li S; Li Y; Chen Z; Dong C; Wang Y; Li H; Chen Y; Zhu H, 2024, TransferFuzz: Fuzzing with Historical Trace for Verifying Propagated Vulnerability Code, http://dx.doi.org/10.48550/arxiv.2411.18347

Zhang X; Zhang C; Li X; Du Z; Mao B; Li Y; Zheng Y; Li Y; Pan L; Liu Y; Deng RH, 2024, A Survey of Protocol Fuzzing, http://dx.doi.org/10.48550/arxiv.2401.01568

Li N; Li Y; Liu Y; Shi L; Wang K; Wang H, 2024, Drowzee: Metamorphic Testing for Fact-Conflicting Hallucination Detection in Large Language Models, http://dx.doi.org/10.48550/arxiv.2405.00648

Liu Y; Ding J; Deng G; Li Y; Zhang T; Sun W; Zheng Y; Ge J; Liu Y, 2024, Image-Based Geolocation Using Large Vision-Language Models, http://dx.doi.org/10.48550/arxiv.2408.09474

Zhang C; Zheng Y; Bai M; Li Y; Ma W; Xie X; Li Y; Sun L; Liu Y, 2024, How Effective Are They? Exploring Large Language Model Based Fuzz Driver Generation, http://dx.doi.org/10.48550/arxiv.2307.12469

Xu Z; Liu Y; Deng G; Wang K; Li Y; Shi L; Picek S, 2024, Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models, http://dx.doi.org/10.48550/arxiv.2407.13796

Deng G; Liu Y; Mayoral-Vilches V; Liu P; Li Y; Xu Y; Zhang T; Liu Y; Pinzger M; Rass S, 2024, PentestGPT: An LLM-empowered Automatic Penetration Testing Tool, http://dx.doi.org/10.48550/arxiv.2308.06782

Xu Z; Liu Y; Deng G; Li Y; Picek S, 2024, A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models, http://dx.doi.org/10.48550/arXiv.2402.13457

Li Y; Liu Y; Deng G; Zhang Y; Song W; Shi L; Wang K; Li Y; Liu Y; Wang H, 2024, Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective Detection, http://dx.doi.org/10.48550/arxiv.2404.09894

Liu Y; Deng G; Xu Z; Li Y; Zheng Y; Zhang Y; Zhao L; Zhang T; Wang K; Liu Y, 2024, Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study, http://dx.doi.org/10.48550/arxiv.2305.13860

Wang G; Li Y; Liu Y; Deng G; Li T; Xu G; Liu Y; Wang H; Wang K, 2024, MeTMaP: Metamorphic Testing for Detecting False Vector Matching Problems in LLM Augmented Generation, http://dx.doi.org/10.48550/arxiv.2402.14480

Deng G; Liu Y; Wang K; Li Y; Zhang T; Liu Y, 2024, Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning, http://dx.doi.org/10.48550/arxiv.2402.08416

Li H; Deng G; Liu Y; Wang K; Li Y; Zhang T; Liu Y; Xu G; Xu G; Wang H, 2024, Digger: Detecting Copyright Content Mis-usage in Large Language Model Training, http://dx.doi.org/10.48550/arxiv.2401.00676

Deng G; Liu Y; Li Y; Wang K; Zhang Y; Li Z; Wang H; Zhang T; Liu Y, 2023, MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots, http://dx.doi.org/10.48550/arxiv.2307.08715

Liu Y; Li Y; Deng G; Juefei-Xu F; Du Y; Zhang C; Liu C; Li Y; Ma L; Liu Y, 2023, ASTER: Automatic Speech Recognition System Accessibility Testing for Stutterers, http://dx.doi.org/10.48550/arxiv.2308.15742

Shi J; Xiao Y; Li Y; Li Y; Yu D; Yu C; Su H; Chen Y; Huo W, 2023, ACETest: Automated Constraint Extraction for Testing Deep Learning Operators, http://dx.doi.org/10.48550/arxiv.2305.17914

Sun W; Fang C; You Y; Miao Y; Liu Y; Li Y; Deng G; Huang S; Chen Y; Zhang Q; Qian H; Liu Y; Chen Z, 2023, Automatic Code Summarization via ChatGPT: How Far Are We?, http://dx.doi.org/10.48550/arxiv.2305.12865

Liu Y; Li Y; Deng G; Liu Y; Wan R; Wu R; Ji D; Xu S; Bao M, 2022, Morest: Model-based RESTful API Testing with Execution Feedback, http://dx.doi.org/10.48550/arxiv.2204.12148

Chen H; Guo S; Xue Y; Sui Y; Zhang C; Li Y; Wang H; Liu Y, 2020, MUZZ: Thread-aware Grey-box Fuzzing for Effective Bug Hunting in Multithreaded Programs, http://dx.doi.org/10.48550/arxiv.2007.15943

Geng S; Li Y; Du Y; Xu J; Liu Y; Mao B, 2020, An Empirical Study on Benchmarks of Artificial Software Vulnerabilities, http://dx.doi.org/10.48550/arxiv.2003.09561

Du X; Chen B; Li Y; Guo J; Zhou Y; Liu Y; Jiang Y, 2020, LEOPARD: Identifying Vulnerable Code for Vulnerability Assessment through Program Metrics, http://dx.doi.org/10.48550/arxiv.1901.11479

Lin X; Deng G; Li Y; Ge J; Ho JWK; Liu Y, GeneRAG: Enhancing Large Language Models with Gene-Related Task by Retrieval-Augmented Generation, http://dx.doi.org/10.1101/2024.06.24.600176


Back to profile page