Select Publications
Journal articles
, 2026, 'Code-Switching Speech Recognition under the Lens: Model- and Data-Centric Perspectives', IEEE Transactions on Audio Speech and Language Processing, 34, pp. 1853 - 1865, http://dx.doi.org/10.1109/TASLPRO.2026.3675776
, 2026, 'Why Pre-trained Models Fail: Feature Entanglement in Multimodal Depression Detection', IEEE Transactions on Affective Computing, http://dx.doi.org/10.1109/TAFFC.2026.3711426
, 2025, 'Aligning Speech to Languages to Enhance Code-Switching Speech Recognition', IEEE Transactions on Audio Speech and Language Processing, 33, pp. 4712 - 4725, http://dx.doi.org/10.1109/TASLPRO.2025.3629290
, 2025, 'Mamba in Speech: Towards an Alternative to Self-Attention', IEEE Transactions on Audio Speech and Language Processing, 33, pp. 1933 - 1948, http://dx.doi.org/10.1109/TASLPRO.2025.3566210
, 2025, 'Selective State Space Model for Monaural Speech Enhancement', IEEE Transactions on Consumer Electronics, 71, pp. 5414 - 5424, http://dx.doi.org/10.1109/TCE.2024.3523297
, 2023, 'Twin-S: a digital twin for skull base surgery', International Journal of Computer Assisted Radiology and Surgery, 18, pp. 1077 - 1084, http://dx.doi.org/10.1007/s11548-023-02863-9
Conference Papers
, 2026, 'Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech', in ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing, IEEE, pp. 22697 - 22701, http://dx.doi.org/10.1109/ICASSP55912.2026.11463835
, 2025, 'Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction', in Proceedings of the Annual Conference of the International Speech Communication Association Interspeech, pp. 4263 - 4267, http://dx.doi.org/10.21437/Interspeech.2025-17
, 2025, 'CASPER: A Large Scale Spontaneous Speech Dataset', in Asru 2025 2025 IEEE Automatic Speech Recognition and Understanding Workshop, http://dx.doi.org/10.1109/ASRU65441.2025.11434634
, 2025, 'Long-Context Modeling Networks for Monaural Speech Enhancement: A Comparative Study', in IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, http://dx.doi.org/10.1109/WASPAA66052.2025.11230983
, 2025, 'Multi-Class Dementia Detection Using Acoustic Features - ICASSP-2025 PROCESS Challenge', in ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings, http://dx.doi.org/10.1109/ICASSP49660.2025.10889847
, 2025, 'Rethinking Mamba in Speech Processing by Self-Supervised Models', in ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings, http://dx.doi.org/10.1109/ICASSP49660.2025.10889111
, 2025, 'SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information', in Proceedings of the Annual Meeting of the Association for Computational Linguistics, pp. 10019 - 10030, http://dx.doi.org/10.18653/v1/2025.findings-acl.521
, 2024, 'Binaural Selective Attention Model for Target Speaker Extraction', in Proceedings of the Annual Conference of the International Speech Communication Association Interspeech, pp. 4323 - 4327, http://dx.doi.org/10.21437/Interspeech.2024-683
, 2024, 'ENHANCING CODE-SWITCHING SPEECH RECOGNITION WITH INTERACTIVE LANGUAGE BIASES', in ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings, pp. 10886 - 10890, http://dx.doi.org/10.1109/ICASSP48485.2024.10448335
, 2024, 'Speaking in Wavelet Domain: A Simple and Efficient Approach to Speed up Speech Diffusion Model', in Emnlp 2024 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the Conference, pp. 159 - 171, http://dx.doi.org/10.18653/v1/2024.emnlp-main.9
, 2024, 'Striking a Balance between Classical and Deep Learning Approaches in Natural Language Processing Pedagogy', in Teachnlp 2024 6th Workshop on Teaching Nlp Proceedings of the Workshop, pp. 23 - 32
, 2024, 'Unidirectional Brain-Computer Interface: Artificial Neural Network Encoding Natural Images to FMRI Response in the Visual Cortex', in ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings, pp. 1851 - 1855, http://dx.doi.org/10.1109/ICASSP48485.2024.10446366
, 2024, 'When LLMs Meet Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection', in Emnlp 2024 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the Conference, pp. 146 - 158, http://dx.doi.org/10.18653/v1/2024.emnlp-main.8
, 2023, 'A New Approach to Extract Fetal Electrocardiogram Using Affine Combination of Adaptive Filters', in ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings, http://dx.doi.org/10.1109/ICASSP49357.2023.10095885
, 2023, 'MERLIon CCS Challenge: A English-Mandarin code-switching child-directed speech corpus for language identification and diarization', in Proceedings of the Annual Conference of the International Speech Communication Association Interspeech, pp. 4109 - 4113, http://dx.doi.org/10.21437/Interspeech.2023-1446
, 2023, 'PQLM - Multilingual Decentralized Portable Quantum Language Model', in ICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings, http://dx.doi.org/10.1109/ICASSP49357.2023.10095215
, 2023, 'A Quantitative Approach to Understand Self-Supervised Models as Cross-lingual Feature Extracters', in Proceedings of the 6th International Conference on Natural Language and Speech Processing (ICNLSP 2023), Association for Computational Linguistics, Online, pp. 200 - 211, https://aclanthology.org/2023.icnlsp-1.20/
Reports
, 2026, Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation, arXiv, arXiv:2605.12034, https://arxiv.org/abs/2605.12034
, 2026, DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action, arXiv, arXiv:2605.20755, https://arxiv.org/abs/2605.20755
, 2026, Step-Audio-R1.5 Technical Report, arXiv, arXiv:2604.25719, https://arxiv.org/abs/2604.25719
, 2026, StepAudio 2.5 Technical Report, arXiv, arXiv:2605.23463, https://arxiv.org/abs/2605.23463
, 2025, Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models, arXiv, arXiv:2510.09592, https://arxiv.org/abs/2510.09592
, 2025, Step-audio 2 technical report, arXiv, arXiv:2507.16632, https://arxiv.org/abs/2507.16632
, 2025, Step-Audio-EditX Technical Report, arXiv, arXiv:2511.03601, https://arxiv.org/abs/2511.03601
, 2025, Step-Audio-R1 Technical Report, arXiv, arXiv:2511.15848, https://arxiv.org/abs/2511.15848
Working Papers
, 2026, DepFlow: Disentangled Speech Generation to Mitigate Semantic Bias in Depression Detection, http://dx.doi.org, https://arxiv.org/abs/2601.00303
, 2026, The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models, http://dx.doi.org, https://arxiv.org/abs/2605.29209
, 2026, Why your tokenizer fails in information fusion: A timing-aware pre-quantization fusion for video-enhanced audio tokenization, http://dx.doi.org, https://arxiv.org/abs/2604.12145
, 2025, Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English, http://dx.doi.org, https://arxiv.org/abs/2505.17076
, 2025, Multi-bench: A multi-turn interactive benchmark for assessing emotional intelligence ability of spoken dialogue models, http://dx.doi.org, https://arxiv.org/abs/2511.00850
, 2025, Step-audio: Unified understanding and generation in intelligent speech interaction, http://dx.doi.org, https://arxiv.org/abs/2502.11946
Preprints
, 2026, The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models, https://arxiv.org/abs/2605.29209v1
, 2026, StepAudio 2.5 Technical Report, https://arxiv.org/abs/2605.23463v1
, 2026, DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action, https://arxiv.org/abs/2605.20755v2
, 2026, Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation, https://arxiv.org/abs/2605.12034v2
, 2026, Step-Audio-R1.5 Technical Report, https://arxiv.org/abs/2604.25719v2
, 2026, Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization, https://arxiv.org/abs/2604.12145v1
, 2026, DepFlow: Disentangled Speech Generation to Mitigate Semantic Bias in Depression Detection, https://arxiv.org/abs/2601.00303v1
, 2025, Step-Audio-R1 Technical Report, https://arxiv.org/abs/2511.15848v2
, 2025, Step-Audio-EditX Technical Report, https://arxiv.org/abs/2511.03601v2
, 2025, MULTI-Bench: A Multi-Turn Interactive Benchmark for Assessing Emotional Intelligence ability of Spoken Dialogue Models, https://arxiv.org/abs/2511.00850v1
, 2025, Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models, https://arxiv.org/abs/2510.09592v2
, 2025, Step-Audio 2 Technical Report, https://arxiv.org/abs/2507.16632v3
, 2025, Distinctive Feature Codec: An Adaptive Efficient Speech Representation for Depression Detection, https://arxiv.org/abs/2505.18516v2