Xuanjun Chen

Xuanjun Chen (陳炫均)

Ph.D. Candidate, NTU EECS · Taipei, Taiwan

I am a Ph.D. candidate at National Taiwan University (NTU), advised by Jyh-Shing Roger Jang and Hung-yi Lee. My research explores the intersection of large language models and speech/audio processing, focusing primarily on training and benchmarking foundation models for audio understanding and generation. I have published 10+ first/co-first papers in top-tier venues (e.g., COLM, TASLP, TIST, ICASSP, INTERSPEECH), with selected honors including the  NVIDIA Academic Grant, the IEEE SPS IEEE Signal Processing Society Scholarship, and a Best Student Paper nomination at ASRU 2025. Additionally, I am actively looking for a 2027 research internship.

Feel free to reach out at d12942018 [at] ntu.edu.tw, or find me on LinkedIn, GitHub, and 𝕏.

Research Highlights * equal contribution, † co-second

See all publications by topic, or my Google Scholar for a full list.

Preprints & Technical Reports

[1]
Can LLM Agents Automate Reinforcement Learning for Text-to-Speech?
Xuanjun Chen, Zixiong Su, Hao Shi, Chang Zeng, Kai Li, Jyh-Shing Roger Jang, Hung-yi Lee
Technical Report, Oct. 2026 bib · arXiv
[2]
MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses
Xuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu, Yinghao Ma, Jyh-Shing Roger Jang, Hung-yi Lee
Technical Report, Sep. 2026 bib · arXiv
[3]
Towards audio language modeling-an overview
Haibin Wu, Xuanjun Chen, Yi-Cheng Lin, Kai-wei Chang, Ho-Lam Chung, Alexander Liu, Hung-yi Lee
Technical Report, Feb. 2024 88 Citations as of Oct. 2026 bib · arXiv · Awesome

International Journals

[1]
CodaRAG: Connecting the Dots with Associativity Inspired by Complementary Learning
Cheng-Yen Li*, Xuanjun Chen*, Claire Lin, Wei-Yu Chen, Wenhua Nie, Hung-yi Lee, Jyh-Shing Roger Jang
ACM Trans. Intell. Syst. Technol. (ACM TIST) 2026 JCR Q1, IF 10.7 bib · arXiv
[2]
CodecFake+: Codec-Based Resynthesized Data as a Proxy for Detecting CodecFake Speech
Xuanjun Chen*, Jiawei Du*, Haibin Wu, Lin Zhang, I-Ming Lin, ..., Jyh-Shing Roger Jang, Hung-yi Lee
IEEE Trans. Audio Speech Lang. Process. (IEEE TASLP) 2026 JCR Q1, IF 5.1 bib · arXiv · IEEE · Project · HF · Code
[3]
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
Ke-Han Lu, Zhehuai Chen, Szu-Wei Fu, ..., Xuanjun Chen, ..., Boris Ginsburg, Yu-Chiang Frank Wang, Hung-yi Lee
IEEE Trans. Audio Speech Lang. Process. (IEEE TASLP) 2026 JCR Q1, IF 5.1 bib · arXiv · IEEE · Code

International Conferences

[1]
Only Ask What You Don't Know: Grounded Delta Planning for Efficient Multi-step RAG
Wei-Chieh Chou*, Xuanjun Chen*, Jian-Ren Lin, Claire Lin, Hung-yi Lee, Jyh-Shing Roger Jang
Conf. Lang. Model. (COLM) 2026 Accept Rate 29.1% bib · arXiv
[2]
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang, ..., Xuanjun Chen, ..., Shinji Watanabe, Hung-yi Lee
Int. Conf. Learn. Represent. (ICLR) 2025 Accept Rate 32.08% bib · arXiv · OpenReview · Code
[3]
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
Haibin Wu*, Ho-Lam Chung*, Yi-Cheng Lin†, Yuan-Kuei Wu†, Xuanjun Chen†, ..., Alexander H. Liu, Hung-yi Lee
Findings Assoc. Comput. Linguist. (ACL) 2024 Accept Rate 22.1% 311 GitHub Stars bib · arXiv · Anthology · Leaderboard · Code · HF
[4]
Joint Fullband-Subband Modeling for High-Resolution SingFake Detection
Xuanjun Chen*, Chia-Yu Hu*, Sung-Feng Huang, Haibin Wu, Hung-yi Lee, Jyh-Shing Roger Jang
INTERSPEECH 2026 (Long Oral) Accept Rate 30% bib · arXiv
[5]
How Does Instrumental Music Help SingFake Detection?
Xuanjun Chen, Chia-Yu Hu, I-Ming Lin, Yi-Cheng Lin, I-Hsiang Chiu, ..., Hung-yi Lee, Jyh-Shing Roger Jang
IEEE Int. Conf. Acoust. Speech Signal Process. (ICASSP) 2026 bib · arXiv · IEEE
[6]
Towards Generalized Source Tracing for Codec-Based Deepfake Speech
Xuanjun Chen*, I-Ming Lin*, Lin Zhang, Haibin Wu, Hung-yi Lee, Jyh-Shing Roger Jang
IEEE Autom. Speech Recognit. Underst. Workshop (IEEE ASRU) 2025 Best Student Paper Nominee (Top 2.5%) bib · arXiv · Code
[7]
Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
Xuanjun Chen*, I-Ming Lin*, Lin Zhang, Jiawei Du, Haibin Wu, Hung-yi Lee, Jyh-Shing Roger Jang
INTERSPEECH 2025 bib · arXiv · ISCA · Code
[8]
Singing Voice Graph Modeling for SingFake Detection
Xuanjun Chen, Haibin Wu, Jyh-Shing Roger Jang, Hung-yi Lee
INTERSPEECH 2024 (Oral) 35 Citations as of Oct. 2026 bib · arXiv · ISCA · Code · Lightning Talk

Experience

  • : Lead Student Researcher, NVIDIA Academic Grant Program, Taipei, Taiwan
  • : Research Intern, Shanda Group, Tokyo, Japan
  • : Ph.D. in Communication EngineeringCommunication Engineering, National Taiwan University ·  GPA 4.3/4.3
  • : M.S. in Computer ScienceComputer Science, National Taiwan University ·  GPA 4.19/4.3
  • : B.S. in Computer ScienceComputer Science, Taiwan Tech ·  GPA 4.11/4.3

Selected Honors & Service

Research and Academic Honors

  • : IEEE Signal Processing Society Scholarship One of dozens of international recipients
  • : NVIDIA Academic Grant Recipient Awarded dedicated GPU compute via competitive global selection
  • : NTU Mr. Wen Tzu-Hsiang Memorial Scholarship One of nine recipients at NTU
  • : Best Student Paper nominee The IEEE Automatic Speech Recognition and Understanding Workshop
  • : Best Paper Award The 37th Conference on Computational Linguistics and Speech Processing
  • : Student Travel Grants Google APAC 2024, ACLCLP 2025 & NSTC Subsidy 2025
  • : CTCI Foundation Science and Technology Scholarship Research Scholarship 2025 & Bursary Award 2024
  • : Ranked 3rd of 42 worldwide On the LA track of ASVspoof 2021 challenge
  • : Taipei City Guangdong Association Scholarship Awarded across six academic years
  • : Certificate of Achievement Top 5% in CSIE, Taiwan Tech, across three semesters
  • : 3rd Prize, SZIIT Academic Award Top 20% of students
  • : National Encouragement Scholarship Ministry of Education

International Event Organizer, Speaker & Reviewer