I am a master student at College of Computer Science and Technology, Zhejiang University, majoring in Computer Science.
Currently I work on the Audio Research Team at Zhejiang University, under the supervision of Prof. Zhou Zhao. Previously I graduated from Turing Class, a program established by Chu Kochen Honors College, with a bachelor’s degree in Artificial Intelligence.
My research interests primarily focus on Multi-Modal Generative AI, specifically in Speech, Spatial Audio, and Singing. I have published papers at top international AI conference, including NeurIPS, ACL, ACM-MM and EMNLP. Currently, I am working on Spatial Audio Generation and Immersive Audio Synthesis.
I am actively looking for academic collaboration, feel free to contact me via email at panch@zju.edu.cn.
🔥 News
- 2026.08 🎉 3 Papers are accepted by EMNLP 2026!
- 2026.08 🚀 We release SwanTale, Hugging Face’s #1 Paper of the Day!
- 2026.06 🚀 We release SwanVoice, Hugging Face’s #3 Paper of the Day!
- 2026.05 🎉 2 Papers are accepted by ICML 2026!
- 2026.04 🎉 Swanbench-Speech is accepted by ACL!
- 2025.09 🎉 MRSAudio is accepted by NeurIPS 2025!
- 2024.10 🎉 I am awarded Chu Kochen Scholarship!
- 2024.09 🎉 GTSinger is accepted by NeurIPS 2024(spotlight) and 1 paper is accepted by EMNLP 2024!
📝 Publications
# denotes co-first authors
🗣 Text-to-Speech

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
Yu Zhang#, Ruiqi Li#, Changhao Pan#, et al.
Project | #1st of the Daily Paper
- SwanTale is a unified model for multi-speaker expressive speech and audio generation across instruct and zero-shot tasks.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
Ruiqi Li#, Yu Zhang#, Changhao Pan#, et al.
Project | #3rd of the Daily Pape
- SwanVoice is a zero-shot TTS model for expressive long-form monologue and dialogue with one to four speakers.

Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
Changhao Pan, Rui Yang, Han Wang, et al.
Project |
- SwanBench-Speech evaluates long-form speech generation across scenario coverage, automatic metrics, and model behavior analysis.
-
PreprintAudio Editing in the Era of Foundation Models: A Survey, Changhao Pan, Yifei Fan, Fan Zhuo, Yifu Chen, Wenxiang Guo, Yu Zhang, et al. | Project -
PreprintVoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching, Wenxiang Guo, Changhao Pan, Ziyue Jiang, Zhou Zhao, Fei Wu. | Project -
ACL 2026Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness, Jingyu Lu, Yuhan Wang, Fan Zhuo, Xize Cheng, Changhao Pan, et al. -
EMNLP 2026Speaking While Listening: A Survey and Empirical Audit of Full-Duplex Spoken Dialogue Systems, Jingyu Lu, Yuhan Wang, Jianming Luo, Yifu Chen, Tianle Liang, Shengpeng Ji, Ziyue Jiang, Xiaoda Yang, Yu Zhang, Changhao Pan et al. | Project
👂 Spatial Audio

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations
Wenxiang Guo#, Changhao Pan#, Zhiyuan Zhu#, Xintong Hu#, et al.
- The largest recorded spatial audio dataset contains four scenarios: daily life, singing, music, and speech, with a total duration of 500 hours.
- Supports multiple spatial audio tasks: audio spatialization, spatial TTA, acoustic event localization and detection(SELD), etc.

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
Ke Lei#, Yu Zhang#, Changhao Pan#, et al.
Project |
- A causal autoregressive diffusion transformer architecture that enables streaming high-quality spatial audio generation.
-
ACM-MM-2025A Multimodal Evaluation Framework for Spatial Audio Playback Systems: From Localization to Listener Preference, Changhao Pan#, Wenxiang Guo#, Yu Zhang#, et al. | Project -
ACMMM-2025ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting, Yu Zhang#, Wenxiang Guo#, Changhao Pan#, et al. |Project
-
EMNLP-2026One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing, Ke Lei, Chenyuhao Wen, Yu Zhang, Wenxiang Guo, Changhao Pan, et al. -
EMNLP-2026CSAVocoder: A Causal Spatial Audio Vocoder Towards Real-Time Spatial Audio Generation, Zhiyuan Zhu, Han Wang, Wenxiang Guo, Yu Zhang, Changhao Pan, et al. -
AACL-IJCNLP-2025ASAudio: A Survey of Advanced Spatial Audio Research, Zhiyuan Zhu, Yu Zhang, Wenxiang Guo, Changhao Pan, Zhou Zhao. | -
PreprintSpatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding, Zhiyuan Zhu, Yixuan Chen, Yiwen Shao, Wenxiang Guo, Changhao Pan, Yu Zhang, et al. |
🎙 Singing Voice Synthesis

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
Yu Zhang, Changhao Pan#, Wenxiang Guo#, et al.

Versatile Framework for Song Generation with Prompt-based Control
Yu Zhang#, Wenxiang Guo#, Changhao Pan#, et al.
| Project
- VersBand is a multi-task song generation framework for synthesizing high-quality, aligned songs with prompt-based control.
-
AACL-IJCNLP-2025(Oral)Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches, Changhao Pan, Dongyu Yao, Yu Zhang, et al. | -
ACL 2025(Findings)STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation, Wenxiang Guo#, Yu Zhang#, Changhao Pan#, et al. | Project | -
ACL 2025(Findings)TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis, Yu Zhang#, Wenxiang Guo#, Changhao Pan#, et al. | Project | -
EMNLP-2024TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control, Yu Zhang, Ziyue Jiang, Ruiqi Li, Changhao Pan, Jinzheng He, Rongjie Huang, Chuxin Wang, Zhou Zhao. | Project | -
AAAI-2025TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow Matching, Wenxiang Guo, Yu Zhang, Changhao Pan, et. al. | Project |
📹 Audio-Visual Generation
-
Technical ReportALIVE: Animate Your World with Lifelike Audio-Video Generation, Ying Guo, Qijun Gan, Yifu Zhang, et al.
Contributor: Changhao Pan -
ICML-2026TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation, Xiaoda yang#, Majun Zhang#, Changhao Pan, et al. -
ArxivImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks, Jiayang Xu, Fan Zhuo, Majun Zhang, Changhao Pan, et al.
Others
IEEE-TVCGInteractive Table Synthesis with Natural Language, Yanwei Huang, Yunfan Zhou, Ran Chen, Changhao Pan, Xinhuan Shu, Di Weng, Yingcai Wu.
🎖 Honors and Awards
- Chu Kochen Scholarship(as undergraduate), 2024
- Highest scholarship at Zhejiang University
- Chinese National Scholarship, 2022, 2023, 2024
- Awarded by Ministry of Education of China; Top 1%; 3 consecutive times
- CCF Outstanding Undergraduate Students, 2024
- Awarded to 100 undergraduates in the field of Computer Scicence
- Top-10 Outstanding Undergraduate Students of the College of Computer Science and Technology, 2024
- Top-10 Outstanding Students of the Chu Kozhen Honors College, 2023
- BaoGang Elite Scholarship, 2023
📖 Educations
- 2025.09 - 2028.06(Expected), Master, College of Computer Science and Technology, Zhejiang University, Hangzhou, Zhejiang
- Major: Computer Science
- 2021.09 - 2025.06, Undergraduate, Chu Kochen Honors College & College of Computer Science and Technology, Zhejiang Univeristy, Hangzhou, Zhejiang
- Major: Artificial Intelligence (Turing Honor Program)
- GPA: 4.85/5.0, 94.04/100, Rank: 1/79
- 2018.09 - 2021.06, Yuying Experimental School, Wenzhou, Zhejiang
💻 Research & Internships
- 2023.01-2023.09 Research Assisant in State Key Lab of CAD&CG at Zhejiang University
. Advisor: Prof. Yingcai Wu (巫英才). - 2023.09-2024.06 Research Assisant in Audio Research Team at Zhejiang University
. Advisor: Prof. Zhou Zhao (赵洲). - 2025.08-Now AI Algorithm Intern at Bytedance Commerical AI Group
, HangZhou. Mentor: Ruiqi Li, Yu Zhang
📚 Academic Service
- Conference Reviewer: NeurIPS 2025 & 2026, ACL 2025 & 2026, AAAI 2026, ACM-MM 2026