📝 Publications

# denotes co-first authors

🗣 Text-to-Speech

Technique Report
SwanTale

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
Yu Zhang#, Ruiqi Li#, Changhao Pan#, et al.

Project | #1st of the Daily Paper

  • SwanTale is a unified model for multi-speaker expressive speech and audio generation across instruct and zero-shot tasks.
Technique Report
SwanVoice

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
Ruiqi Li#, Yu Zhang#, Changhao Pan#, et al.

Project | #3rd of the Daily Pape

  • SwanVoice is a zero-shot TTS model for expressive long-form monologue and dialogue with one to four speakers.
ACL 2026
SwanBench-Speech

Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
Changhao Pan, Rui Yang, Han Wang, et al.

Project | GitHub

  • SwanBench-Speech evaluates long-form speech generation across scenario coverage, automatic metrics, and model behavior analysis.

👂 Spatial Audio

NeurIPS 2025
sym

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations
Wenxiang Guo#, Changhao Pan#, Zhiyuan Zhu#, Xintong Hu#, et al.

Hugging Face Demo

  • The largest recorded spatial audio dataset contains four scenarios: daily life, singing, music, and speech, with a total duration of 500 hours.
  • Supports multiple spatial audio tasks: audio spatialization, spatial TTA, acoustic event localization and detection(SELD), etc.
ICML 2026
SwanSphere

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
Ke Lei#, Yu Zhang#, Changhao Pan#, et al.

Project | GitHub

  • A causal autoregressive diffusion transformer architecture that enables streaming high-quality spatial audio generation.

🎙 Singing Voice Synthesis

NeurIPS 2024(Spotlight)
sym

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
Yu Zhang, Changhao Pan#, Wenxiang Guo#, et al.

Hugging Face Demo

  • GTSinger is a large Global, multi-Technique, free-to-use, high-quality singing corpus with realistic music scores, designed for all singing tasks.
  • Our work is promoted by multiple media and forums, such as weixin, weixin, and zhihu.
EMNLP 2025
sym

Versatile Framework for Song Generation with Prompt-based Control
Yu Zhang#, Wenxiang Guo#, Changhao Pan#, et al.

| Project

  • VersBand is a multi-task song generation framework for synthesizing high-quality, aligned songs with prompt-based control.

📹 Audio-Visual Generation

Others