Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision Paper • 2401.00273 • Published Dec 30, 2023 • 1
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks Paper • 2411.05361 • Published Nov 8, 2024 • 6
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment Paper • 2507.02768 • Published Jul 3, 2025 • 19
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling Paper • 2506.00736 • Published May 31, 2025 • 11
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages Paper • 2310.03018 • Published Mar 18, 2024
Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis Paper • 2607.06027 • Published Jul 11 • 1
FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation Paper • 2607.10421 • Published Jul 11 • 2
Improving generalizability of distilled self-supervised speech processing models under distorted settings Paper • 2210.07978 • Published Oct 20, 2022
Improving Distortion Robustness of Self-supervised Speech Processing Tasks with Domain Adaptation Paper • 2203.16104 • Published Jul 25, 2022
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Paper • 2507.09834 • Published Jul 14, 2025 • 1
Ensemble knowledge distillation of self-supervised speech models Paper • 2302.12757 • Published Feb 24, 2023
Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis Paper • 2607.06027 • Published Jul 11 • 1
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision Paper • 2401.00273 • Published Dec 30, 2023 • 1
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks Paper • 2411.05361 • Published Nov 8, 2024 • 6
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Paper • 2507.09834 • Published Jul 14, 2025 • 1
FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation Paper • 2607.10421 • Published Jul 11 • 2