Accès ouvert
2026
conference-paper
OpenAlex
Ryo Aihara, Yoshiki Masuyama, Francesco Paissan, François G. Germain et autres
Neural audio codecs (NACs) provide compact representations that can be leveraged in many downstream applications, in particular large language models. Yet most NACs encode mixtures of multiple sources in an entangled manner, which may impede efficient downstream processing in applications that need …
us
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Ryo Aihara, Yoshiki Masuyama, François G. Germain, Gordon Wichern et autres
Neural audio codecs (NACs), which use neural networks to generate compact audio representations, have garnered interest for their applicability to many downstream tasks, especially quantized codecs due to their compatibility with large language models. However, unlike text, speech conveys not only linguistic …
us
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Francesco Paissan, Gordon Wichern, Yoshiki Masuyama, Ryo Aihara et autres
Time-Frequency (TF) dual-path models are currently among the best performing audio source separation network architectures, achieving state-of-the-art performance in speech enhancement, music source separation, and cinematic audio source separation. While they are characterized by a relatively low parameter count, they still require …
us, jp
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Yoshiki Masuyama, François G. Germain, Gordon Wichern, Christopher Ick et autres
This paper presents a physics-informed neural network (PINN) for modeling first-order Ambisonic (FOA) room impulse responses (RIRs). PINNs have demonstrated promising performance in sound field interpolation by combining the powerful modeling capability of neural networks and the physical principles of sound propagation. …
us
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Takao Kawamura, Yoshiki Masuyama, Nobutaka Ono
In this study, we propose an unsupervised domain adaptation method for multi-channel acoustic scene classification (ASC). Multi-channel ASC leverages spatial features and provides advantages to distinguish scenes with similar spectral features. Meanwhile, it is sensitive to the mismatch of the spatial features …
jp
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Boboš et autres
International audience
2025
conference-paper
OpenAlex
Haici Yang, Gordon Wichern, Ryo Aihara, Yoshiki Masuyama et autres
2025
conference-paper
OpenAlex
Christopher Ick, Gordon Wichern, Yoshiki Masuyama, François G. Germain et autres
Accès ouvert
2025
preprint
OpenAlex
Samuele Cornell, Christoph Boeddeker, Tae‐Jin Park, He Huang et autres
The CHiME-7 and 8 distant speech recognition (DASR) challenges focus on multi-channel, generalizable, joint automatic speech recognition (ASR) and diarization of conversational speech. With participation from 9 teams submitting 32 diverse systems, these challenges have contributed to state-of-the-art research in the field. …
Accès ouvert
2025
preprint
OpenAlex
Francesco Paissan, Gordon Wichern, Yoshiki Masuyama, Ryo Aihara et autres
Time-Frequency (TF) dual-path models are currently among the best performing audio source separation network architectures, achieving state-of-the-art performance in speech enhancement, music source separation, and cinematic audio source separation. While they are characterized by a relatively low parameter count, they still require …
Accès ouvert
2025
preprint
OpenAlex
Yoshiki Masuyama, François G. Germain, Gordon Wichern, Christopher Ick et autres
This paper presents a physics-informed neural network (PINN) for modeling first-order Ambisonic (FOA) room impulse responses (RIRs). PINNs have demonstrated promising performance in sound field interpolation by combining the powerful modeling capability of neural networks and the physical principles of sound propagation. …
Accès ouvert
2025
preprint
OpenAlex
Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Boboš et autres
We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge distillation objectives: one from a pre-trained speech encoder (HuBERT) for phoneme-level structure, and …