Accès ouvert
2026
preprint
OpenAlex
Yoshiki Mitsui, Ryo Aihara, Tatsuhiko Saito, Yoshiki Masuyama et autres
Task-aware unified source separation (TUSS) enables a single model to handle diverse separation tasks by conditioning on input prompts. However, conventional TUSS does not account for downstream task requirements, such as whether the enhanced speech will be used for human listening or …
Accès ouvert
2026
preprint
OpenAlex
Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama, Christoph Boeddeker et autres
In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farther away …
Accès ouvert
2026
preprint
OpenAlex
Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama, Christoph Boeddeker et autres
In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farther away …
us, jp
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern, Christoph Boeddeker et autres
We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs (NABEATs) as a BEATs-based realization of this framework. Audio SSL models are designed to handle a wide …
Accès ouvert
2026
preprint
OpenAlex
Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern, Christoph Boeddeker et autres
We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs (NABEATs) as a BEATs-based realization of this framework. Audio SSL models are designed to handle a wide …
us, jp
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Dominik Klement, Yoshiki Masuyama, Christoph Boeddeker, Kohei Saijo et autres
Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE Challenge aims to advance performance on real-world far-field noisy and reverberant recordings. This technical report describes MERL's submission to the …
Accès ouvert
2026
preprint
OpenAlex
Dominik Klement, Yoshiki Masuyama, Christoph Boeddeker, Kohei Saijo et autres
Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE Challenge aims to advance performance on real-world far-field noisy and reverberant recordings. This technical report describes MERL's submission to the …
us, cz, jp
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Julius Richter, Yoshiki Masuyama, Christoph Boeddeker, Takahiro Edo et autres
We propose a plug-and-play framework for speech enhancement and separation that augments predictive methods with a generative speech prior. Our approach, termed Stochastic Interpolant Prior for Speech (SIPS), builds on stochastic interpolants and leverages their flexibility to bridge predictive and generative modeling. …
Accès ouvert
2026
preprint
OpenAlex
Julius Richter, Yoshiki Masuyama, Christoph Boeddeker, Takahiro Edo et autres
We propose a plug-and-play framework for speech enhancement and separation that augments predictive methods with a generative speech prior. Our approach, termed Stochastic Interpolant Prior for Speech (SIPS), builds on stochastic interpolants and leverages their flexibility to bridge predictive and generative modeling. …
us
(code pays fourni par la source)
2026
conference-paper
OpenAlex
Adrian Meise, Tobias Cord-Landwehr, Christoph Boeddeker, Marc Delcroix et autres
Sound capture by microphone arrays opens the possibility to exploit spatial, in addition to spectral, information for diarization and signal enhancement, two important tasks in meeting transcription. However, there is no one-to-one mapping of positions in space to speakers if speakers move. …
de, jp
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Adrian Meise, Tobias Cord-Landwehr, Christoph Boeddeker, Marc Delcroix et autres
Sound capture by microphone arrays opens the possibility to exploit spatial, in addition to spectral, information for diarization and signal enhancement, two important tasks in meeting transcription. However, there is no one-to-one mapping of positions in space to speakers if speakers move. …
Accès ouvert
2026
preprint
OpenAlex
Adrian Meise, Tobias Cord-Landwehr, Christoph Boeddeker, Marc Delcroix et autres
Sound capture by microphone arrays opens the possibility to exploit spatial, in addition to spectral, information for diarization and signal enhancement, two important tasks in meeting transcription. However, there is no one-to-one mapping of positions in space to speakers if speakers move. …
de, jp
(code pays fourni par la source)