Accès ouvert
2026
preprint
OpenAlex
Francesco Paissan, Luca Della Libera, Mirco Ravanelli, Cem Subakan
Neural audio codecs provide compact discrete representations for speech generation and manipulation. However, most codecs organize tokens as frame-level sequences, making it difficult to study or intervene on global factors of variation. In this work, we propose the Latent Audio Tokenizer for …
Accès ouvert
2026
preprint
OpenAlex
Francesco Paissan, Luca Della Libera, Mirco Ravanelli, Cem Subakan
Neural audio codecs provide compact discrete representations for speech generation and manipulation. However, most codecs organize tokens as frame-level sequences, making it difficult to study or intervene on global factors of variation. In this work, we propose the Latent Audio Tokenizer for …
ca
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Francesco Paissan, Mirco Ravanelli, Cem Subakan
Free-form, text-based audio editing remains a persistent challenge, despite progress in inversion-based neural methods. Current approaches rely on slow inversion procedures, limiting their practicality. We present a virtual-consistency based audio editing system that bypasses inversion by adapting the sampling process of diffusion …
ca
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Ryo Aihara, Yoshiki Masuyama, Francesco Paissan, François G. Germain et autres
Neural audio codecs (NACs) provide compact representations that can be leveraged in many downstream applications, in particular large language models. Yet most NACs encode mixtures of multiple sources in an entangled manner, which may impede efficient downstream processing in applications that need …
us
(code pays fourni par la source)
2026
conference-paper
OpenAlex
Yoshiki Masuyama, Kohei Saijo, Francesco Paissan, Jiangyu Han et autres
Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel systems have been extended to deal with a variable …
us, jp, cz
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Gabriele Santini, Francesco Paissan, Elisabetta Farella
We introduce a probabilistic dynamic quantization method for neural networks that combines the adaptability of dynamic quantization with the low memory footprint of static approaches. Our technique uses a surrogate probabilistic model of pre-activation statistics to estimate quantization parameters before layer execution, …
it
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Francesco Paissan, Gordon Wichern, Yoshiki Masuyama, Ryo Aihara et autres
Time-Frequency (TF) dual-path models are currently among the best performing audio source separation network architectures, achieving state-of-the-art performance in speech enhancement, music source separation, and cinematic audio source separation. While they are characterized by a relatively low parameter count, they still require …
us, jp
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Michele Romani, Francesco Paissan, Andrea Fossà, Elisabetta Farella
Brain-Computer Interfaces (BCIs) suffer from high inter-subject variability and limited labeled data, often requiring lengthy calibration phases. In this work, we present an end-to-end approach that explicitly models the subject dependency using lightweight convolutional neural networks (CNNs) conditioned on the subject's identity. …
Accès ouvert
2025
preprint
OpenAlex
Francesco Paissan, Gordon Wichern, Yoshiki Masuyama, Ryo Aihara et autres
Time-Frequency (TF) dual-path models are currently among the best performing audio source separation network architectures, achieving state-of-the-art performance in speech enhancement, music source separation, and cinematic audio source separation. While they are characterized by a relatively low parameter count, they still require …
2025
conference-paper
OpenAlex
Manuel Barusco, Francesco Borsatti, Davide Dalle Pezze, Francesco Paissan et autres
Visual Anomaly Detection (VAD) has gained significant research attention for its ability to identify anomalous images and pinpoint the specific areas responsible for the anomaly. A key advantage of VAD is its unsupervised nature, which eliminates the need for costly and time-consuming …
it
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Gabriele Santini, Francesco Paissan, Elisabetta Farella
We propose a probabilistic framework for dynamic quantization of neural networks that allows for a computationally efficient input-adaptive rescaling of the quantization parameters. Our framework applies a probabilistic model to the network's pre-activations through a lightweight surrogate, enabling the adaptive adjustment of …
2025
conference-paper
OpenAlex
Eleonora Mancini, Francesco Paissan, Paolo Torroni, Mirco Ravanelli et autres
Speech impairments in Parkinson’s disease (PD) provide significant early indicators for diagnosis. While models for speech-based PD detection have shown strong performance, their interpretability remains underexplored. This study systematically evaluates several explainability methods to identify PD-specific speech features, aiming to support the …
it, ca
(code pays fourni par la source)