Accès ouvert
2023
preprint
OpenAlex
Gemini Robotics Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac et autres
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging from complex reasoning tasks to on-device memory-constrained …
2023
conference-paper
OpenAlex
Mingqiu Wang, Wei Han, Izhak Shafran, Zelin Wu et autres
We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM freezes the pretrained foundation models to maximally preserves their capabilities, and only trains a simple adapter …
us
(code pays fourni par la source)
Accès ouvert
2023
preprint
OpenAlex
Mingqiu Wang, Wei Han, Izhak Shafran, Zelin Wu et autres
We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM freezes the pretrained foundation models to maximally preserves their capabilities, and only trains a simple adapter …
Accès ouvert
2023
preprint
OpenAlex
Nanxin Chen, Izhak Shafran, Yu Zhang, Chung‐Cheng Chiu et autres
Large pre-trained speech models are widely used as the de-facto paradigm, especially in scenarios when there is a limited amount of labeled data available. However, finetuning all parameters from the self-supervised learned model can be computationally expensive, and becomes infeasiable as the …
Accès ouvert
2023
preprint
OpenAlex
Zih-Ching Chen, Chao-Han Huck Yang, Bo Li, Yu Zhang et autres
In this work, we introduce a "score-based assessment" framework for estimating the transferability of pre-trained speech models (PSMs) for fine-tuning target tasks. We leverage upon two representation theories, Bayesian likelihood estimation and optimal transport, to generate rank scores for the PSM candidates …
Accès ouvert
2023
preprint
OpenAlex
Chao-Han Huck Yang, Bo Li, Yu Zhang, Nanxin Chen et autres
In this work, we propose a new parameter-efficient learning framework based on neural model reprogramming for cross-lingual speech recognition, which can \textbf{re-purpose} well-trained English automatic speech recognition (ASR) models to recognize the other languages. We design different auxiliary neural architectures focusing on …
Accès ouvert
2022
preprint
OpenAlex
Chao-Han Huck Yang, Bo Li, Yu Zhang, Nanxin Chen et autres
We propose a quantum kernel learning (QKL) framework to address the inherent data sparsity issues often encountered in training large-scare acoustic models in low-resource scenarios. We project acoustic features based on classical-to-quantum feature encoding. Different from existing quantum convolution techniques, we utilize …
Accès ouvert
2022
preprint
OpenAlex
Nobuyuki Morioka, Heiga Zen, Nanxin Chen, Yu Zhang et autres
Adapting a neural text-to-speech (TTS) model to a target speaker typically involves fine-tuning most if not all of the parameters of a pretrained multi-speaker backbone model. However, serving hundreds of fine-tuned neural TTS models is expensive as each of them requires significant …
Accès ouvert
2022
preprint
OpenAlex
Zhehuai Chen, Ankur Bapna, Andrew E. Rosenberg, Yu Zhang et autres
Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matched joint speech and text model can be leveraged to train a massively multilingual ASR model without any supervised (manually …
2022
conference-paper
OpenAlex
Yan Tong, Zhaoyu Ku, Nanxin Chen, Hu Sheng
VCB is an important component to ensure the safe and smooth operation of the power system. As an important driving part of the vacuum circuit breaker, the operating mechanism is prone to mechanical failure, which leads to power grid accidents. This paper …
cn
(code pays fourni par la source)
Accès ouvert
2022
dataset
OpenAlex
Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang et autres
This is preprocessed data and pretrained models from two of our papers: "Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings," by Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi. (ICASSP 2020) https://arxiv.org/abs/1910.10838 "Pretraining Strategies, Waveform …
jp, us
(code pays fourni par la source)
Accès ouvert
2022
dataset
OpenAlex
Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang et autres
This is preprocessed data and pretrained models from two of our papers: "Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings," by Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi. (ICASSP 2020) https://arxiv.org/abs/1910.10838 "Pretraining Strategies, Waveform …
jp, us
(code pays fourni par la source)