Aller au contenu principal
Profil bibliographique

Nanxin Chen

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

56Publications signalées
4918Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

Speech Recognition and SynthesisSpeech and Audio ProcessingMusic and Audio ProcessingNatural Language Processing TechniquesTopic Modeling

Les publications récentes

Accès ouvert 2023 preprint OpenAlex

Gemini: A Family of Highly Capable Multimodal Models

Gemini Robotics Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac et autres

This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitable for applications ranging from complex reasoning tasks to on-device memory-constrained …

843 citations arXiv (Cornell University)
2023 conference-paper OpenAlex

SLM: Bridge the Thin Gap Between Speech and Text Foundation Models

Mingqiu Wang, Wei Han, Izhak Shafran, Zelin Wu et autres

We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM freezes the pretrained foundation models to maximally preserves their capabilities, and only trains a simple adapter …

us (code pays fourni par la source)

29 citations
Accès ouvert 2023 preprint OpenAlex

SLM: Bridge the thin gap between speech and text foundation models

Mingqiu Wang, Wei Han, Izhak Shafran, Zelin Wu et autres

We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM freezes the pretrained foundation models to maximally preserves their capabilities, and only trains a simple adapter …

1 citation arXiv (Cornell University)
Accès ouvert 2023 preprint OpenAlex

Efficient Adapters for Giant Speech Models

Nanxin Chen, Izhak Shafran, Yu Zhang, Chung‐Cheng Chiu et autres

Large pre-trained speech models are widely used as the de-facto paradigm, especially in scenarios when there is a limited amount of labeled data available. However, finetuning all parameters from the self-supervised learned model can be computationally expensive, and becomes infeasiable as the …

4 citations arXiv (Cornell University)
Accès ouvert 2023 preprint OpenAlex

How to Estimate Model Transferability of Pre-Trained Speech Models?

Zih-Ching Chen, Chao-Han Huck Yang, Bo Li, Yu Zhang et autres

In this work, we introduce a "score-based assessment" framework for estimating the transferability of pre-trained speech models (PSMs) for fine-tuning target tasks. We leverage upon two representation theories, Bayesian likelihood estimation and optimal transport, to generate rank scores for the PSM candidates …

0 citations arXiv (Cornell University)
Accès ouvert 2023 preprint OpenAlex

From English to More Languages: Parameter-Efficient Model Reprogramming for Cross-Lingual Speech Recognition

Chao-Han Huck Yang, Bo Li, Yu Zhang, Nanxin Chen et autres

In this work, we propose a new parameter-efficient learning framework based on neural model reprogramming for cross-lingual speech recognition, which can \textbf{re-purpose} well-trained English automatic speech recognition (ASR) models to recognize the other languages. We design different auxiliary neural architectures focusing on …

0 citations arXiv (Cornell University)
Accès ouvert 2022 preprint OpenAlex

A Quantum Kernel Learning Approach to Acoustic Modeling for Spoken Command Recognition

Chao-Han Huck Yang, Bo Li, Yu Zhang, Nanxin Chen et autres

We propose a quantum kernel learning (QKL) framework to address the inherent data sparsity issues often encountered in training large-scare acoustic models in low-resource scenarios. We project acoustic features based on classical-to-quantum feature encoding. Different from existing quantum convolution techniques, we utilize …

2 citations arXiv (Cornell University)
Accès ouvert 2022 preprint OpenAlex

Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR

Zhehuai Chen, Ankur Bapna, Andrew E. Rosenberg, Yu Zhang et autres

Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matched joint speech and text model can be leveraged to train a massively multilingual ASR model without any supervised (manually …

0 citations arXiv (Cornell University)
Accès ouvert 2022 dataset OpenAlex

Preprocessed Data and Pretrained Models for Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings

Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang et autres

This is preprocessed data and pretrained models from two of our papers: "Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings," by Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi. (ICASSP 2020) https://arxiv.org/abs/1910.10838 "Pretraining Strategies, Waveform …

jp, us (code pays fourni par la source)

0 citations Zenodo (CERN European Organization for Nuclear Research)
Accès ouvert 2022 dataset OpenAlex

Preprocessed Data and Pretrained Models for Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings

Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang et autres

This is preprocessed data and pretrained models from two of our papers: "Zero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings," by Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi. (ICASSP 2020) https://arxiv.org/abs/1910.10838 "Pretraining Strategies, Waveform …

jp, us (code pays fourni par la source)

0 citations Zenodo (CERN European Organization for Nuclear Research)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.