Aller au contenu principal
Profil bibliographique

Minglun Han

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

19Publications signalées
398Citations signalées
2Affiliations récentes

Les institutions déclarées

Les domaines associés

Speech Recognition and SynthesisNatural Language Processing TechniquesSpeech and Audio ProcessingMusic and Audio ProcessingTopic Modeling

Les publications récentes

2025 conference-paper OpenAlex

Integrate-and-Fire Compressor: Learning to Compress Context for LLMs Adaptively

Yunlong Zhao, Xiyun Li, Ziyi Wang, Haoran Wu et autres

Large language models (LLMs) face significant challenges in long-context modeling due to increased inference costs, higher latency, and performance degradation caused by information redundancy. Context compression offers a promising solution, but existing methods often rely on fixed strategies that don’t adapt to …

cn (code pays fourni par la source)

0 citations
Accès ouvert 2024 preprint OpenAlex

NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training

Minglun Han, Ye Bai, Chen Shen, Youjia Huang et autres

Speech self-supervised pre-training can effectively improve the performance of downstream tasks. However, previous self-supervised learning (SSL) methods for speech, such as HuBERT and BEST-RQ, focus on utilizing non-causal encoders with bidirectional context, and lack sufficient support for downstream streaming models. To address …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Ye Bai, Jingping Chen, Jitong Chen, Wei Chen et autres

Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios. Classic end-to-end models fused with extra language models perform well, but mainly in …

6 citations arXiv (Cornell University)
2024 conference-paper OpenAlex

ViLaS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition

Ziyi Ni, Minglun Han, Feilong Chen, Linghui Meng et autres

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived from human lip motions. In fact, context-dependent visual and linguistic cues …

cn (code pays fourni par la source)

4 citations
Accès ouvert 2023 conference-paper OpenAlex

Enhancing Visual Question Answering via Deconstructing Questions and Explicating Answers

Feilong Chen, Minglun Han, Jing Shi, Shuang Xu et autres

A compositional question refers to a question that involves multiple visual objects, as well as their attributes and relationships, which requires compositional reasoning to answer.Existing VQA models can well answer a compositional question, but few works can give the reasoning process and …

cn (code pays fourni par la source)

0 citations
Accès ouvert 2023 conference-paper OpenAlex

Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition

Qingyu Wang, Tielin Zhang, Minglun Han, Yi Wang et autres

The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still relatively simple compared to that in the biological brain. Further research on more types of neurons with …

cn (code pays fourni par la source)

28 citations Proceedings of the AAAI Conference on Artificial Intelligence
Accès ouvert 2023 preprint OpenAlex

VILAS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition

Ziyi Ni, Minglun Han, Feilong Chen, Linghui Meng et autres

Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived from human lip motions. In fact, context-dependent visual and linguistic cues …

0 citations arXiv (Cornell University)
Accès ouvert 2023 preprint OpenAlex

X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang et autres

Large language models (LLMs) have demonstrated remarkable language abilities. GPT-4, based on advanced LLMs, exhibits extraordinary multimodal capabilities beyond previous visual language models. We attribute this to the use of more advanced LLMs compared with previous multimodal models. Unfortunately, the model architecture …

22 citations arXiv (Cornell University)
2023 conference-paper OpenAlex

Matching-Based Term Semantics Pre-Training for Spoken Patient Query Understanding

Zefa Hu, Xiuyi Chen, Haoran Wu, Minglun Han et autres

Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficient term semantics learning makes existing approaches hard to capture semantically identical but colloquial expressions of terms …

cn (code pays fourni par la source)

4 citations
Accès ouvert 2023 preprint OpenAlex

Matching-based Term Semantics Pre-training for Spoken Patient Query Understanding

Zefa Hu, Xiuyi Chen, Haoran Wu, Minglun Han et autres

Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficient term semantics learning makes existing approaches hard to capture semantically identical but colloquial expressions of terms …

0 citations arXiv (Cornell University)
Accès ouvert 2023 preprint OpenAlex

Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition

Minglun Han, Qingyu Wang, Tielin Zhang, Yi Wang et autres

The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still relatively simple compared to that in the biological brain. Further research on more types of neurons with …

1 citation arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.