2025
conference-paper
OpenAlex
Yunlong Zhao, Xiyun Li, Ziyi Wang, Haoran Wu et autres
Large language models (LLMs) face significant challenges in long-context modeling due to increased inference costs, higher latency, and performance degradation caused by information redundancy. Context compression offers a promising solution, but existing methods often rely on fixed strategies that don’t adapt to …
cn
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Minglun Han, Ye Bai, Chen Shen, Youjia Huang et autres
Speech self-supervised pre-training can effectively improve the performance of downstream tasks. However, previous self-supervised learning (SSL) methods for speech, such as HuBERT and BEST-RQ, focus on utilizing non-causal encoders with bidirectional context, and lack sufficient support for downstream streaming models. To address …
Accès ouvert
2024
preprint
OpenAlex
Ye Bai, Jingping Chen, Jitong Chen, Wei Chen et autres
Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios. Classic end-to-end models fused with extra language models perform well, but mainly in …
2024
conference-paper
OpenAlex
Ziyi Ni, Minglun Han, Feilong Chen, Linghui Meng et autres
Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived from human lip motions. In fact, context-dependent visual and linguistic cues …
cn
(code pays fourni par la source)
Accès ouvert
2023
conference-paper
OpenAlex
Feilong Chen, Minglun Han, Jing Shi, Shuang Xu et autres
A compositional question refers to a question that involves multiple visual objects, as well as their attributes and relationships, which requires compositional reasoning to answer.Existing VQA models can well answer a compositional question, but few works can give the reasoning process and …
cn
(code pays fourni par la source)
2023
conference-paper
OpenAlex
Minglun Han, Feilong Chen, Jing Cheng Shi, Shuang Xu et autres
Accès ouvert
2023
conference-paper
OpenAlex
Qingyu Wang, Tielin Zhang, Minglun Han, Yi Wang et autres
The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still relatively simple compared to that in the biological brain. Further research on more types of neurons with …
cn
(code pays fourni par la source)
Accès ouvert
2023
preprint
OpenAlex
Ziyi Ni, Minglun Han, Feilong Chen, Linghui Meng et autres
Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived from human lip motions. In fact, context-dependent visual and linguistic cues …
Accès ouvert
2023
preprint
OpenAlex
Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang et autres
Large language models (LLMs) have demonstrated remarkable language abilities. GPT-4, based on advanced LLMs, exhibits extraordinary multimodal capabilities beyond previous visual language models. We attribute this to the use of more advanced LLMs compared with previous multimodal models. Unfortunately, the model architecture …
2023
conference-paper
OpenAlex
Zefa Hu, Xiuyi Chen, Haoran Wu, Minglun Han et autres
Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficient term semantics learning makes existing approaches hard to capture semantically identical but colloquial expressions of terms …
cn
(code pays fourni par la source)
Accès ouvert
2023
preprint
OpenAlex
Zefa Hu, Xiuyi Chen, Haoran Wu, Minglun Han et autres
Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficient term semantics learning makes existing approaches hard to capture semantically identical but colloquial expressions of terms …
Accès ouvert
2023
preprint
OpenAlex
Minglun Han, Qingyu Wang, Tielin Zhang, Yi Wang et autres
The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still relatively simple compared to that in the biological brain. Further research on more types of neurons with …