Accès ouvert
2026
preprint
OpenAlex
Nethmi Muthugala, Supryadi, Surangika Ranathunga, Nisansa de Silva et autres
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamics. Existing benchmarks overlook Sri …
Accès ouvert
2026
preprint
OpenAlex
Nethmi Muthugala, Supryadi, Surangika Ranathunga, Nisansa de Silva et autres
Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamics. Existing benchmarks overlook Sri …
cn, nz, lk, us
(code pays fourni par la source)
Accès ouvert
2026
conference-paper
OpenAlex
Duc-Tuan Truong, Tianchi Liu, Junjie Li, Ruijie Tao et autres
In speech deepfake detection (SDD), data augmentation (DA) is commonly used to improve model generalization across varied speech conditions and spoofing attacks. However, during training, the backpropagated gradients from original and augmented inputs may misalign, which can result in conflicting parameter updates. …
sg, hk
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Jiadong Wang, Ke Zhang, Xinyuan Qian, Ruijie Tao et autres
Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance, the issue of degraded visual inputs has received relatively little …
Accès ouvert
2026
preprint
OpenAlex
Jiadong Wang, Ke Zhang, Xinyuan Qian, Ruijie Tao et autres
Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance, the issue of degraded visual inputs has received relatively little …
de
(code pays fourni par la source)
2026
article
OpenAlex
Changjiang Wu, Xin Li, Dongfang Wu, Ruijie Tao et autres
) was carried out to synthesize selenotetrazoles under mild conditions. When the solvent system and additives were adjusted, this reaction also efficiently afforded selenocarbamates. The protocol features broad substrate scope and excellent functional-group tolerance, and a plausible reaction mechanism is proposed.
cn
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Mehmet Sinan Yıldırım, Ruijie Tao, Wupeng Wang, Junyi Ao et autres
Target Language Extraction aims to extract speech in a specific language from a mixture waveform that contains multiple speakers speaking different languages. The human auditory system is adept at performing this task with the knowledge of the particular language. However, the performance …
2025
conference-paper
OpenAlex
Ruijie Tao, Zhan Shi, Yidi Jiang, Tianchi Liu et autres
Modern speaker recognition system relies on abundant and balanced datasets for classification training. However, diverse defective datasets, such as partially-labelled, small-scale, and imbalanced datasets, are common in real-world applications. Previous works usually studied specific solutions for each scenario from the algorithm perspective. …
sg, cn
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Jiawei Zhang, Tianhao Zhang, Jun Wang, Ruijie Tao et autres
Controlling the spatial and stylistic characteristics of synthesized speech is essential for immersive and personalized applications such as virtual reality, gaming, and human-computer interaction. While recent Text-to-speech (TTS) systems have explored multi-modal conditioning, they often suffer from poor reverberation fidelity or degraded …
cn, sg
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Mehmet Sinan Yıldırım, Ruijie Tao, Wupeng Wang, Junyi Ao et autres
Target Language Extraction aims to extract speech in a specific language from a mixture waveform that contains multiple speakers speaking different languages. The human auditory system is adept at performing this task with the knowledge of the particular language. However, the performance …
sg, cn
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Tianchi Liu, Ruijie Tao, Qiongqiong Wang, Yidi Jiang et autres
The success of deep learning-based speaker verification systems is largely attributed to access to large-scale and diverse speaker identity data. However, collecting data from more identities is expensive, challenging, and often limited by privacy concerns. To address this limitation, we propose INSIDE …
sg, cn
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Duc-Tuan Truong, Tianchi Liu, Ruijie Tao, Junjie Li et autres
Recent work shows that one-class learning can detect unseen deepfake attacks by modeling a compact distribution of bona fide speech around a single centroid. However, the single-centroid assumption can oversimplify the bona fide speech representation and overlook useful cues, such as speech …