Accès ouvert
2026
preprint
OpenAlex
Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres et autres
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside …
Accès ouvert
2025
dissertation
OpenAlex
Yassir Fathullah
Transformer-based autoregressive sequence models have revolutionised natural language processing and speech processing, achieving state-of-the-art performance on a wide range of tasks. However, their deployment in real-world scenarios, especially safety-critical applications like autonomous systems or medical diagnosis, necessitates not only high accuracy but …
Accès ouvert
2025
preprint
OpenAlex
Yassir Fathullah, Mark Gales
This paper explores generalised probabilistic modelling and uncertainty estimation in comparative LLM-as-a-judge frameworks. We show that existing Product-of-Experts methods are specific cases of a broader framework, enabling diverse modelling options. Furthermore, we propose improved uncertainty estimates for individual comparisons, enabling more efficient …
Accès ouvert
2025
conference-paper
OpenAlex
Rao Ma, Mengjie Qian, Yassir Fathullah, Siyuan Tang et autres
Rao Ma, Mengjie Qian, Yassir Fathullah, Siyuan Tang, Mark Gales, Kate Knill. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers). 2025.
gb
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Rao Ma, Mengjie Qian, Yassir Fathullah, Siyuan Tang et autres
There has been increasing interest in building multilingual foundation models for NLP and speech research. This paper examines how to expand the speech translation capability of these models with restricted data. Whisper, a speech foundation model with strong performance on speech recognition …
Accès ouvert
2024
preprint
OpenAlex
Rao Ma, Mengjie Qian, Yassir Fathullah, Siyuan Tang et autres
There has been increasing interest in building multilingual foundation models for NLP and speech research. This paper examines how to expand the speech translation capability of these models with restricted data. Whisper, a speech foundation model with strong performance on speech recognition …
Accès ouvert
2024
preprint
OpenAlex
Adian Liusie, Vatsal Raina, Yassir Fathullah, Mark Gales
LLM-as-a-judge approaches are a practical and effective way of assessing a range of text tasks. However, when using pairwise comparisons to rank a set of candidates, the computational cost scales quadratically with the number of candidates, which has practical limitations. This paper …
Accès ouvert
2024
preprint
OpenAlex
Yassir Fathullah, Mark Gales
Encoder-decoder foundation models have displayed state-of-the-art performance on a range of autoregressive sequence tasks. This paper proposes a simple and lightweight modification to such systems to control the behaviour according to a specific attribute of interest. This paper proposes a novel inference-efficient …
Accès ouvert
2024
preprint
OpenAlex
Adian Liusie, Yassir Fathullah, Mark Gales
Large Language Models (LLMs) have demonstrated impressive zero-shot capabilities and versatility in NLP tasks, however they sometimes fail to maintain crucial invariances for specific tasks. One example is permutation sensitivity, where LLMs' outputs may significantly vary depending on the order of the …
2024
conference-paper
OpenAlex
Egor Lakomkin, Chunyang Wu, Yassir Fathullah, Ozlem Kalinli et autres
In recent years, Large Language Models (LLMs) have garnered significant attention from the research community due to their exceptional performance and generalization capabilities. In this paper, we introduce a novel method for contextualizing speech recognition models incorporating LLMs. Our approach casts speech …
2024
conference-paper
OpenAlex
Yuan Shangguan, Haichuan Yang, Danni Li, Chunyang Wu et autres
Automatic Speech Recognition (ASR) models need to be optimized for specific hardware before they can be deployed on devices. This can be done by tuning the model’s hyperparameters or exploring variations in its architecture. Re-training and re-validating models after making these changes …
2024
conference-paper
OpenAlex
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia et autres
Large language models (LLMs) have proven themselves highly flexible, able to solve a wide range of generative tasks, such as abstractive summarization and open-ended question answering. In this paper we extend the capabilities of LLM by directly attaching a small audio encoder …
gb
(code pays fourni par la source)