Accès ouvert
2026
preprint
OpenAlex
Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald et autres
Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their implicit reward model and decrease the likelihood of preferred responses. This results in a …
Accès ouvert
2026
preprint
OpenAlex
Shawn Im, Federico Danieli, Skyler Seto, Barry-John Theobald et autres
Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs have been observed to over-optimize their implicit reward model and decrease the likelihood of preferred responses. This results in a …
Accès ouvert
2026
preprint
OpenAlex
Keitaro Sakamoto, Pierre Ablin, Federico Danieli, Marco Cuturi
The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software platforms. They also allow for more open research and testing, to the extent that users …
Accès ouvert
2026
preprint
OpenAlex
Keitaro Sakamoto, Pierre Ablin, Federico Danieli, Marco Cuturi
The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software platforms. They also allow for more open research and testing, to the extent that users …
Accès ouvert
2026
preprint
OpenAlex
Lucas Monteiro Paes, Natalie Mackraz, Barry-John Theobald, Federico Danieli
Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common alignment approaches are (i) reinforcement learning, which maximizes the expected reward under a KL-divergence constraint, and (ii) best-of-$N$ alignment, which …
Accès ouvert
2026
preprint
OpenAlex
Lucas Monteiro Paes, Natalie Mackraz, Barry-John Theobald, Federico Danieli
Language model (LM) alignment improves model outputs to reflect human preferences while preserving the capabilities of the base model. The most common alignment approaches are (i) reinforcement learning, which maximizes the expected reward under a KL-divergence constraint, and (ii) best-of-$N$ alignment, which …
Accès ouvert
2026
preprint
OpenAlex
Abhinav Moudgil, Ningyuan Huang, Eeshan Gunesh Dhekane, Pau Rodríguez et autres
State Space Models (SSMs) such as Mamba have become a popular alternative to Transformer models, due to their reduced memory consumption and higher throughput at generation compared to their Attention-based counterparts. On the other hand, the community has built up a considerable …
Accès ouvert
2026
preprint
OpenAlex
Abhinav Moudgil, Ningyuan Huang, Eeshan Gunesh Dhekane, Pau Rodríguez et autres
State Space Models (SSMs) such as Mamba have become a popular alternative to Transformer models, due to their reduced memory consumption and higher throughput at generation compared to their Attention-based counterparts. On the other hand, the community has built up a considerable …
us, Algérie
(code pays fourni par la source)
2026
article
OpenAlex
Ningyuan Huang, Federico Danieli
us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Emily Cheng, Carmen Amo Alonso, Federico Danieli, Arno Blaas et autres
As generative models become ubiquitous, there is a critical need for fine-grained control over the generation process. Yet, while controlled generation methods from prompting to fine-tuning proliferate, a fundamental question remains unanswered: are these models truly controllable in the first place? In …
Accès ouvert
2026
preprint
OpenAlex
Emily Cheng, Carmen Amo Alonso, Federico Danieli, Arno Blaas et autres
As generative models become ubiquitous, there is a critical need for fine-grained control over the generation process. Yet, while controlled generation methods from prompting to fine-tuning proliferate, a fundamental question remains unanswered: are these models truly controllable in the first place? In …
2025
article
OpenAlex
Federico Danieli, Ben S. Southworth, Jacob B. Schroder
ABSTRACT This work develops an all‐at‐once space‐time preconditioning approach for resistive magnetohydrodynamics (MHD). We consider parallel‐in‐time due to the long time domains required to capture the physics of interest, as well as the complexity of the underlying system, and thereby the computational …
gb, us
(code pays fourni par la source)