Accès ouvert
2026
preprint
OpenAlex
Keitaro Sakamoto, Pierre Ablin, Federico Danieli, Marco Cuturi
The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software platforms. They also allow for more open research and testing, to the extent that users …
Accès ouvert
2026
preprint
OpenAlex
Eleonora Gualdoni, Sonia Laguna, Louis Bethune, Joao Monteiro et autres
Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving performance on constrained domains, such as general knowledge, instruction following, or safety evaluations. Existing data mixing strategies rely on fixed heuristics or adaptive rules that cannot explicitly enforce …
Accès ouvert
2026
preprint
OpenAlex
Eleonora Gualdoni, Sonia Laguna, Louis Béthune, João Monteiro et autres
Multi-domain fine-tuning of large language models requires improving performance on target domains while preserving performance on constrained domains, such as general knowledge, instruction following, or safety evaluations. Existing data mixing strategies rely on fixed heuristics or adaptive rules that cannot explicitly enforce …
Accès ouvert
2026
preprint
OpenAlex
Keitaro Sakamoto, Pierre Ablin, Federico Danieli, Marco Cuturi
The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software platforms. They also allow for more open research and testing, to the extent that users …
Accès ouvert
2026
preprint
OpenAlex
Joao Monteiro, Michal Klein, Pierre Ablin, Marco Cuturi
Evaluating softmax attention over a fixed long context requires reading every cached key-value pair for each new query token. For a given context (a book, a manual, a legal corpus) the attention output is a deterministic function of the query. We propose …
Accès ouvert
2026
preprint
OpenAlex
João Monteiro, Michal Klein, Pierre Ablin, Marco Cuturi
Evaluating softmax attention over a fixed long context requires reading every cached key-value pair for each new query token. For a given context (a book, a manual, a legal corpus) the attention output is a deterministic function of the query. We propose …
Accès ouvert
2026
preprint
OpenAlex
Oscar Davis, Anastasiia Filippova, Pierre Ablin, Victor Turrisi et autres
Continuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they unlock a host of advantages currently reserved for continuous modalities, including accelerated sampling and tilting. Recently, several works have demonstrated the possibility …
Accès ouvert
2026
preprint
OpenAlex
Oscar Davis, Anastasiia Filippova, Pierre Ablin, Victor Turrisi et autres
Continuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they unlock a host of advantages currently reserved for continuous modalities, including accelerated sampling and tilting. Recently, several works have demonstrated the possibility …
il, gb
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Valentino Maiorca, Eleonora Gualdoni, Xavier Suau, Marco Cuturi et autres
As foundation models grow in capability, the ability to efficiently and reliably control their behavior becomes critical. Fine-tuning these models can be costly, and while prompting can be practical for controllability, it remains fragile due to models' high sensitivity to exact prompt …
Accès ouvert
2026
preprint
OpenAlex
Valentino Maiorca, Eleonora Gualdoni, Xavier Suau, Marco Cuturi et autres
As foundation models grow in capability, the ability to efficiently and reliably control their behavior becomes critical. Fine-tuning these models can be costly, and while prompting can be practical for controllability, it remains fragile due to models' high sensitivity to exact prompt …
Accès ouvert
2026
preprint
OpenAlex
Anastasiia Filippova, David Grangier, Marco Cuturi, João Monteiro
Serving transformer language models with high throughput requires caching Key-Values (KVs) to avoid redundant computation during autoregressive generation. The memory footprint of KV caching is significant and heavily impacts serving costs. This work proposes to lessen these memory requirements. While recent work …
Accès ouvert
2026
preprint
OpenAlex
Anastasiia Filippova, David Grangier, Marco Cuturi, João Monteiro
Serving transformer language models with high throughput requires caching Key-Values (KVs) to avoid redundant computation during autoregressive generation. The memory footprint of KV caching is significant and heavily impacts serving costs. This work proposes to lessen these memory requirements. While recent work …