Accès ouvert
2026
preprint
OpenAlex
Xavier Suau, Alex Ferrando de las Morenas, Luca Zappella, Samy Bengio
When language models reason in chain-of-thought or exchange free-text intermediates, they serialize structured information into natural language. How much tree-structured compositional content survives this bottleneck? We propose a round-trip protocol that answers this question empirically for tree-structured expressions. A generator converts a …
Accès ouvert
2026
preprint
OpenAlex
Iuri Macocco, Pau Rodríguez, Arno Blaas, Luca Zappella et autres
Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. Current approaches to conditioning are often evaluated with a narrow focus on their effectiveness at injecting …
Accès ouvert
2026
preprint
OpenAlex
Iuri Macocco, Pau Rodríguez, Arno Blaas, Luca Zappella et autres
Controlling the output of Large Language Models (LLMs) is a central challenge for their reliable deployment, yet a clear understanding of the involved trade-offs remains elusive. Current approaches to conditioning are often evaluated with a narrow focus on their effectiveness at injecting …
es, gb
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Valentino Maiorca, Eleonora Gualdoni, Xavier Suau, Marco Cuturi et autres
As foundation models grow in capability, the ability to efficiently and reliably control their behavior becomes critical. Fine-tuning these models can be costly, and while prompting can be practical for controllability, it remains fragile due to models' high sensitivity to exact prompt …
Accès ouvert
2026
preprint
OpenAlex
Valentino Maiorca, Eleonora Gualdoni, Xavier Suau, Marco Cuturi et autres
As foundation models grow in capability, the ability to efficiently and reliably control their behavior becomes critical. Fine-tuning these models can be costly, and while prompting can be practical for controllability, it remains fragile due to models' high sensitivity to exact prompt …
Accès ouvert
2026
preprint
OpenAlex
Zihuiwen Ye, Lukas Aichberger, Michael Kirchhof, Sinead Williamson et autres
Large Language Models (LLMs) are increasingly deployed to autonomously solve real-world tasks. A key ingredient for this is the LLM Function-Calling paradigm, a widely used approach for equipping LLMs with tool-use capabilities. However, an LLM calling functions incorrectly can have severe implications, …
Accès ouvert
2026
preprint
OpenAlex
Zihuiwen Ye, Lukas Aichberger, Michael Kirchhof, Sinead Williamson et autres
Large Language Models (LLMs) are increasingly deployed to autonomously solve real-world tasks. A key ingredient for this is the LLM Function-Calling paradigm, a widely used approach for equipping LLMs with tool-use capabilities. However, an LLM calling functions incorrectly can have severe implications, …
gb
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Abhinav Moudgil, Ningyuan Teresa Huang, Eeshan Gunesh Dhekane, Pau Riera Rodriguez et autres
State Space Models (SSMs) such as Mamba have become a popular alternative to Transformer models, due to their reduced memory consumption and higher throughput at generation compared to their Attention-based counterparts. On the other hand, the community has built up a considerable …
Accès ouvert
2026
preprint
OpenAlex
Abhinav Moudgil, Ningyuan Teresa Huang, Eeshan Gunesh Dhekane, Pau Riera Rodriguez et autres
State Space Models (SSMs) such as Mamba have become a popular alternative to Transformer models, due to their reduced memory consumption and higher throughput at generation compared to their Attention-based counterparts. On the other hand, the community has built up a considerable …
us, Algérie
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Louis Bethune, Victor Turrisi, Bruno Mlodozeniec, Pau Rodriguez Lopez et autres
Discrete diffusion models have emerged as strong alternatives to autoregressive language models, with recent work initializing and fine-tuning a base unimodal model for bimodal generation. Diverging from previous approaches, we introduce the first tri-modal masked diffusion model pretrained from scratch on text, …
Accès ouvert
2026
preprint
OpenAlex
Louis Bethune, Victor Turrisi, Bruno Mlodozeniec, Pau Rodriguez Lopez et autres
Discrete diffusion models have emerged as strong alternatives to autoregressive language models, with recent work initializing and fine-tuning a base unimodal model for bimodal generation. Diverging from previous approaches, we introduce the first tri-modal masked diffusion model pretrained from scratch on text, …
Accès ouvert
2026
preprint
OpenAlex
Emily Cheng, Carmen Amo Alonso, Federico Danieli, Arno Blaas et autres
As generative models become ubiquitous, there is a critical need for fine-grained control over the generation process. Yet, while controlled generation methods from prompting to fine-tuning proliferate, a fundamental question remains unanswered: are these models truly controllable in the first place? In …