Accès ouvert
2026
preprint
OpenAlex
Gustavo Penha, Juan Elenter, Claudia Hauff, Hugues Bouchard et autres
Recent work has focused on improving explicit natural-language descriptive reasoning traces for generative recommendation. This includes systems that augment semantic ID (SID) prediction with chain-of-thought reasoning. However, because SIDs are opaque learned identifiers rather than natural language, they require costly alignment before …
Accès ouvert
2026
conference-paper
OpenAlex
Ali Vardasbi, Gustavo Penha, Enrico Palumbo, Claudia Hauff et autres
Large-scale search systems evolve faster than human quality assurance scales, especially for long-tail intents and multilingual queries. LLM-as-a-judge approaches are a scalable alternative for evaluating the relevance of search engine result pages (SERPs), but judgments based solely on semantic similarity or world …
nl
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Ali Vardasbi, Gustavo Penha, Enrico Palumbo, Claudia Hauff et autres
Large-scale search systems evolve faster than human quality assurance can scale, especially for long-tail intents and multilingual queries. LLM-as-a-judge approaches provide a scalable alternative for evaluating the relevance of search engine result pages (SERPs), but judgments based solely on semantic similarity or …
Accès ouvert
2026
preprint
OpenAlex
Ali Vardasbi, Gustavo Penha, Enrico Palumbo, Claudia Hauff et autres
Large-scale search systems evolve faster than human quality assurance can scale, especially for long-tail intents and multilingual queries. LLM-as-a-judge approaches provide a scalable alternative for evaluating the relevance of search engine result pages (SERPs), but judgments based solely on semantic similarity or …
2026
article
OpenAlex
Leif Azzopardi, Charles L. A. Clarke, Claudia Hauff, Yubin Kim et autres
The Third Search Futures Workshop [Azzopardi et al., 2026], in conjunction with the Forty-eight European Conference on Information Retrieval (ECIR) 2026, looked into the future of search to ask questions such as: • How can we navigate data privacy in large language …
gb, ca, nl, us, au, de, cn, fi
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Maria Movin, Claudia Hauff, Aron Henriksson, Panagiotis Papapetrou
LLM-driven GUI agents are increasingly used in production systems to automate workflows and simulate users for evaluation and optimization. Yet most GUI-agent evaluations emphasize task success and provide limited evidence on whether agents interact in human-like ways. We present a trace-level evaluation …
Accès ouvert
2026
preprint
OpenAlex
Maria Movin, Claudia Hauff, Aron Henriksson, Panagiotis Papapetrou
LLM-driven GUI agents are increasingly used in production systems to automate workflows and simulate users for evaluation and optimization. Yet most GUI-agent evaluations emphasize task success and provide limited evidence on whether agents interact in human-like ways. We present a trace-level evaluation …
se
(code pays fourni par la source)
2026
conference-paper
OpenAlex
Ali Vardasbi, Gustavo Penha, Claudia Hauff, Hugues Bouchard
nl
(code pays fourni par la source)
2026
conference-paper
OpenAlex
Leif A. Azzopardi, Charles L. A. Clarke, Claudia Hauff, Yubin Kim et autres
gb, ca, nl, au, de
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Ali Vardasbi, Gustavo Penha, Claudia Hauff, Hugues Bouchard
When using LLMs to rank items based on given criteria, or evaluate answers, the order of candidate items can influence the model's final decision. This sensitivity to item positioning in a LLM's prompt is known as position bias. Prior research shows that …
Accès ouvert
2025
conference-paper
OpenAlex
Gianluca Demartini, Claudia Hauff, Matthew Lease, Stefano Mizzaro et autres
Contains fulltext : 321417.pdf (Publisher’s version ) (Open Access)
au, nl, us, it
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Maria Movin, Claudia Hauff
LLMs have been explored for their use in IR as end-to-end rankers, rerankers and assessors. Recently, the exploration of the prompt-and-predict paradigm for reranking in combination with highly performant LLMs have drawn the attention of researchers. Instead of training or fine-tuning a …
se
(code pays fourni par la source)