Aller au contenu principal
Accès ouvert déclaré 2025 conference-paper

Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations

3Citations signalées, ce qui n’est pas une note de qualité
2Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : se. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

LLMs have been explored for their use in IR as end-to-end rankers, rerankers and assessors. Recently, the exploration of the prompt-and-predict paradigm for reranking in combination with highly performant LLMs have drawn the attention of researchers. Instead of training or fine-tuning a reranker, LLMs are prompted in a zero-shot manner to produce relevance scores, pairwise preferences, or reranked lists. Existing research, though, has been confined to unstructured text corpora, leaving a gap in our understanding: to what extent do the findings of zero-shot LLM rerankers established on plain text corpora hold for datasets containing predominantly precomputed ranking features as is common in industrial settings? We explore this question via an empirical study on one public learning-to-rank dataset (MSLR-WEB10K) and two datasets collected from an audio streaming platform's search logs. Our results paint a differentiated picture: On average, there remains a significant performance gap: prompting the high-capacity LLM GPT-4 results in up to 16% lower NDCG@10 compared to the traditional supervised learning-to-rank (LTR) approach LambdaMART on the public MSLR-WEB10K dataset. However, when focusing only on a subset of hard queries-i.e. queries where the LTR approach ranks a non-relevant document at the top-the zero-shot LLM reranking outperforms the LTR baseline. We confirm the same trends on two proprietary audio search datasets. We also provide insights into prompt design choices and their impact on LLM reranking. We show that LLMs remain brittle, with the same strategies sometimes helping or hurting depending on the model size and dataset.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Zero-Shot Reranking with Large Language Models and Precomputed Ranking Features: Opportunities and Limitations
Date Crossref
13/07/2025
Éditeur
ACM
Type
proceedings-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Stockholm University Sweden and Department of Computer and Systems Sciences pays non établi dans la notice
    Université ou école supérieure
  • Spotify (Sweden) pays non établi dans la notice
    Entreprise
  • Spotify AB pays non établi dans la notice
    Institution

Sweden and Department of Computer and Systems Sciences — Stockholm University, Spotify (Sweden) et Spotify AB.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Topic ModelingNatural Language Processing TechniquesAdvanced Graph Neural Networks

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.