Aller au contenu principal
Profil bibliographique

Aviv Rosenberg

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

38Publications signalées
149Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

Reinforcement Learning in RoboticsAdvanced Bandit Algorithms ResearchMachine Learning and AlgorithmsAdversarial Robustness in Machine LearningExplainable Artificial Intelligence (XAI)

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias

Ofek I. Cohen, Lior Shani, Aviv Rosenberg, Ankur Samanta et autres

Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. However, such adaptation remains non-trivial: it often requires operationally challenging fine-tuning of open-source models or ad hoc …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias

Ofek I. Cohen, Lior Shani, Aviv Rosenberg, Ankur Samanta et autres

Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. However, such adaptation remains non-trivial: it often requires operationally challenging fine-tuning of open-source models or ad hoc …

us, il (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Quantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning

Asaf Cassel, Aviv Rosenberg

Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates to drive exploration. Although theoretically sound, such estimates are hard to compute in practical settings and therefore offer limited insight for designing exploration heuristics. Meanwhile, ensembling has emerged as …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Quantile of Means: A Bonus-Free Ensemble Method for Minimax Optimal Reinforcement Learning

Asaf Cassel, Aviv Rosenberg

Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates to drive exploration. Although theoretically sound, such estimates are hard to compute in practical settings and therefore offer limited insight for designing exploration heuristics. Meanwhile, ensembling has emerged as …

us (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation

Orin Levy, Aviv Rosenberg, Alon Cohen, Yishay Mansour

We introduce \texttt{OPO-CMDP}, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. Our approach achieves a high probability regret bound of $\widetilde{O}(H^4\sqrt{T|S||A|\log(|\mathcal{F}||\mathcal{P}|)}),$ where $S$ and $A$ denote the state and action spaces, $H$ the …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Near-Optimal Regret for Policy Optimization in Contextual MDPs with General Offline Function Approximation

Orin Levy, Aviv Rosenberg, Alon Cohen, Yishay Mansour

We introduce \texttt{OPO-CMDP}, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. Our approach achieves a high probability regret bound of $\widetilde{O}(H^4\sqrt{T|S||A|\log(|\mathcal{F}||\mathcal{P}|)}),$ where $S$ and $A$ denote the state and action spaces, $H$ the …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

Online Weighted Paging with Unknown Weights

Orin Levy, Noam Touitou, Aviv Rosenberg

Online paging is a fundamental problem in the field of online algorithms, in which one maintains a cache of $k$ slots as requests for fetching pages arrive online. In the weighted variant of this problem, each page has its own fetching cost; …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

Building Math Agents with Multi-Turn Iterative Preference Learning

Wei Xiong, Chengshuai Shi, Jiaming Shen, Aviv Rosenberg et autres

Recent studies have shown that large language models' (LLMs) mathematical problem-solving capabilities can be enhanced by integrating external tools, such as code interpreters, and employing multi-turn Chain-of-Thought (CoT) reasoning. While current methods focus on synthetic data generation and Supervised Fine-Tuning (SFT), this …

1 citation arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

Multi-turn Reinforcement Learning from Preference Human Feedback

Lior Shani, Aviv Rosenberg, Asaf Cassel, Oran Lang et autres

Reinforcement Learning from Human Feedback (RLHF) has become the standard approach for aligning Large Language Models (LLMs) with human preferences, allowing LLMs to demonstrate remarkable abilities in various tasks. Existing methods work by emulating the preferences at the single decision (turn) level, …

1 citation arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.