Accès ouvert
2026
preprint
OpenAlex
Ofek I. Cohen, Lior Shani, Aviv Rosenberg, Ankur Samanta et autres
Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. However, such adaptation remains non-trivial: it often requires operationally challenging fine-tuning of open-source models or ad hoc …
Accès ouvert
2026
preprint
OpenAlex
Ofek I. Cohen, Lior Shani, Aviv Rosenberg, Ankur Samanta et autres
Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. However, such adaptation remains non-trivial: it often requires operationally challenging fine-tuning of open-source models or ad hoc …
us, il
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Asaf Cassel, Aviv Rosenberg
Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates to drive exploration. Although theoretically sound, such estimates are hard to compute in practical settings and therefore offer limited insight for designing exploration heuristics. Meanwhile, ensembling has emerged as …
Accès ouvert
2026
preprint
OpenAlex
Asaf Cassel, Aviv Rosenberg
Optimal Reinforcement Learning (RL) algorithms typically rely on carefully constructed count-based uncertainty estimates to drive exploration. Although theoretically sound, such estimates are hard to compute in practical settings and therefore offer limited insight for designing exploration heuristics. Meanwhile, ensembling has emerged as …
us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Orin Levy, Aviv Rosenberg, Alon Cohen, Yishay Mansour
We introduce \texttt{OPO-CMDP}, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. Our approach achieves a high probability regret bound of $\widetilde{O}(H^4\sqrt{T|S||A|\log(|\mathcal{F}||\mathcal{P}|)}),$ where $S$ and $A$ denote the state and action spaces, $H$ the …
Accès ouvert
2026
preprint
OpenAlex
Orin Levy, Aviv Rosenberg, Alon Cohen, Yishay Mansour
We introduce \texttt{OPO-CMDP}, the first policy optimization algorithm for stochastic Contextual Markov Decision Process (CMDPs) under general offline function approximation. Our approach achieves a high probability regret bound of $\widetilde{O}(H^4\sqrt{T|S||A|\log(|\mathcal{F}||\mathcal{P}|)}),$ where $S$ and $A$ denote the state and action spaces, $H$ the …
Accès ouvert
2024
preprint
OpenAlex
Orin Levy, Noam Touitou, Aviv Rosenberg
Online paging is a fundamental problem in the field of online algorithms, in which one maintains a cache of $k$ slots as requests for fetching pages arrive online. In the weighted variant of this problem, each page has its own fetching cost; …
Accès ouvert
2024
preprint
OpenAlex
Wei Xiong, Chengshuai Shi, Jiaming Shen, Aviv Rosenberg et autres
Recent studies have shown that large language models' (LLMs) mathematical problem-solving capabilities can be enhanced by integrating external tools, such as code interpreters, and employing multi-turn Chain-of-Thought (CoT) reasoning. While current methods focus on synthetic data generation and Supervised Fine-Tuning (SFT), this …
Accès ouvert
2024
preprint
OpenAlex
Asaf Cassel, Aviv Rosenberg
Policy Optimization (PO) methods are among the most popular Reinforcement Learning (RL) algorithms in practice. Recently, Sherman et al. [2023a] proposed a PO-based algorithm with rate-optimal regret guarantees under the linear Markov Decision Process (MDP) model. However, their algorithm relies on a …
Accès ouvert
2024
preprint
OpenAlex
Lior Shani, Aviv Rosenberg, Asaf Cassel, Oran Lang et autres
Reinforcement Learning from Human Feedback (RLHF) has become the standard approach for aligning Large Language Models (LLMs) with human preferences, allowing LLMs to demonstrate remarkable abilities in various tasks. Existing methods work by emulating the preferences at the single decision (turn) level, …
Accès ouvert
2024
preprint
OpenAlex
Asaf Cassel, Haipeng Luo, Aviv Rosenberg, Dmitry Sotnikov
In many real-world applications, it is hard to provide a reward signal in each step of a Reinforcement Learning (RL) process and more natural to give feedback when an episode ends. To this end, we study the recently proposed model of RL …
2024
conference-paper
OpenAlex
Lior Shani, Aviv Rosenberg, Asaf Cassel, Oran Lang et autres