Accès ouvert
2026
preprint
OpenAlex
Ariel Shaulov, Eitan Shaar, Amit Edenzon, Gal Chechik et autres
Text-to-video diffusion models generate realistic videos, but often fail on prompts requiring fine-grained compositional understanding, such as relations between entities, attributes, actions, and motion directions. We hypothesize that these failures need not be addressed by retraining the generator, but can instead be …
Accès ouvert
2026
preprint
OpenAlex
Ariel Shaulov, Eitan Shaar, Amit Edenzon, Gal Chechik et autres
Text-to-video diffusion models generate realistic videos, but often fail on prompts requiring fine-grained compositional understanding, such as relations between entities, attributes, actions, and motion directions. We hypothesize that these failures need not be addressed by retraining the generator, but can instead be …
il, gb
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Eitan Shaar, Ariel Shaulov, Yalcin Tur, Gal Chechik et autres
Adversarial attacks are a central tool for probing the robustness of modern vision models, yet most methods optimize perturbations directly in pixel space under $\ell_\infty$ or $\ell_2$ constraints. While effective in white-box settings, pixel-space optimization often produces high-frequency, texture-like noise that is …
Accès ouvert
2026
preprint
OpenAlex
Eitan Shaar, Ariel Shaulov, Yalcin Tur, Gal Chechik et autres
Adversarial attacks are a central tool for probing the robustness of modern vision models, yet most methods optimize perturbations directly in pixel space under $\ell_\infty$ or $\ell_2$ constraints. While effective in white-box settings, pixel-space optimization often produces high-frequency, texture-like noise that is …
il, us, gb
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Ariel Shaulov, Eitan Shaar, Amit Edenzon, Lior Wolf
Auto-regressive video generation enables long video synthesis by iteratively conditioning each new batch of frames on previously generated content. However, recent work has shown that such pipelines suffer from severe temporal drift, where errors accumulate and amplify over long horizons. We hypothesize …
Accès ouvert
2026
preprint
OpenAlex
Ariel Shaulov, Eitan Shaar, Amit Edenzon, Lior Wolf
Auto-regressive video generation enables long video synthesis by iteratively conditioning each new batch of frames on previously generated content. However, recent work has shown that such pipelines suffer from severe temporal drift, where errors accumulate and amplify over long horizons. We hypothesize …
2025
conference-paper
OpenAlex
Eitan Shaar, Ariel Shaulov, Gal Chechik, Lior Wolf
In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing approaches are constrained by the vocabulary available in their training data. This limitation significantly impedes their capacity to …
il
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Ariel Shaulov, Tal Shaharabany, Eitan Shaar, Gal Chechik et autres
Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions and contexts. This limitation hinders performance in tasks like audio or video …
il
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Aviv Shamsian, Eitan Shaar, Aviv Navon, Gal Chechik et autres
Machine unlearning aims to remove the influence of problematic training data after a model has been trained. The primary challenge in machine unlearning is ensuring that the process effectively removes specified data without compromising the model's overall performance on the remaining dataset. …
Accès ouvert
2025
preprint
OpenAlex
Ariel Shaulov, Tal Shaharabany, Eitan Shaar, Gal Chechik et autres
Most current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions and contexts. This limitation hinders performance in tasks like audio or video …
Accès ouvert
2023
preprint
OpenAlex
Lior Bracha, Eitan Shaar, Aviv Shamsian, Ethan Fetaya et autres
Referring Expressions Generation (REG) aims to produce textual descriptions that unambiguously identifies specific objects within a visual scene. Traditionally, this has been achieved through supervised learning methods, which perform well on specific data distributions but often struggle to generalize to new images …