Accès ouvert
2026
preprint
OpenAlex
Federico Spurio, Olga Zatsarynna, Lars Doorenbos, Emad Bahrami et autres
Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecting mistakes in videos, with current methods mostly focusing on closed-set protocols. While successful in controlled …
Accès ouvert
2026
article
OpenAlex
Federico Spurio, Emad Bahrami, Olga Zatsarynna, Yazan Abu Farha et autres
We introduce Action Discovery, a novel task that addresses the challenge of discovering actions in long, untrimmed videos where only a subset of the present actions have been annotated. The goalis thus to discover new actions in the video segments that have …
de, ps, be
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Emad Bahrami, Olga Zatsarynna, Parth Pathak, Sunando Sengupta et autres
We introduce STRIVE (SpatioTemporal Reinforcement with Importance-aware Variant Exploration), a structured reinforcement learning framework for video question answering. While group-based policy optimization methods have shown promise in large multimodal models, they often suffer from low reward variance when responses exhibit similar correctness, …
Accès ouvert
2026
article
OpenAlex
Emad Bahrami, Olga Zatsarynna, Gianpiero Francesca, Jüergen Gall
Abstract While there has been substantial progress in temporal action segmentation, the challenge to generalize to unseen views remains unaddressed. Hence, we define a protocol for unseen view action segmentation where camera views for evaluating the model are unavailable during training. This …
de, be
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Federico Spurio, Emad Bahrami, Gianpiero Francesca, Jüergen Gall
In this work, we address unsupervised temporal action segmentation, which segments a set of long, untrimmed videos into semantically meaningful segments that are consistent across videos. While recent approaches combine representation learning and clustering in a single step for this task, they …
us, ch
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Olga Zatsarynna, Emad Bahrami, Yazan Abu Farha, Gianpiero Francesca et autres
Long-term dense action anticipation is very challenging since it requires predicting actions and their durations several minutes into the future based on provided video observations. To model the uncertainty of future outcomes, stochastic models predict several potential future action sequences for the …
Accès ouvert
2024
preprint
OpenAlex
Federico Spurio, Emad Bahrami, Gianpiero Francesca, Jüergen Gall
In this work, we address unsupervised temporal action segmentation, which segments a set of long, untrimmed videos into semantically meaningful segments that are consistent across videos. While recent approaches combine representation learning and clustering in a single step for this task, they …
2024
conference-paper
OpenAlex
Olga Zatsarynna, Emad Bahrami, Yazan Abu Farha, Gianpiero Francesca et autres
de, ps, be
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Olga Zatsarynna, Emad Bahrami, Yazan Abu Farha, Gianpiero Francesca et autres
Long-term action anticipation has become an important task for many applications such as autonomous driving and human-robot interaction. Unlike short-term anticipation, predicting more actions into the future imposes a real challenge with the increasing uncertainty in longer horizons. While there has been …
2023
conference-paper
OpenAlex
Emad Bahrami, Gianpiero Francesca, Jüergen Gall
Modeling long-term context in videos is crucial for many fine-grained tasks including temporal action segmentation. An interesting question that is still open is how much long-term temporal context is needed for optimal performance. While transformers can model the long-term context of a …
de, be
(code pays fourni par la source)
Accès ouvert
2023
preprint
OpenAlex
Emad Bahrami, Gianpiero Francesca, Jüergen Gall
Modeling long-term context in videos is crucial for many fine-grained tasks including temporal action segmentation. An interesting question that is still open is how much long-term temporal context is needed for optimal performance. While transformers can model the long-term context of a …
Accès ouvert
2022
preprint
OpenAlex
Yaser Souri, Yazan Abu Farha, Emad Bahrami, Gianpiero Francesca et autres
Action segmentation is the task of predicting an action label for each frame of an untrimmed video. As obtaining annotations to train an approach for action segmentation in a fully supervised way is expensive, various approaches have been proposed to train action …