Accès ouvert
2026
preprint
OpenAlex
Bonan Ding, Umair Nawaz, Ufaq Khan, Abdelrahman Shaker et autres
Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evidence. In such a scenario, uniform fusion induces modality interference, allowing irrelevant channels to distract the …
Accès ouvert
2026
preprint
OpenAlex
Bonan Ding, Umair Nawaz, Ufaq Khan, Abdelrahman Shaker et autres
Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evidence. In such a scenario, uniform fusion induces modality interference, allowing irrelevant channels to distract the …
cn, ca, se
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Ahmed Heakl, Abdelrahman Shaker, Youssef Mohamed, Rania Elbadry et autres
When a model produces a correct solution under reinforcement learning with verifiable rewards (RLVR), every token receives the same reward signal regardless of whether it was a decisive reasoning step or a grammatical filler. A natural fix is to condition the model …
Accès ouvert
2026
preprint
OpenAlex
Ahmed Heakl, Abdelrahman Shaker, Youssef Mohamed, Rania Elbadry et autres
When a model produces a correct solution under reinforcement learning with verifiable rewards (RLVR), every token receives the same reward signal regardless of whether it was a decisive reasoning step or a grammatical filler. A natural fix is to condition the model …
se, au
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Umair Nawaz, Ahmed Heakl, Ufaq Khan, Abdelrahman Shaker et autres
Diffusion Transformers (DiTs) power high-fidelity video world models but remain computationally expensive due to sequential denoising and costly spatio-temporal attention. Training-free feature caching accelerates inference by reusing intermediate activations across denoising steps; however, existing methods largely rely on a Zero-Order Hold assumption …
Accès ouvert
2026
preprint
OpenAlex
Abdelrahman Shaker, Ahmed Heakl, Jaseel Muhammad, Ritesh Thawkar et autres
Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on edge devices. We present Mobile-O, a compact vision-language-diffusion model that brings unified multimodal intelligence to a mobile …
Accès ouvert
2026
preprint
OpenAlex
Abdelrahman Shaker, Ahmed Heakl, Jaseel Muhammad, Ritesh Thawkar et autres
Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on edge devices. We present Mobile-O, a compact vision-language-diffusion model that brings unified multimodal intelligence to a mobile …
Accès ouvert
2025
preprint
OpenAlex
Omkar Thawakar, Shravan Venkatraman, Ritesh Thawkar, Abdelrahman Shaker et autres
Recent advances in large multimodal models (LMMs) have enabled impressive reasoning and perception abilities, yet most existing training pipelines still depend on human-curated data or externally verified reward models, limiting their autonomy and scalability. In this work, we strive to improve LMM …
Accès ouvert
2025
preprint
OpenAlex
Tajamul Ashraf, Umair Nawaz, Abdelrahman Shaker, Rao Muhammad Anwer et autres
Vision language models (VLMs) are increasingly deployed as controllers with access to external tools for complex reasoning and decision-making, yet their effectiveness remains limited by the scarcity of high-quality multimodal trajectories and the cost of manual annotation. We address this challenge with …
2025
conference-paper
OpenAlex
Abdelrahman Shaker, Syed Talal Wasim, Salman A. Khan, Jüergen Gall et autres
State-Space models (SSMs) have recently shown promise in capturing long-range dependencies with subquadratic computational complexity, making them attractive for various applications. However, purely SSM-based models face critical challenges related to stability and achieving state-of-the-art performance in computer vision tasks. Our paper addresses …
ae, de
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Hanoona Rasheed, Abdelrahman Shaker, Muhammad Maaz, Ming–Hsuan Yang et autres
Mathematical reasoning in real-world video settings presents a fundamentally different challenge than in static images or text. It requires interpreting fine-grained visual information, accurately reading handwritten or digital text, and integrating spoken cues, often dispersed non-linearly over time. In such multimodal contexts, …
Accès ouvert
2025
preprint
OpenAlex
Ghazi Shazan Ahmad, Ahmed Heakl, Hanan Gani, Abdelrahman Shaker et autres
Spatio-temporal localization is vital for precise interactions across diverse domains, from biological research to autonomous navigation and interactive interfaces. Current video-based approaches, while proficient in tracking, lack the sophisticated reasoning capabilities of large language models, limiting their contextual understanding and generalization. We …