Aller au contenu principal
Profil bibliographique

Abdelrahman Shaker

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

36Publications signalées
1740Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

Multimodal Machine Learning ApplicationsAdvanced Neural Network ApplicationsDomain Adaptation and Few-Shot LearningAdvanced Image and Video Retrieval TechniquesTopic Modeling

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos

Bonan Ding, Umair Nawaz, Ufaq Khan, Abdelrahman Shaker et autres

Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evidence. In such a scenario, uniform fusion induces modality interference, allowing irrelevant channels to distract the …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos

Bonan Ding, Umair Nawaz, Ufaq Khan, Abdelrahman Shaker et autres

Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, depth map, or dense temporal evidence. In such a scenario, uniform fusion induces modality interference, allowing irrelevant channels to distract the …

cn, ca, se (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

Ahmed Heakl, Abdelrahman Shaker, Youssef Mohamed, Rania Elbadry et autres

When a model produces a correct solution under reinforcement learning with verifiable rewards (RLVR), every token receives the same reward signal regardless of whether it was a decisive reasoning step or a grammatical filler. A natural fix is to condition the model …

se, au (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

WorldCache: Content-Aware Caching for Accelerated Video World Models

Umair Nawaz, Ahmed Heakl, Ufaq Khan, Abdelrahman Shaker et autres

Diffusion Transformers (DiTs) power high-fidelity video world models but remain computationally expensive due to sequential denoising and costly spatio-temporal attention. Training-free feature caching accelerates inference by reusing intermediate activations across denoising steps; however, existing methods largely rely on a Zero-Order Hold assumption …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device

Abdelrahman Shaker, Ahmed Heakl, Jaseel Muhammad, Ritesh Thawkar et autres

Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on edge devices. We present Mobile-O, a compact vision-language-diffusion model that brings unified multimodal intelligence to a mobile …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device

Abdelrahman Shaker, Ahmed Heakl, Jaseel Muhammad, Ritesh Thawkar et autres

Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on edge devices. We present Mobile-O, a compact vision-language-diffusion model that brings unified multimodal intelligence to a mobile …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards

Omkar Thawakar, Shravan Venkatraman, Ritesh Thawkar, Abdelrahman Shaker et autres

Recent advances in large multimodal models (LMMs) have enabled impressive reasoning and perception abilities, yet most existing training pipelines still depend on human-curated data or externally verified reward models, limiting their autonomy and scalability. In this work, we strive to improve LMM …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning

Tajamul Ashraf, Umair Nawaz, Abdelrahman Shaker, Rao Muhammad Anwer et autres

Vision language models (VLMs) are increasingly deployed as controllers with access to external tools for complex reasoning and decision-making, yet their effectiveness remains limited by the scarcity of high-quality multimodal trajectories and the cost of manual annotation. We address this challenge with …

0 citations arXiv (Cornell University)
2025 conference-paper OpenAlex

GroupMamba: Efficient Group-Based Visual State Space Model

Abdelrahman Shaker, Syed Talal Wasim, Salman A. Khan, Jüergen Gall et autres

State-Space models (SSMs) have recently shown promise in capturing long-range dependencies with subquadratic computational complexity, making them attractive for various applications. However, purely SSM-based models face critical challenges related to stability and achieving state-of-the-art performance in computer vision tasks. Our paper addresses …

ae, de (code pays fourni par la source)

22 citations
Accès ouvert 2025 preprint OpenAlex

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos

Hanoona Rasheed, Abdelrahman Shaker, Muhammad Maaz, Ming–Hsuan Yang et autres

Mathematical reasoning in real-world video settings presents a fundamentally different challenge than in static images or text. It requires interpreting fine-grained visual information, accurately reading handwritten or digital text, and integrating spoken cues, often dispersed non-linearly over time. In such multimodal contexts, …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

VideoMolmo: Spatio-Temporal Grounding Meets Pointing

Ghazi Shazan Ahmad, Ahmed Heakl, Hanan Gani, Abdelrahman Shaker et autres

Spatio-temporal localization is vital for precise interactions across diverse domains, from biological research to autonomous navigation and interactive interfaces. Current video-based approaches, while proficient in tracking, lack the sophisticated reasoning capabilities of large language models, limiting their contextual understanding and generalization. We …

0 citations arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.