Aller au contenu principal
Profil bibliographique

Muhammad Abdullah Jamal

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

27Publications signalées
346Citations signalées
3Affiliations récentes

Les institutions déclarées

Les domaines associés

Surgical Simulation and TrainingMultimodal Machine Learning ApplicationsDomain Adaptation and Few-Shot LearningAdvanced Neural Network ApplicationsGenerative Adversarial Networks and Image Synthesis

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models

John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu et autres

Vision foundation models (VFMs) are valuable in data-scarce domains such as surgery, where a single pretrained backbone can provide rich representations for many downstream tasks. Yet the dominant self-supervised pretraining paradigm uses only RGB images, leaving readily available complementary signals, such as …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning

Alejandra Beatriz Pérez, Anita Rau, Lee W. White, Busisiwe Mlambo et autres

Surgeons don't just see -- they interpret. When an expert observes a surgical scene, they understand not only what instrument is being used, but why it was chosen, what risk it poses, and what comes next. Current surgical AI cannot answer such …

us (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training

John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Chinedu Nwoye et autres

Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlooking the complex 3D geometry inherent to surgical environments. Although several architectures support multimodal or geometry-aware inputs in general computer …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training

John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Chinedu Nwoye et autres

Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlooking the complex 3D geometry inherent to surgical environments. Although several architectures support multimodal or geometry-aware inputs in general computer …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

Mitigating Surgical Data Imbalance with Dual-Prediction Video Diffusion Model

Danush Kumar Venkatesh, Adam Schmidt, Muhammad Abdullah Jamal, Omid Mohareri

Surgical video datasets are essential for scene understanding, enabling procedural modeling and intra-operative support. However, these datasets are often heavily imbalanced, with rare actions and tools under-represented, which limits the robustness of downstream models. We address this challenge with $SurgiFlowVid$, a sparse …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning

Alejandra Perez, Chinedu Innocent Nwoye, Ramtin Raji Kermani, Omid Mohareri et autres

Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without relying on expert-labeled datasets. However, progress in surgical VLP remains constrained by the limited scale, procedural diversity, semantic quality, and …

0 citations arXiv (Cornell University)
2025 conference-paper OpenAlex

Multi-Modal Contrastive Masked Autoencoders: A Two-Stage Progressive Pre-training Approach for RGBD Datasets

Muhammad Abdullah Jamal, Omid Mohareri

In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Autoencoder and Denoising techniques. Our proposed approach consists of two stages. In the first stage, we pre-train the …

us (code pays fourni par la source)

1 citation
Accès ouvert 2024 preprint OpenAlex

VidLPRO: A $\underline{Vid}$eo-$\underline{L}$anguage $\underline{P}$re-training Framework for $\underline{Ro}$botic and Laparoscopic Surgery

Mohammadmahdi Honarmand, Muhammad Abdullah Jamal, Omid Mohareri

We introduce VidLPRO, a novel video-language (VL) pre-training framework designed specifically for robotic and laparoscopic surgery. While existing surgical VL models primarily rely on contrastive learning, we propose a more comprehensive approach to capture the intricate temporal dynamics and align video with …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders

Muhammad Abdullah Jamal, Omid Mohareri

In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Autoencoder and Denoising techniques. Our proposed approach consists of two stages. In the first stage, we pre-train the …

0 citations arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.