Accès ouvert
2026
preprint
OpenAlex
John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu et autres
Vision foundation models (VFMs) are valuable in data-scarce domains such as surgery, where a single pretrained backbone can provide rich representations for many downstream tasks. Yet the dominant self-supervised pretraining paradigm uses only RGB images, leaving readily available complementary signals, such as …
Accès ouvert
2026
article
OpenAlex
Idris Hamoud, Vinkle Srivastav, Muhammad Abdullah Jamal, Didier Mutter et autres
fr, in, us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Alejandra Beatriz Pérez, Anita Rau, Lee W. White, Busisiwe Mlambo et autres
Surgeons don't just see -- they interpret. When an expert observes a surgical scene, they understand not only what instrument is being used, but why it was chosen, what risk it poses, and what comes next. Current surgical AI cannot answer such …
Accès ouvert
2026
preprint
OpenAlex
Alejandra Beatriz Pérez, Anita Rau, Lee W. White, Busisiwe Mlambo et autres
Surgeons don't just see -- they interpret. When an expert observes a surgical scene, they understand not only what instrument is being used, but why it was chosen, what risk it poses, and what comes next. Current surgical AI cannot answer such …
us
(code pays fourni par la source)
2026
article
OpenAlex
Alejandra Perez, Chinedu Nwoye, Ramtin Raji Kermani, Omid Mohareri et autres
co, ca, us, ch
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Chinedu Nwoye et autres
Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlooking the complex 3D geometry inherent to surgical environments. Although several architectures support multimodal or geometry-aware inputs in general computer …
Accès ouvert
2026
preprint
OpenAlex
John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Chinedu Nwoye et autres
Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlooking the complex 3D geometry inherent to surgical environments. Although several architectures support multimodal or geometry-aware inputs in general computer …
Accès ouvert
2025
preprint
OpenAlex
Danush Kumar Venkatesh, Adam Schmidt, Muhammad Abdullah Jamal, Omid Mohareri
Surgical video datasets are essential for scene understanding, enabling procedural modeling and intra-operative support. However, these datasets are often heavily imbalanced, with rare actions and tools under-represented, which limits the robustness of downstream models. We address this challenge with $SurgiFlowVid$, a sparse …
Accès ouvert
2025
preprint
OpenAlex
Alejandra Perez, Chinedu Innocent Nwoye, Ramtin Raji Kermani, Omid Mohareri et autres
Vision-language pre-training (VLP) offers unique advantages for surgery by aligning language with surgical videos, enabling workflow understanding and transfer across tasks without relying on expert-labeled datasets. However, progress in surgical VLP remains constrained by the limited scale, procedural diversity, semantic quality, and …
2025
conference-paper
OpenAlex
Muhammad Abdullah Jamal, Omid Mohareri
In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Autoencoder and Denoising techniques. Our proposed approach consists of two stages. In the first stage, we pre-train the …
us
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Mohammadmahdi Honarmand, Muhammad Abdullah Jamal, Omid Mohareri
We introduce VidLPRO, a novel video-language (VL) pre-training framework designed specifically for robotic and laparoscopic surgery. While existing surgical VL models primarily rely on contrastive learning, we propose a more comprehensive approach to capture the intricate temporal dynamics and align video with …
Accès ouvert
2024
preprint
OpenAlex
Muhammad Abdullah Jamal, Omid Mohareri
In this paper, we propose a new progressive pre-training method for image understanding tasks which leverages RGB-D datasets. The method utilizes Multi-Modal Contrastive Masked Autoencoder and Denoising techniques. Our proposed approach consists of two stages. In the first stage, we pre-train the …