Aller au contenu principal
Profil bibliographique

Hanrong Ye

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

48Publications signalées
259Citations signalées
2Affiliations récentes

Les institutions déclarées

Les domaines associés

Advanced Neural Network ApplicationsMultimodal Machine Learning ApplicationsDomain Adaptation and Few-Shot LearningSpeech Recognition and SynthesisHuman Pose and Action Recognition

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics

Junqi Liu, Yufan He, Yexiao He, Pengfei Guo et autres

Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, K. S. Jayakumar, Lasha Koroshinadze et autres

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Sreyan Ghosh, Arushi Goel, K. S. Jayakumar, Lasha Koroshinadze et autres

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over …

gb, us (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

Shufan Li, Greg Heinrich, Hanrong Ye, Yonggan Fu et autres

We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared with prior work on masked image generation, Nemotron-Labs-Diffusion-Image addresses two key challenges. First, unlike continuous diffusion models which progressively refine latent representations across the entire image, standard …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

Shufan Li, Greg Heinrich, Hanrong Ye, Yonggan Fu et autres

We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared with prior work on masked image generation, Nemotron-Labs-Diffusion-Image addresses two key challenges. First, unlike continuous diffusion models which progressively refine latent representations across the entire image, standard …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Vesta: A Generalist Embodied Reasoning Model

Johan Björck, Zhiqi Li, Yunze Man, Jing Wang et autres

Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deploying a multi-model stack is computationally expensive and prone to cascading errors. We present Vesta, a unified embodied generalist that …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Vesta: A Generalist Embodied Reasoning Model

Johan Björck, Zhiqi Li, Yunze Man, Yi-Xiang Wang et autres

Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deploying a multi-model stack is computationally expensive and prone to cascading errors. We present Vesta, a unified embodied generalist that …

us, hk, sg, de (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders

Yitong Jiang, Hongjun Wang, Collin McCarthy, Hanrong Ye et autres

Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale pretraining. Subquadratic alternatives such as linear attention and state-space models reduce this cost, but often serialize images into 1D token streams …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders

Yitong Jiang, Hongjun Wang, Collin McCarthy, Hanrong Ye et autres

Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale pretraining. Subquadratic alternatives such as linear attention and state-space models reduce this cost, but often serialize images into 1D token streams …

hk, us (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search

Dongyun Zou, Zhuoyang Zhang, Junyu Chen, Wenkun He et autres

We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while achieving substantially higher inference efficiency on high-resolution images. At the core of our approach is Post-Training Attention Search, a …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search

Dongyun Zou, Zhuoyang Zhang, Junyu Chen, Wenkun He et autres

We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while achieving substantially higher inference efficiency on high-resolution images. At the core of our approach is Post-Training Attention Search, a …

0 citations arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.