Accès ouvert
2026
preprint
OpenAlex
NVIDIA, :, Aaron Blakeman, Austin Thomas et autres
We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M tokens, and post-trained using Supervised Fine …
Accès ouvert
2026
preprint
OpenAlex
NVIDIA, :, Aaron Blakeman, Austin Thomas et autres
We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 trillion text tokens, then extended the context length to 1M tokens, and post-trained using Supervised Fine …
Accès ouvert
2026
preprint
OpenAlex
NVIDIA, :, Amala Sanjay Deshmukh, Kateryna Chumachenko et autres
We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its predecessor, Nemotron Nano V2 VL, …
Accès ouvert
2026
preprint
OpenAlex
NVIDIA, :, Amala Sanjay Deshmukh, Kateryna Chumachenko et autres
We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its predecessor, Nemotron Nano V2 VL, …
2026
conference-paper
OpenAlex
Vignesh Prabhakar, Md Amirul Islam, Adam Atanas, Yao‐Ting Wang et autres
Large Language Models (LLMs) have demonstrated remarkable potential in advancing scientific knowledge and addressing complex challenges. In this work, we introduce OmniScience, a specialized large reasoning model for general science, developed through three key components: (1) domain adaptive pretraining on a carefully …
Accès ouvert
2025
preprint
OpenAlex
NVIDIA, Amala Sanjay Deshmukh, Kateryna Chumachenko, Tuomas Rintamaki et autres
We introduce Nemotron Nano V2 VL, the latest model of the Nemotron vision-language series designed for strong real-world document understanding, long video comprehension, and reasoning tasks. Nemotron Nano V2 VL delivers significant improvements over our previous model, Llama-3.1-Nemotron-Nano-VL-8B, across all vision and …
Accès ouvert
2024
preprint
OpenAlex
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Bo Liu et autres
The impressive capabilities of recent language models can be largely attributed to the multi-trillion token pretraining datasets that they are trained on. However, model developers fail to disclose their construction methodology which has lead to a lack of open information on how …
Accès ouvert
2024
preprint
OpenAlex
Nvidia, Bo Adler, Niket Agarwal, Ashwath Aithal et autres
We release the Nemotron-4 340B model family, including Nemotron-4-340B-Base, Nemotron-4-340B-Instruct, and Nemotron-4-340B-Reward. Our models are open access under the NVIDIA Open Model License Agreement, a permissive model license that allows distribution, modification, and use of the models and its outputs. These models …
Accès ouvert
2024
preprint
OpenAlex
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Mostofa Patwary et autres
We introduce Nemotron-4 15B, a 15-billion-parameter large multilingual language model trained on 8 trillion text tokens. Nemotron-4 15B demonstrates strong performance when assessed on English, multilingual, and coding tasks: it outperforms all existing similarly-sized open models on 4 out of 7 downstream …
Accès ouvert
2024
article
OpenAlex
Micha Livne, Zulfat Miftahutdinov, Elena Tutubalina, Maksim Kuznetsov et autres
Large Language Models (LLMs) have substantially driven scientific progress in various domains, and many papers have demonstrated their ability to tackle complex problems with creative solutions. Our paper introduces a new foundation model, nach0, capable of solving various chemical and biological tasks: …
us, ca
(code pays fourni par la source)
Accès ouvert
2024
conference-paper
OpenAlex
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Bo Liu et autres
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Bo Liu, Aastha Jhunjhunwala, Zhilin Wang, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Accès ouvert
2023
preprint
OpenAlex
Micha Livne, Zulfat Miftahutdinov, Elena Tutubalina, Maksim Kuznetsov et autres
Large Language Models (LLMs) have substantially driven scientific progress in various domains, and many papers have demonstrated their ability to tackle complex problems with creative solutions. Our paper introduces a new foundation model, nach0, capable of solving various chemical and biological tasks: …
us, hk
(code pays fourni par la source)