Accès ouvert
2026
preprint
OpenAlex
Akhiad Bercovich, Nir Ailon, Vladimir Anisimov, Tomer Asida et autres
Reasoning-focused LLMs improve answer quality by generating longer reasoning traces, but the additional tokens dramatically increase serving cost, motivating inference optimization. We extend and apply Puzzle, a post-training neural architecture search (NAS) framework, to gpt-oss-120B to produce gpt-oss-puzzle-88B, a deployment-optimized derivative. Our …
Accès ouvert
2026
preprint
OpenAlex
Akhiad Bercovich, Nir Ailon, Vladimir Anisimov, Tomer Asida et autres
Reasoning-focused LLMs improve answer quality by generating longer reasoning traces, but the additional tokens dramatically increase serving cost, motivating inference optimization. We extend and apply Puzzle, a post-training neural architecture search (NAS) framework, to gpt-oss-120B to produce gpt-oss-puzzle-88B, a deployment-optimized derivative. Our …
Accès ouvert
2026
preprint
OpenAlex
Meng Xin, Sweta Priyadarshi, Jingyu Xin, Bilal Kartal et autres
This technical report presents quantization-aware distillation (QAD) and our best practices for recovering accuracy of NVFP4-quantized large language models (LLMs) and vision-language models (VLMs). QAD distills a full-precision teacher model into a quantized student model using a KL divergence loss. While applying …
Accès ouvert
2026
preprint
OpenAlex
Meng Xin, Sweta Priyadarshi, Jingyu Xin, Bilal Kartal et autres
This technical report presents quantization-aware distillation (QAD) and our best practices for recovering accuracy of NVFP4-quantized large language models (LLMs) and vision-language models (VLMs). QAD distills a full-precision teacher model into a quantized student model using a KL divergence loss. While applying …
Accès ouvert
2025
preprint
OpenAlex
NVIDIA, :, Aaron Blakeman, Aaron Grattafiori et autres
We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to provide best-in-class throughput and context lengths of up to …
Accès ouvert
2025
preprint
OpenAlex
NVIDIA, :, Aaron Blakeman, Aaron Grattafiori et autres
We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to provide best-in-class throughput and context lengths of up to …
Accès ouvert
2025
preprint
OpenAlex
NVIDIA, :, Aaron Blakeman, Aaron Grattafiori et autres
We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and large-scale RL on …
Accès ouvert
2025
preprint
OpenAlex
NVIDIA, :, Aaron Blakeman, Aaron Grattafiori et autres
We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and large-scale RL on …