Spatiotemporal distribution of mixture homogeneity: Mechanism analysis of injection pulse width on combustion characteristics in hydrogen engines
Kan Zhu, Yunhua Zhang, Diming Lou, Liang Fang et autres
cn (code pays fourni par la source)
Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.
Kan Zhu, Yunhua Zhang, Diming Lou, Liang Fang et autres
cn (code pays fourni par la source)
Kan Zhu, Diming Lou, Yunhua Zhang, Liang Fang et autres
NVIDIA, :, Aaron Blakeman, Aaron Grattafiori et autres
We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to provide best-in-class throughput and context lengths of up to …
NVIDIA, :, Aaron Blakeman, Aaron Grattafiori et autres
We introduce the Nemotron 3 family of models - Nano, Super, and Ultra. These models deliver strong agentic, reasoning, and conversational capabilities. The Nemotron 3 family uses a Mixture-of-Experts hybrid Mamba-Transformer architecture to provide best-in-class throughput and context lengths of up to …
NVIDIA, :, Aaron Blakeman, Aaron Grattafiori et autres
We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and large-scale RL on …
NVIDIA, :, Aaron Blakeman, Aaron Grattafiori et autres
We present Nemotron 3 Nano 30B-A3B, a Mixture-of-Experts hybrid Mamba-Transformer language model. Nemotron 3 Nano was pretrained on 25 trillion text tokens, including more than 3 trillion new unique tokens over Nemotron 2, followed by supervised fine tuning and large-scale RL on …
Kan Zhu, Diming Lou, Yunhua Zhang, Liang Fang et autres
cn (code pays fourni par la source)
Yilong Zhao, Jiaming Tang, Kan Zhu, Zihao Ye et autres
Reasoning language models have demonstrated remarkable capabilities on challenging tasks by generating elaborate chain-of-thought (CoT) solutions. However, such lengthy generation shifts the inference bottleneck from compute-bound to memory-bound. To generate each token, the model applies full attention to all previously generated tokens, …
Xinyue Rao, Diming Lou, Yunhua Zhang, Zhiwei Wang et autres
cn (code pays fourni par la source)
Kan Zhu, Haiyang Shi, Lei Xu, Jiaxin Shan et autres
Advances in Large Language Models (LLMs) have led to a surge of LLM-powered applications. These applications have diverse token-generation latency requirements. As a result, simply classifying workloads as latency-sensitive (LS) or best-effort (BE) overlooks the nuances within the latency-sensitive category and results …
Kan Zhu, Tian Tang, Qiang Xu, Yile Gu et autres
Long-context models are essential for many applications but face inefficiencies in loading large KV caches during decoding. Prior methods enforce fixed token budgets for sparse attention, assuming a set number of tokens can approximate full attention. However, these methods overlook variations in …
Yilong Zhao, Shuo Yang, Kan Zhu, Lianmin Zheng et autres
Offline batch inference, which leverages the flexibility of request batching to achieve higher throughput and lower costs, is becoming more popular for latency-insensitive applications. Meanwhile, recent progress in model capability and modality makes requests more diverse in compute and memory demands, creating …
BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.
L'essentiel de l'actu tech du Burkina & d'Afrique, chaque semaine dans votre boîte mail.
Gratuit · sans spam · désinscription en un clic