Aller au contenu principal
Accès ouvert déclaré 2026 article

Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities

2Citations signalées, ce qui n’est pas une note de qualité
5Institutions déclarées
3Pays d’affiliation déclarés

Rattachement africain : cn, hk, sg. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Large language models (LLMs) have achieved impressive performance across various domains. However, the substantial hardware resources required for their training present a significant barrier to efficiency and scalability. To mitigate this challenge, low-precision training techniques have been widely adopted, leading to notable advancements in training efficiency. Despite these gains, low-precision training involves several components, such as weights, activations, and gradients, each of which can be represented in different numerical formats. The resulting diversity has created a fragmented landscape in low-precision training research, making it difficult for researchers to gain a unified overview of the field. This survey provides a comprehensive review of existing low-precision training methods. To systematically organize these approaches, we categorize them into three primary groups based on their underlying numerical formats, which is a key factor influencing hardware compatibility, computational efficiency, and ease of reference for readers. The categories are (1) fixed-point and integer-based methods, (2) floating-point-based methods, and (3) customized format-based methods. Additionally, we discuss quantization-aware training approaches, which share key similarities with low-precision training during forward propagation. Beyond efficiency, we examine robustness and deployment reliability under low precision. Finally, we highlight several promising research directions to advance this field. A collection of papers discussed in this survey is provided in Awesome-Low-Precision-Training.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
Date Crossref
01/01/2026
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Beijing Institute of Technology pays non établi dans la notice
    Université ou école supérieure
  • City University of Hong Kong pays non établi dans la notice
    Université ou école supérieure
  • Sun Yat-sen University pays non établi dans la notice
    Université ou école supérieure
  • Wuhan University pays non établi dans la notice
    Université ou école supérieure
  • Nanyang Technological University pays non établi dans la notice
    Université ou école supérieure

Beijing Institute of Technology, City University of Hong Kong et Sun Yat-sen University, avec 2 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Topic ModelingNatural Language Processing Techniques

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.