Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
Rattachement africain : cn, hk, sg. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Large language models (LLMs) have achieved impressive performance across various domains. However, the substantial hardware resources required for their training present a significant barrier to efficiency and scalability. To mitigate this challenge, low-precision training techniques have been widely adopted, leading to notable advancements in training efficiency. Despite these gains, low-precision training involves several components, such as weights, activations, and gradients, each of which can be represented in different numerical formats. The resulting diversity has created a fragmented landscape in low-precision training research, making it difficult for researchers to gain a unified overview of the field. This survey provides a comprehensive review of existing low-precision training methods. To systematically organize these approaches, we categorize them into three primary groups based on their underlying numerical formats, which is a key factor influencing hardware compatibility, computational efficiency, and ease of reference for readers. The categories are (1) fixed-point and integer-based methods, (2) floating-point-based methods, and (3) customized format-based methods. Additionally, we discuss quantization-aware training approaches, which share key similarities with low-precision training during forward propagation. Beyond efficiency, we examine robustness and deployment reliability under low precision. Finally, we highlight several promising research directions to advance this field. A collection of papers discussed in this survey is provided in Awesome-Low-Precision-Training.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
- Date Crossref
- 01/01/2026
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Beijing Institute of Technology pays non établi dans la noticeUniversité ou école supérieure
-
City University of Hong Kong pays non établi dans la noticeUniversité ou école supérieure
-
Sun Yat-sen University pays non établi dans la noticeUniversité ou école supérieure
-
Wuhan University pays non établi dans la noticeUniversité ou école supérieure
-
Nanyang Technological University pays non établi dans la noticeUniversité ou école supérieure
Beijing Institute of Technology, City University of Hong Kong et Sun Yat-sen University, avec 2 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.