Low-rank adaptation with gating mechanisms in large language models, an improved method for fine-tuning: G-LoRA
Le résumé fourni par la source
The gating mechanism in typical models such as Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) effectively controls the flow of information in neural network structures. The Low-Rank Adaptation (LoRA) method involves incorporating two structures, A and B, alongside a pre-trained model. Typically, at the beginning of training, the parameters of these structures are initialized with Gaussian distribution and zeros, respectively. The output dimension of A and the input dimension of B are much smaller than the original model's input and output dimensions. Considering the characteristics of the gating mechanism, A, and B structures, a fusion is performed to achieve control over the information in A and B structures. Experimental results on large Transformer-based models show that, for the same hyperparameters, the LoRA structure with gating mechanism (G-LoRA) provides significant improvements in certain tasks.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Low-rank adaptation with gating mechanisms in large language models, an improved method for fine-tuning: G-LoRA
- Date Crossref
- 16/10/2024
- Éditeur
- SPIE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.