2026
article
OpenAlex
Wei Wang, Zhiquan Lai, Dongsheng Li, Shengwei Li et autres
The size of deep learning models has been increasing to enhance model quality. The linear increase in training computation budgets with model size means that training an extremely large-scale model is exceedingly time-consuming. Recently, the Mixture of Experts (MoE) has drawn significant …
cn
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Zhiquan Lai, Weijie Liu, Shengwei Li, Wei Wang et autres
Multimodal large language models (MLLMs) achieve strong performance across diverse AI applications but remain costly to train. With advancing hardware, heterogeneous clusters are increasingly used for MLLM training. However, existing methods often ignore MLLMs' hybrid-module structure and varying parameter states, leading to …
cn
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Zhiquan Lai, Shengwei Li, Weijie Liu, Wei Wang et autres
Mixture-of-Experts (MoE) has been extensively adopted for its incredible capability to expand model scale with a sub-linear increase in computational requirement. Training MoE models requires substantial computing nodes and extended periods, necessitating reliable distributed training systems. Checkpointing is a common approach to …
cn
(code pays fourni par la source)
2025
article
OpenAlex
Xiaoge Deng, Li Shen, Shengwei Li, Dacheng Tao
Stochastic gradient descent (SGD) performed in an asynchronous manner plays a crucial role in training large-scale machine learning models. However, the generalization performance of asynchronous delayed SGD, which is an essential metric for assessing machine learning algorithms, has rarely been explored. Existing …
cn, sg
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Shengwei Li, Yinshan Wang, Jiayin Li, Yize Li et autres
This paper proposes a novel dual-channel metal detection method that integrates the Frost Optimization Algorithm with a BiLSTM-Attention framework to enhance detection accuracy, material discrimination, and anti-interference performance in complex environments. Utilizing the RIME algorithm for multi-frequency scanning, the proposed method achieves …
cn
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Shengwei Li, Xiangyu Dai, Wei Huang, Xiangfeng Kong
cn
(code pays fourni par la source)
2024
article
OpenAlex
Weijie Liu, Zhiquan Lai, Shengwei Li, Keshi Ge et autres
Recently, the data-parallel pipeline approach has been widely used in training DNN models on commodity GPU servers. However, there are still three challenges for hybrid parallelism on commodity GPU servers: i) a balanced model partition is crucial for efficiency, whereas prior works …
cn
(code pays fourni par la source)
2024
conference-paper
OpenAlex
Z. Zheng, Qingyuan Xia, Shengwei Li, Bohai Deng
In this paper, an improved hybrid algorithm based on particle swarm optimization (PSO) and genetic algorithm (GA) is formed to solve the problem of "premature" and slow convergence. First of all, this paper improves the multi-task Collaborative assignment model (CMTAP), comprehensively considers …
cn
(code pays fourni par la source)
2024
conference-paper
OpenAlex
Yanqi Hao, Zhiquan Lai, Wei Wang, Shengwei Li et autres
Large deep neural network (DNN) models have demonstrated exceptional performance across diverse downstream tasks. Sharded data parallelism (SDP) has been widely used to reduce the memory footprint of model states. In a DNN training cluster, a device usually has multiple inter-device links …
cn
(code pays fourni par la source)
2024
article
OpenAlex
Shengwei Li, Zeping Tong, Muhammad Haroon
cn, pk
(code pays fourni par la source)
2024
article
OpenAlex
Shengwei Li, Kai Lü, Zhiquan Lai, Weijie Liu et autres
The transformer-based deep neural network (DNN) models have shown considerable success across diverse tasks, prompting widespread adoption of distributed training methods such as data parallelism and pipeline parallelism. With the increasing parameter number, hybrid parallel training becomes imperative to scale training. The …
cn
(code pays fourni par la source)
2023
conference-paper
OpenAlex
Zhiquan Lai, Yanqi Hao, Shengwei Li, Dongsheng Li
Multidimensional parallel training has been widely applied to train large-scale deep learning models like GPT-3. The efficiency of parameter communication among training devices/processes is often the performance bottleneck of large model training. Analysis of parameter communication mode and traffic has important reference …
cn
(code pays fourni par la source)