Accès ouvert déclaré
2025
preprint
Huawei Cloud Model-as-a-Service on the CloudMatrix384 SuperPod
Ao Xiao, Bangzheng He, Baoquan Zhang, Baoxing Huai, Bingji Wang, Bo Wang, Bo Xu, Chan K. Yang, Changhong Liu, Cheng Cui, Chenyu Zhu, Cong Feng, Daohui Wang, Duo Zhao, Fu Wang, Gangqiang Zhang, Guanjie Chen, Guodong Yang, Haifeng Li, Haley Li, Hao Feng, Hao Huang, Hao Xu, Hengrui Ma, Hui Liu, Jia Li, Jiang Liu, Jiang Xu, Jie Meng, Jinhan Xin, Junhao Hu, Jiani Chen, Yu Lan, Liang Liu, Lu Zhou, Meina Han, Mingyu Deng, Naitian Deng, Pei Zhao, Peng Pan, Pengfei Shen, Ping Li, Qi Zhang, Qian Wang, Qingyi Zhang, Qunchao Fu, Ruimin Gao, Shaochun Li, Long Sheng, Shuo Li, Siqi Wan, Shuai Shen, Song Zhang, Tao Xu, Tianlin Du, Ting Chen, W.-G. Wu, Wei Jiang, Weiwei Chen, Wen Peng, Wenli Zhou, Wenquan Yang, Wenxin Liang, Xiang Liu, Xiaoli Zhou, Xin Jin, Xinyu Duan, Xu Li, Xu Zhang, Xusheng Chen, Yikang Shan, Yang Gan, Yao Lu, Yi Deng, Yi Zheng, Xiong Ying, Yingfei Zheng, Yizhou Shan, Yong Gao, Yongqiang Yang, Yueqiu Gong, Yue Yu, Yuetao Chen, Yukun Zhu, Yulong He, Yan Zhao, Yuting Wu, Zhaoyang Ji, Zhefeng Wang, Zheng Wang, Zhenan Fan, Zhenhua Yang, Zhenli Sheng, Zhibin Yu, Zhigang Ji, Zhihao Ren, Zhixia Liu, Zhiyu Dong, LI Zhong-hua, Yu Zhou, Zhibin Shen, Zhuwei Peng, Zi Ye, Zihao Xiang, Zheng Fu, Zixuan Zhang
0Citations signalées, ce qui n’est pas une note de qualité
0Institutions déclarées
0Pays d’affiliation déclarés
Le résumé fourni par la source
Scaled-out MoE LLMs and scaled-up SuperPods create new systems challenges for production Model-as-a-Service (MaaS), requiring disaggregation, low-latency communication, and decentralized serving. This report presents xDeepServe, the production serving system behind Huawei Cloud's MaaS offering on CloudMatrix384, a 48-server SuperPod with 384 Ascend 910C chips connected by a high-bandwidth UB fabric and global shared memory. It serves models including DeepSeek, Kimi, GLM, Qwen, and MiniMax, among others. xDeepServe is built around Transformerless, a disaggregated execution architecture that decomposes transformer inference into modular units -- attention, feedforward, and MoE -- and supports disaggregated Prefill-Decode and MoE-Attention deployments. To enable disaggregation, we develop XCCL, a memory-semantic communication layer providing microsecond-level point-to-point and scalable all-to-all primitives, and we extend FlowServe with decentralized DP groups and techniques to mitigate stragglers and synchronization variance. In a peak decoding configuration, xDeepServe reaches 2400 tokens/s per Ascend 910C chip at ~50ms time-per-output-token (TPOT).
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
La source scientifique ouverte est momentanément indisponible.
Les sujets associés
Parallel Computing and Optimization TechniquesBig Data and Digital EconomyCloud Computing and Resource Management