From Nimitz to NetPila: The Evolution of Production-Scale Container Network
Jiaqi Gao, Chao Qin, Sheng Cheng, Jiamin Cao et autres
cn, us (code pays fourni par la source)
Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.
Jiaqi Gao, Chao Qin, Sheng Cheng, Jiamin Cao et autres
cn, us (code pays fourni par la source)
Zhaochen Zhang, Jiaqi Gao, Sheng Cheng, Peiwen Yu et autres
cn (code pays fourni par la source)
Jiamin Cao, Qingxu Li, Yaozhong Liu, Jiaqi Gao et autres
The Mixture-of-Experts (MoE) architecture scales large language models (LLMs) to trillions of parameters by activating only a small subset of experts per token. In practice, MoE inference is commonly deployed with Expert Parallelism (EP), which places whole experts on different GPUs to …
cn (code pays fourni par la source)
Chenxu Wang, Xumiao Zhang, Xianshang Lin, Xuan Zeng et autres
Root causing and failure localization are critical to maintain reliability in cloud network operations. When an incident is reported, network operators must review massive volumes of monitoring data and identify the root cause (i.e., error device) as fast as possible, making it …
cn, us (code pays fourni par la source)
Yifei Yuan, Fangdan Ye, Yifan Li, Mengqi Liu et autres
The network verification system Hoyan has been deployed for Alibaba Cloud's wide-area network (WAN) for years and achieved considerable success in preventing misconfiguration-caused network incidents. However, recent years have seen the emergence of new challenges in scalability, usability, and accuracy for Hoyan. …
us, cn (code pays fourni par la source)
Huanwu Hu, Yifan Li, Yunguang Li, Bingchuan Tian et autres
For providers operating large-scale global networks, the timeliness of network failure recovery significantly affects the reliability of network services. Ideally, a network monitoring system should have enough coverage to detect even minor issues, but high coverage means alert floods during severe network …
cn, us (code pays fourni par la source)
Jie Lu, Jiaqi Gao, Fei Feng, Zhiqiang Hé et autres
The rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), …
cn, us (code pays fourni par la source)
Jiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu et autres
The performance of collective communication schedules is crucial for the efficiency of machine learning jobs and GPU cluster utilization. Existing open-source collective communication libraries (such as NCCL and RCCL) rely on fixed schedules and cannot adjust to varying topology and model requirements. …
cn, us (code pays fourni par la source)
Yu Guan, Zhiyu Yin, Haoyu Chen, Sheng Cheng et autres
Troubleshooting performance problems of large model training (LMT) is immensely challenging, due to unprecedented scales of modern GPU clusters, the complexity of software-hardware interactions, and the data intensity of the training process. Existing troubleshooting approaches designed for traditional distributed systems or datacenter …
Jianbo Dong, Bin Luo, Jun Zhang, Pengcheng Zhang et autres
The emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, the efficiency of large-scale distributed training systems is often suboptimal due to the increased …
us (code pays fourni par la source)
Kun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao et autres
This paper presents HPN, Alibaba Cloud's data center network for large language model (LLM) training. Due to the differences between LLMs and general cloud computing (e.g., in terms of traffic patterns and fault tolerance), traditional data center networks are not well-suited for …
cn, us (code pays fourni par la source)
Ruihan Li, Yifei Yuan, Fangdan Ye, Mengqi Liu et autres
This paper presents YU, the first verification system for checking traffic load properties under arbitrary failure scenarios that can scale to production Wide Area Networks (WANs). Building a practical YU requires us to address two challenges in terms of generality and efficiency. …
cn, us (code pays fourni par la source)
BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.
L'essentiel de l'actu tech du Burkina & d'Afrique, chaque semaine dans votre boîte mail.
Gratuit · sans spam · désinscription en un clic