Aller au contenu principal
Profil bibliographique

Dennis Cai

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

34Publications signalées
773Citations signalées
2Affiliations récentes

Les institutions déclarées

Les domaines associés

Cloud Computing and Resource ManagementSoftware-Defined Networks and 5GSoftware System Performance and ReliabilityAdvanced Optical Network TechnologiesNetwork Traffic and Congestion Control

Les publications récentes

Accès ouvert 2026 conference-paper OpenAlex

Balancing and Beyond: Communication-Centric Optimizations in Expert Parallelism

Jiamin Cao, Qingxu Li, Yaozhong Liu, Jiaqi Gao et autres

The Mixture-of-Experts (MoE) architecture scales large language models (LLMs) to trillions of parameters by activating only a small subset of experts per token. In practice, MoE inference is commonly deployed with Expert Parallelism (EP), which places whole experts on different GPUs to …

cn (code pays fourni par la source)

0 citations
Accès ouvert 2025 conference-paper OpenAlex

Towards LLM-Based Failure Localization in Production-Scale Networks

Chenxu Wang, Xumiao Zhang, Xianshang Lin, Xuan Zeng et autres

Root causing and failure localization are critical to maintain reliability in cloud network operations. When an incident is reported, network operators must review massive volumes of monitoring data and identify the root cause (i.e., error device) as fast as possible, making it …

cn, us (code pays fourni par la source)

13 citations
Accès ouvert 2025 conference-paper OpenAlex

New Evolution of Hoyan: Enhancing Scalability, Usability, and Accuracy for Alibaba's Global WAN Verification

Yifei Yuan, Fangdan Ye, Yifan Li, Mengqi Liu et autres

The network verification system Hoyan has been deployed for Alibaba Cloud's wide-area network (WAN) for years and achieved considerable success in preventing misconfiguration-caused network incidents. However, recent years have seen the emergence of new challenges in scalability, usability, and accuracy for Hoyan. …

us, cn (code pays fourni par la source)

5 citations
Accès ouvert 2025 conference-paper OpenAlex

SkyNet: Analyzing Alert Flooding from Severe Network Failures in Large Cloud Infrastructures

Huanwu Hu, Yifan Li, Yunguang Li, Bingchuan Tian et autres

For providers operating large-scale global networks, the timeliness of network failure recovery significantly affects the reliability of network services. Ideally, a network monitoring system should have enough coverage to detect even minor issues, but high coverage means alert floods during severe network …

cn, us (code pays fourni par la source)

3 citations
Accès ouvert 2025 conference-paper OpenAlex

Alibaba Stellar: A New Generation RDMA Network for Cloud AI

Jie Lu, Jiaqi Gao, Fei Feng, Zhiqiang Hé et autres

The rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), …

cn, us (code pays fourni par la source)

14 citations
Accès ouvert 2025 conference-paper OpenAlex

SyCCL: Exploiting Symmetry for Efficient Collective Communication Scheduling

Jiamin Cao, Shangfeng Shi, Jiaqi Gao, Weisen Liu et autres

The performance of collective communication schedules is crucial for the efficiency of machine learning jobs and GPU cluster utilization. Existing open-source collective communication libraries (such as NCCL and RCCL) rely on fixed schedules and cannot adjust to varying topology and model requirements. …

cn, us (code pays fourni par la source)

14 citations
Accès ouvert 2025 preprint OpenAlex

EROICA: Online Performance Troubleshooting for Large-scale Model Training

Yu Guan, Zhiyu Yin, Haoyu Chen, Sheng Cheng et autres

Troubleshooting performance problems of large model training (LMT) is immensely challenging, due to unprecedented scales of modern GPU clusters, the complexity of software-hardware interactions, and the data intensity of the training process. Existing troubleshooting approaches designed for traditional distributed systems or datacenter …

0 citations arXiv (Cornell University)
2025 conference-paper OpenAlex

Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization

Jianbo Dong, Bin Luo, Jun Zhang, Pengcheng Zhang et autres

The emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single model. Unfortunately, the efficiency of large-scale distributed training systems is often suboptimal due to the increased …

us (code pays fourni par la source)

3 citations
2024 conference-paper OpenAlex

A General and Efficient Approach to Verifying Traffic Load Properties under Arbitrary k Failures

Ruihan Li, Yifei Yuan, Fangdan Ye, Mengqi Liu et autres

This paper presents YU, the first verification system for checking traffic load properties under arbitrary failure scenarios that can scale to production Wide Area Networks (WANs). Building a practical YU requires us to address two challenges in terms of generality and efficiency. …

cn, us (code pays fourni par la source)

9 citations

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.