Aller au contenu principal
Profil bibliographique

Huayou Su

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

77Publications signalées
563Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

Video Coding and Compression TechnologiesParallel Computing and Optimization TechniquesAdvanced Vision and ImagingAdvanced Data Storage TechnologiesAdvanced Graph Neural Networks

Les publications récentes

Accès ouvert 2026 conference-paper OpenAlex

TileGEMM: Boosting the Performance of GEMM on AMX-Powered CPUs by Exploiting Data Reuse

Kangkang Chen, Huayou Su, Menghan Jia, Yong Dou

General Matrix Multiplication (GEMM) is the cornerstone of high-performance computing and deep learning. Its efficiency significantly influences the performance of applications ranging from large language models to scientific simulations. Intel Advanced Matrix Extensions (AMX) significantly boost matrix operations throughput, yet existing implementations …

cn (code pays fourni par la source)

0 citations
Accès ouvert 2026 article OpenAlex

Eliminating Redundant Feature Tests in Decision Tree and Random Forest Inference on SQL Predicates

Mingxi Liu, Zhengyuan Ding, Chenyang Zhang, Qingfeng Pan et autres

In-database prediction queries that apply machine learning (ML) pipelines to perform data analysis are prevalent in many applications. Since data stored in databases is typically tabular, tree-based models are particularly well-suited and thus widely adopted for such tasks. When ML inference with …

cn, us (code pays fourni par la source)

0 citations Proceedings of the ACM on Management of Data
Accès ouvert 2026 article OpenAlex

Quantitative Analysis and Performance Optimization of Graph Neural Networks on Multi-core CPUs

Huayou Su, Xi Yang, Zitong An, Yong Dou et autres

Graph Neural Networks (GNNs) are becoming increasingly popular in graph data processing due to their excellent performance in feature extraction on graph datasets. Compared to GPUs, CPUs are more widely accessible and serve as a practical platform for GNN inference. However, achieving …

cn (code pays fourni par la source)

0 citations ACM Transactions on Architecture and Code Optimization
2026 article OpenAlex

Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-Scale MoE Models

Wei Wang, Zhiquan Lai, Dongsheng Li, Shengwei Li et autres

The size of deep learning models has been increasing to enhance model quality. The linear increase in training computation budgets with model size means that training an extremely large-scale model is exceedingly time-consuming. Recently, the Mixture of Experts (MoE) has drawn significant …

cn (code pays fourni par la source)

0 citations IEEE Transactions on Parallel and Distributed Systems

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.