Accès ouvert
2026
preprint
OpenAlex
Alexander Du, Jianjun Ou, Danyang Zhuo, Matthew Lentz
Large language models are increasingly used for code generation, but many generated programs fail to compile, a prerequisite for further correctness checks such as unit tests. Existing solutions for repairing static errors are costly in both latency and token consumption. Post-hoc repair …
Accès ouvert
2026
preprint
OpenAlex
Alexander Du, Jianjun Ou, Danyang Zhuo, Matthew Lentz
Large language models are increasingly used for code generation, but many generated programs fail to compile, a prerequisite for further correctness checks such as unit tests. Existing solutions for repairing static errors are costly in both latency and token consumption. Post-hoc repair …
us
(code pays fourni par la source)
Accès ouvert
2026
article
OpenAlex
Yicheng Jin, Wenjun Hu, Bruce MacDowell Maggs, Xiao Zhang et autres
Embedding-based dense retrieval has become the cornerstone of many critical applications, where approximate nearest neighbor search (ANNS) queries are often combined with filters on labels such as dates and price ranges. Graph-based indexes achieve state-of-the-art performance on unfiltered ANNS but encounter connectivity …
us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Yicheng Jin, Yongji Wu, Wenjun Hu, Bruce MacDowell Maggs et autres
Embedding-based dense retrieval has become the cornerstone of many critical applications, where approximate nearest neighbor search (ANNS) queries are often combined with filters on labels such as dates and price ranges. Graph-based indexes achieve state-of-the-art performance on unfiltered ANNS but encounter connectivity …
Accès ouvert
2026
preprint
OpenAlex
Yicheng Jin, Yongji Wu, Wenjun Hu, Bruce MacDowell Maggs et autres
Embedding-based dense retrieval has become the cornerstone of many critical applications, where approximate nearest neighbor search (ANNS) queries are often combined with filters on labels such as dates and price ranges. Graph-based indexes achieve state-of-the-art performance on unfiltered ANNS but encounter connectivity …
Accès ouvert
2025
preprint
OpenAlex
Zhenzhou Qi, Yiming Li, Chung-Hsuan Tung, Danyang Zhuo et autres
Emerging virtualized radio access networks (vRANs) demand flexible and efficient baseband processing across heterogeneous compute substrates. In this paper, we present DecodeX, a unified benchmarking framework for evaluating low-density parity-check (LDPC) decoding acceleration across different hardware platforms. DecodeX integrates a comprehensive suite …
Accès ouvert
2025
article
OpenAlex
Ceyu Xu, Yongji Wu, X.D. Yang, Beidi Chen et autres
As the parameter size of large language models (LLMs) continues to expand, the need for a large memory footprint and high communication bandwidth have become significant bottlenecks for the training and inference of LLMs.To mitigate these bottlenecks, various tensor compression techniques have …
us, hk
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Yixiao Wang, Qinsi Wang, Ting Xin Jiang, Zhixu Du et autres
Singular Value Decomposition (SVD) has recently seen a surge of interest as a simple yet powerful tool for large language models (LLMs) compression, with a growing number of works demonstrating 20-80% parameter reductions at minimal accuracy loss. Previous SVD-based approaches have focused …
Accès ouvert
2025
preprint
OpenAlex
Mark Lee, Chang Lan, Tom Gunter, John Peebles et autres
AXLearn is a production system which facilitates scalable and high-performance training of large deep learning models. Compared to other state-of-art deep learning systems, AXLearn has a unique focus on modularity and support for hardware-agnostic training. AXLearn's internal interfaces between software components follow …
Accès ouvert
2025
conference-paper
OpenAlex
Jianxing Qin, Alexander Du, Danfeng Zhang, Matthew Lentz et autres
Large language models (LLMs) have demonstrated remarkable coding capabilities. They excel in code synthesis benchmarks across diverse domains and have become ubiquitous in coding tools. Recently, they have also shown promise in generating mathematical proofs and small software programs. In this paper, …
us
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Xiangfeng Zhu, Yang Zhou, Yuyao Wang, Xiangyu Gao et autres
Fast and efficient RPCs are key to the performance of applications based on microservices. But RPC communication suffers from significant overhead today because it relies on the standard, layered protocol stack and loose coupling between the end host and in-network proxies that …
us
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Jingrong Chen, Yongji Wu, Liang Luo, Zhaodong Wang et autres
Modern machine learning (ML) training workloads place substantial demands on both computational and communication resources. Consequently, accurate performance estimation has become increasingly critical for guiding system design decisions, such as the selection of parallelization strategies, cluster configurations, and hardware provisioning. Existing simulation-based …