Aller au contenu principal
Profil bibliographique

Ceyu Xu

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

28Publications signalées
108Citations signalées
2Affiliations récentes

Les institutions déclarées

Les domaines associés

Parallel Computing and Optimization TechniquesEmbedded Systems Design TechniquesCloud Computing and Resource ManagementTopic ModelingNatural Language Processing Techniques

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

Memory Compression for High-Fanout Agent Sandboxes

Mengming Li, Ceyu Xu, Qijun Zhang, Jiangnan Yu et autres

High-fanout agent workloads create a growing memory bottleneck because a single task may spawn many concurrent sandbox sessions. Yet these sandboxes are far from independent: they originate from a shared template and execute related trajectories, exposing substantial template-relative and cross-sandbox memory redundancy. …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Why Do Prefetchers Fail? Let Agents Answer

Xudong Sun, Ceyu Xu, Ningzhi Ai, Zeyu Zhu et autres

Hardware prefetchers are crucial to processor performance, yet their design remains labor-intensive and expert-driven. Architects inspect execution and memory-access traces, identify patterns, translate them into online hardware heuristics, and evaluate them in simulation, often with no guarantee of improvement. Human experts cannot …

0 citations arXiv (Cornell University)
2026 conference-paper OpenAlex

ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory Accesses

Mengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang et autres

Irregular memory accesses pose challenges for effective and efficient data prefetching. While temporal prefetchers have recently shown promise for irregular memory access patterns, their effectiveness fundamentally depends on temporal address recurrence and large metadata storage. When memory addresses exhibit weak or no …

hk (code pays fourni par la source)

0 citations
Accès ouvert 2026 preprint OpenAlex

STS: Efficient Sparse Attention with Speculative Token Sparsity

Ceyu Xu, Jiangnan Yu, Yongji Wu, Yuan Xie

The quadratic complexity of attention imposes severe memory and computational bottlenecks on Large Language Model (LLM) inference. This challenge is particularly acute for emerging agentic applications that require processing multi-million token sequences. We propose STS, a sparse attention mechanism that requires no …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory Accesses

Mengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang et autres

Irregular memory accesses pose challenges for effective and efficient data prefetching. While temporal prefetchers have recently shown promise for irregular memory access patterns, their effectiveness fundamentally depends on temporal address recurrence and large metadata storage. When memory addresses exhibit weak or no …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

ICP: Exploiting Instruction Correlation for Prefetching Irregular Memory Accesses

Mengming Li, Chenlu Miao, Buqing Xu, Qijun Zhang et autres

Irregular memory accesses pose challenges for effective and efficient data prefetching. While temporal prefetchers have recently shown promise for irregular memory access patterns, their effectiveness fundamentally depends on temporal address recurrence and large metadata storage. When memory addresses exhibit weak or no …

hk (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2026 conference-paper OpenAlex

PF-LLM: L arge L anguage M odel Hinted Hardware P re f etching

Ceyu Xu, Xiangfeng Sun, Weihang Li, Chen Bai et autres

Hardware data prefetching is a critical technique for mitigating memory latency in modern processors. While sophisticated hardware prefetching algorithms exist, their exclusive reliance on runtime information limits their ability to adapt quickly and comprehend broader program context. Our key insight is that …

hk, us (code pays fourni par la source)

0 citations
Accès ouvert 2026 preprint OpenAlex

VersaQ-3D: Architecture Support for Visual Geometry Grounded Transformers via Versatile Quantization

Yipu Zhang, Jintao Cheng, Xingyu Liu, Zeyu Li et autres

3D reconstruction and view synthesis are fundamental to AR/VR, robotics, and digital twins. The Visual Geometry Grounded Transformer (VGGT) enables strong feed-forward 3D reconstruction while its billion-parameter scale limits on-device deployment. LLM-oriented quantization methods fail on VGGT due to saturated activation channels …

0 citations arXiv (Cornell University)
2026 article OpenAlex

BISC: Rethinking Instruction Set Abstractions at Block Granularity

An Zhong, Hanzhi Hu, Zirui Ma, Mengzhu Li et autres

Most widely deployed ISAs for general-purpose processors expose a centralized register file and instruction-granular control flow, forcing modern CPUs to rely on complex microarchitectural mechanisms to reconcile execution with architectural state. Even in simple in-order designs, these abstractions impose back-and-forth execution overheads, …

sa, cn, hk (code pays fourni par la source)

0 citations IEEE Computer Architecture Letters
Accès ouvert 2025 article OpenAlex

LLM.265: Video Codecs Are Secretly Tensor Codecs

Ceyu Xu, Yongji Wu, X.D. Yang, Beidi Chen et autres

As the parameter size of large language models (LLMs) continues to expand, the need for a large memory footprint and high communication bandwidth have become significant bottlenecks for the training and inference of LLMs.To mitigate these bottlenecks, various tensor compression techniques have …

us, hk (code pays fourni par la source)

6 citations IEEE Micro

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.