2026
conference-paper
OpenAlex
Chang-Hyo Yu, Jaewan Bae, Jinseok Kim, Hongyun Kim et autres
A 4nm-based quad-chiplet with an advanced packaged LLM accelerator achieving 56.8TPS on LLaMA v3.3 70B with single-batch 2k/2k input/output sequences. The architecture combines chiplet-based design, low-latency die-to-die interfaces, unified mixed-precision compute, holistic synchronization, and HBM3E with advanced power schemes to sustain bandwidth, …
2024
conference-paper
OpenAlex
Chang-Hyo Yu, Hyoeun Kim, Sungho Shin, Kyeongryeol Bong et autres
The growing computational demands of AI inference have led to widespread use of hardware accelerators for different platforms, spanning from edge to the datacenter/cloud. Certain AI application areas, such as in high-frequency trading (HFT) [1–2], have a hard inference latency deadline for …
2022
article
OpenAlex
Prasetiyo Prasetiyo, Seongmin Hong, Yashael Faith Arthanto, Joo-Young Kim
Modern deep convolutional neural networks (CNNs) suffer from high computational complexity due to excessive convolution operations. Recently, fast convolution algorithms such as fast Fourier transform (FFT) and Winograd transform have gained attention to address this problem. They reduce the number of multiplications …
kr
(code pays fourni par la source)
2022
conference-paper
OpenAlex
Yashael Faith Arthanto, David Ojika, Joo-Young Kim
By providing highly efficient one-sided communication with globally shared memory space, Partitioned Global Address Space (PGAS) has become one of the most promising parallel computing models in high-performance computing (HPC). Meanwhile, FPGA is getting attention as an alternative compute platform for HPC …
kr
(code pays fourni par la source)
Accès ouvert
2022
preprint
OpenAlex
Yashael Faith Arthanto, David Ojika, Joo-Young Kim
By providing highly efficient one-sided communication with globally shared memory space, Partitioned Global Address Space (PGAS) has become one of the most promising parallel computing models in high-performance computing (HPC). Meanwhile, FPGA is getting attention as an alternative compute platform for HPC …
2019
conference-paper
OpenAlex
Joshua Gunawan, Teresia R. S. Putri, Yashael Faith Arthanto, Trio Adiono
Emotion recognition from speech feature is one of the application where the system needs temporal information in order to produce a correct prediction. On the other hand, recurrent neural network has the advantage of retaining temporal information. This paper proposed a hardware …
id
(code pays fourni par la source)
2019
conference-paper
OpenAlex
Yashael Faith Arthanto, Sayyid Irsyadul Ibad, Trio Adiono
Multiplier is one of the most critical block which is extensively used in many digital systems. However, multiplier consumes large chip area and needs long processing time. Therefore, to improve system computational performance, high speed multiplier design is required. Radix-4 booth multiplier …
id
(code pays fourni par la source)