Accès ouvert
2026
preprint
OpenAlex
Janghwan Lee, Sihwa Lee, Jinseok Kim, Yongjik Kim et autres
Large Reasoning Models (LRMs) achieve strong problem-solving through long chain-of-thought, but their deployment is constrained by the high cost of full-precision inference and growing KV cache footprints. Microscaled FP4 formats enable efficient FP4 deployment; however, fully quantizing weights, activations, and KV caches …
Accès ouvert
2026
preprint
OpenAlex
Janghwan Lee, S Lee, Jinseok Kim, Yong Wook Kim et autres
Large Reasoning Models (LRMs) achieve strong problem-solving through long chain-of-thought, but their deployment is constrained by the high cost of full-precision inference and growing KV cache footprints. Microscaled FP4 formats enable efficient FP4 deployment; however, fully quantizing weights, activations, and KV caches …
2026
conference-paper
OpenAlex
Chang-Hyo Yu, Jaewan Bae, Jinseok Kim, Hongyun Kim et autres
A 4nm-based quad-chiplet with an advanced packaged LLM accelerator achieving 56.8TPS on LLaMA v3.3 70B with single-batch 2k/2k input/output sequences. The architecture combines chiplet-based design, low-latency die-to-die interfaces, unified mixed-precision compute, holistic synchronization, and HBM3E with advanced power schemes to sustain bandwidth, …
kr
(code pays fourni par la source)
2025
article
OpenAlex
Yong Jun Jeong, He Young Kang, Jinwook Oh, Gwang‐Bok Kim et autres
In this paper, we present a high-PPF optical synapse transistor with an IGZO/PbS QD/Ga2O3structure for NIR sensing applications. The fabricated device exhibited superior NIR detection capabilities, with responsivities of 214.5 ± 4.1, 202.4 ± 5.6, and 188.5 ± 5.8 A/W and high …
kr
(code pays fourni par la source)
2025
conference-paper
OpenAlex
Hyunje Jo, Han-Sok Suh, Hyungseok Heo, Jinseok Kim et autres
This paper proposes a novel approach to accelerate large Neural Processing Unit (NPU) simulations on FPGA through Chain-based Time-Division Multiplexing (CTDM) and its automatic compiler. CTDM replaces repeated logic patterns with a single logic pattern and register chains, which can take advantage …
kr, us, ch
(code pays fourni par la source)
Accès ouvert
2025
conference-paper
OpenAlex
Janghwan Lee, Jiwoong Park, Jin-Seok Kim, Jungju Oh et autres
As large language models (LLMs) grow in parameter size and context length, computation precision has been reduced from 16-bit to 4bit to improve inference efficiency.However, this reduction causes accuracy degradation due to activation outliers.Rotation-based INT4 methods address this via matrix calibration, but …
kr, gb
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Janghwan Lee, Jiwoong Park, Jin-Seok Kim, Jungju Oh et autres
As large language models (LLMs) grow in parameter size and context length, computation precision has been reduced from 16-bit to 4-bit to improve inference efficiency. However, this reduction causes accuracy degradation due to activation outliers. Rotation-based INT4 methods address this via matrix …
2024
article
OpenAlex
Yong Jun Jeong, Gwang‐Bok Kim, Min Jae Kim, Jinwook Oh et autres
The development of broadband photosensors has become crucial in various fields. Indium–gallium–zinc oxide (IGZO, In:Ga:Zn = 1:1:1) phototransistors with PbS quantum dots (QDs) have shown promising features for such sensors, such as reasonable mobility, low leakage current, good photosensitivity, and low-cost fabrication. …
kr
(code pays fourni par la source)
2024
conference-paper
OpenAlex
Chang-Hyo Yu, Hyoeun Kim, Sungho Shin, Kyeongryeol Bong et autres
The growing computational demands of AI inference have led to widespread use of hardware accelerators for different platforms, spanning from edge to the datacenter/cloud. Certain AI application areas, such as in high-frequency trading (HFT) [1–2], have a hard inference latency deadline for …
2023
conference-paper
OpenAlex
Sungyeob Yoo, Hyunsung Kim, Jinseok Kim, Sunghyun Park et autres
Recent research shows that artificial intelligence (AI) algorithms can dramatically improve the profitability of high-frequency trading (HFT) with accurate market prediction, overcoming the limitation of conventional latency-oriented approaches. However, it is challenging to integrate the computationally intensive AI algorithm into the existing …
kr, ca, gb
(code pays fourni par la source)
2022
conference-paper
OpenAlex
Hyunsung Kim, Sungyeob Yoo, Jaewan Bae, Kyeongryeol Bong et autres
We present the world’s first AI-enabled high-frequency trading (HFT) system, LightTrader , which integrates the custom AI accelerators and the FPGA-based conventional HFT pipeline for the low-latency-high-throughput trading solutions with a reduced query miss rate. For better utilization, adaptive job scheduling methods …
gb, ca
(code pays fourni par la source)
2021
article
OpenAlex
Sae Kyu Lee, Ankur Agrawal, Joel A. Silberman, Matthew M. Ziegler et autres
Reduced precision computation is a key enabling factor for energy-efficient acceleration of deep learning (DL) applications. This article presents a 7-nm four-core mixed-precision artificial intelligence (AI) chip that supports four compute precisions—FP16, Hybrid-FP8 (HFP8), INT4, and INT2—to support diverse application demands for …
us, de, kr
(code pays fourni par la source)