Accès ouvert
2026
preprint
OpenAlex
Hongkang Yang, Zhewei Xu, Feiyu Xiong, E Weinan
As part of a series on first-principles modeling of cognitive functions, this paper attempts to provide a mathematical formulation of thinking and perception. It formally derives slow thinking or more generally, active perception, and encompasses the design, training and inference of slow …
Accès ouvert
2026
preprint
OpenAlex
Liangkai Hang, Junjie Yao, Zhiyu Li, Feiyu Xiong et autres
Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered. Although progress is usually attributed to scale, data and architecture, we show that parameter initialization is a gene-like determinant of training …
Accès ouvert
2026
preprint
OpenAlex
Liangkai Hang, Junjie Yao, Zhiyu Li, Feiyu Xiong et autres
Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered. Although progress is usually attributed to scale, data and architecture, we show that parameter initialization is a gene-like determinant of training …
cn, at, kp
(code pays fourni par la source)
Accès ouvert
2026
article
OpenAlex
Hongkang Yang, Z Xu, Feiyu Xiong, E Weinan
As part of a series on first-principles modeling of cognitive functions, this paper attempts to provide a mathematical formulation of thinking and perception. It formally derives slow thinking or more generally, active perception, and encompasses the design, training and inference of slow …
at, cn
(code pays fourni par la source)
Accès ouvert
2026
article
OpenAlex
Hongkang Yang, Zhiqin John Xu, Feiyu Xiong, Weinan Ee
Accès ouvert
2025
preprint
OpenAlex
Zhiyu Li, Chunyan Xi, Chunyu Li, Shichao Song et autres
Large Language Models (LLMs) have become an essential infrastructure for Artificial General Intelligence (AGI), yet their lack of well-defined memory management systems hinders the development of long-context reasoning, continual personalization, and knowledge consistency.Existing models mainly rely on static parameters and short-lived contextual …
Accès ouvert
2025
preprint
OpenAlex
Zhiwei Bai, Zhangchen Zhou, Jiajie Zhao, Xiaolong Li et autres
Loss spikes commonly emerge during neural network training with the Adam optimizer across diverse architectures and scales, yet their underlying mechanism remains elusive. While previous explanations attribute these phenomena to sharper loss landscapes at lower loss, we show that landscape geometry alone …
Accès ouvert
2025
preprint
OpenAlex
Liangkai Hang, Junjie Yao, Zhiwei Bai, Tianyi Chen et autres
The reasoning ability of large language models (LLMs) has been rapidly advancing in recent years, attracting interest in more fundamental approaches that can reliably enhance their generalizability. This work demonstrates that model complexity control, conveniently implementable by adjusting the initialization rate and …
Accès ouvert
2025
preprint
OpenAlex
Zhiyu Li, Shichao Song, Hanyu Wang, Simin Niu et autres
Large Language Models (LLMs) have emerged as foundational infrastructure in the pursuit of Artificial General Intelligence (AGI). Despite their remarkable capabilities in language perception and generation, current LLMs fundamentally lack a unified and structured architecture for handling memory. They primarily rely on …
Accès ouvert
2024
article
OpenAlex
Hongkang Yang, Zehao Lin, Wenjin Wang, Hao Wu et autres
The training and inference of large language models (LLMs) are together a costly process that transports knowledge from raw data to meaningful computation. Inspired by the memory hierarchy of the human brain, we reduce this cost by equipping LLMs with explicit memory, …
cn
(code pays fourni par la source)
Accès ouvert
2024
preprint
OpenAlex
Hongkang Yang, Zehao Lin, Wenjin Wang, Hao Wu et autres
The training and inference of large language models (LLMs) are together a costly process that transports knowledge from raw data to meaningful computation. Inspired by the memory hierarchy of the human brain, we reduce this cost by equipping LLMs with explicit memory, …
Accès ouvert
2023
article
OpenAlex
Cheng Wang, Hongkang Yang, Hongting Hua, Yaosuo Xue
Dual-T-type modular multilevel converter (DTMMC) is an alternative to the existing uninterrupted power supply (UPS) due to its advantage in flexible leg reusing, multiple power ports, and high operating efficiency. In this article, a DTMMC-UPS, which has series configurations on both input/output …
cn, us
(code pays fourni par la source)