Accès ouvert
2026
preprint
OpenAlex
Nipun Ghanghas, Siddharth Dhanpal, Shravan Hanasoge, Praneeth Netrapalli et autres
Gravity-mode period spacings (DPi_1) of red giants probe the stellar core directly, constraining its structure, mass and evolutionary state. Their measurement requires resolving narrow, densely spaced mixed modes and has so far relied on the four-year baseline of Kepler. Recovering DPi_1 from …
Accès ouvert
2026
preprint
OpenAlex
Nipun Ghanghas, Siddharth Dhanpal, Shravan Hanasoge, Praneeth Netrapalli et autres
Gravity-mode period spacings (DPi_1) of red giants probe the stellar core directly, constraining its structure, mass and evolutionary state. Their measurement requires resolving narrow, densely spaced mixed modes and has so far relied on the four-year baseline of Kepler. Recovering DPi_1 from …
in, us
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Suvrat Raju, Praneeth Netrapalli
We study the error rate of LLMs on tasks like arithmetic that require a deterministic output, and repetitive processing of tokens drawn from a small set of alternatives. We argue that incorrect predictions arise when small errors in the attention mechanism accumulate …
Accès ouvert
2026
preprint
OpenAlex
Suvrat Raju, Praneeth Netrapalli
We study the error rate of LLMs on tasks like arithmetic that require a deterministic output, and repetitive processing of tokens drawn from a small set of alternatives. We argue that incorrect predictions arise when small errors in the attention mechanism accumulate …
Accès ouvert
2026
dataset
OpenAlex
Suvrat Raju, Praneeth Netrapalli
This dataset accompanies the paper "A model of errors in transformers." Filenames The data is provided in 25 csv files. Each filename comprises a model name and a task name, as specified in the main text. The acronyms for models are the …
in
(code pays fourni par la source)
Accès ouvert
2026
dataset
OpenAlex
Suvrat Raju, Praneeth Netrapalli
This dataset accompanies the paper "A model of errors in transformers." Filenames The data is provided in 25 csv files. Each filename comprises a model name and a task name, as specified in the main text. The acronyms for models are the …
in
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Devvrit Khatri, Pranamya Kulkarni, Nilesh Gupta, Yerram Varun et autres
Large Language Models (LLMs) have been shown to be able to learn different tasks without explicit finetuning when given many input-output examples / demonstrations through In-Context Learning (ICL). Increasing the number of examples, called ``shots'', improves downstream task performance but incurs higher …
2025
conference-paper
OpenAlex
Jatin Alla, Yashas Samaga, Ashwin Vaswani, Praneeth Netrapalli et autres
us
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Chong You, Kan Wu, Z. Jiao, Lin Chen et autres
The discovery of the lazy neuron phenomenon in trained Transformers, where the vast majority of neurons in their feed-forward networks (FFN) are inactive for each token, has spurred tremendous interests in activation sparsity for enhancing large model efficiency. While notable progress has …
Accès ouvert
2025
preprint
OpenAlex
Yashas Samaga, Varun Yerram, Spandana Raj Babbula, Prateek Jain et autres
We consider the Top-$K$ selection problem, which aims to identify the largest $K$ elements in an array. Top-$K$ selection arises in many machine learning algorithms and often becomes a bottleneck on accelerators, which are optimized for dense matrix multiplications. To address this …
2025
conference-paper
OpenAlex
Chong You, Kan Wu, Zhipeng Jia, Lin Chen et autres
us, cn, jp, gb
(code pays fourni par la source)
2024
conference-paper
OpenAlex
Ashwin Vaswani, Yashas Samaga, Gaurav Aggarwal, Praneeth Netrapalli et autres
us
(code pays fourni par la source)