Aller au contenu principal
Profil bibliographique

Mostafa Elhoushi

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

50Publications signalées
439Citations signalées
1Affiliations récentes

Les institutions déclarées

Les domaines associés

Advanced Neural Network ApplicationsNatural Language Processing TechniquesParallel Computing and Optimization TechniquesTopic ModelingIndoor and Outdoor Localization Technologies

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference

Mostafa Elhoushi, Alex Pretko, Nolan Dey, Bin Zhang et autres

Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout - particularly layer dropout - has largely disappeared …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Self-drafting with Uno on a mixture of experts: serving and training an adapter for Gemma 4 26B A4B

Subham Sekhar Sahoo, Lingjie Chen, Khiem Pham, Jonathan Geuter et autres

Large Language Models (LLMs) owe much of their success to next-token prediction (NTP), but their autoregressive (AR) structure requires slow, sequential token generation. To overcome this bottleneck, we introduce diffusion-augmented LLMs, a new class of models that defines an AR model distribution …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

Amr Hegazy, Amr Alanwar, Mostafa Elhoushi

Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM

Everlyn Asiko Chimoto, Mostafa Elhoushi, Bruce A. Bassett

Quantization is an effective technique for reducing the storage footprint and computational costs of Large Language Models (LLMs), but it often results in performance degradation. Existing post-training quantization methods typically use small, English-only calibration sets; however, their impact on multilingual models remains …

0 citations arXiv (Cornell University)
Accès ouvert 2026 preprint OpenAlex

Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM

Everlyn Asiko Chimoto, Mostafa Elhoushi, Bruce A. Bassett

Quantization is an effective technique for reducing the storage footprint and computational costs of Large Language Models (LLMs), but it often results in performance degradation. Existing post-training quantization methods typically use small, English-only calibration sets; however, their impact on multilingual models remains …

us (code pays fourni par la source)

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls

Feiyang Kang, Newsha Ardalani, Michael Kuchnik, Youssef Emad et autres

Training data plays a crucial role in Large Language Models (LLM) scaling, yet high quality data is of limited supply. Synthetic data techniques offer a potential path toward sidestepping these limitations. We conduct a large-scale empirical investigation (>1000 LLMs with >100k GPU …

0 citations arXiv (Cornell University)
2025 article OpenAlex

Characterizing and Efficiently Accelerating Multimodal Generation Model Inference

Alicia Golden, Anna Sun, Basil Hosmer, Bilge Acun et autres

Generative artificial intelligence (AI) technology is revolutionizing the computing industry, posing new system design and optimization opportunities. In particular, AI’s ability to understand and respond in multiple modalities comes with significant system resource demands. To sustainably scale generative AI capabilities to billions …

us, gb (code pays fourni par la source)

5 citations IEEE Micro
Accès ouvert 2025 preprint OpenAlex

any4: Learned 4-bit Numeric Representation for LLMs

Mostafa Elhoushi, Jeff Johnson

We present any4, a learned 4-bit weight quantization solution for large language models (LLMs) providing arbitrary numeric representations without requiring pre-processing of weights or activations. any4 yields higher accuracy compared to other related 4-bit numeric representation types: int4, fp4 and nf4, as …

1 citation arXiv (Cornell University)
2025 conference-paper OpenAlex

PETAH: Parameter Efficient Task Adaptation for Hybrid Transformers

Maximilian Augustin, Syed Shakib Sarwar, Mostafa Elhoushi, Yuecheng Li et autres

Transformers have revolutionized natural language processing (NLP) and are increasingly influential in computer vision tasks. Despite their strong performance and multitasking capabilities, transformers' high computational demands limit their applicability in resource-constrained environments, where convolutional or hybrid models (combining convolution and attention layers) …

de, us (code pays fourni par la source)

0 citations
2025 conference-paper OpenAlex

Semi-Structured Sparsity Using Dynamic Percentile

Omar Sabra, Mostafa Elhoushi, Mohamed Taher, M. Watheq El‐Kharashi

The computational efficiency of deep neural networks is limited by hardware accelerators and sparsity integration. Leveraging sparsity can result in significant computational savings by reducing the number of operations and memory accesses required. In recent years, N:M sparsity, where N out of …

Égypte (code pays fourni par la source)

0 citations

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.