Aller au contenu principal
Profil bibliographique

Utku Evci

Informations fournies par OpenAlex. Research Africa ne déduit ni nationalité, ni poste, ni coordonnées personnelles.

33Publications signalées
810Citations signalées
0Affiliations récentes

Les domaines associés

Domain Adaptation and Few-Shot LearningAdvanced Neural Network ApplicationsMachine Learning and Data ClassificationStochastic Gradient Optimization TechniquesMultimodal Machine Learning Applications

Les publications récentes

Accès ouvert 2026 preprint OpenAlex

Gemma 4 Technical Report

Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres et autres

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside …

3 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

Spark Transformer: Reactivating Sparsity in FFN and Attention

Chong You, Kan Wu, Z. Jiao, Lin Chen et autres

The discovery of the lazy neuron phenomenon in trained Transformers, where the vast majority of neurons in their feed-forward networks (FFN) are inactive for each token, has spurred tremendous interests in activation sparsity for enhancing large model efficiency. While notable progress has …

0 citations arXiv (Cornell University)
Accès ouvert 2025 preprint OpenAlex

The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws

Jin Tian, Ahmed Imtiaz Humayun, Utku Evci, Suvinay Subramanian et autres

Pruning eliminates unnecessary parameters in neural networks; it offers a promising solution to the growing computational demands of large language models (LLMs). While many focus on post-training pruning, sparse pre-training--which combines pruning and pre-training into a single phase--provides a simpler alternative. In …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

Learning Fine-grained Parameter Sharing via Sparse Tensor Decomposition

Cem Üyük, Mike Lasby, Mohamed Yassin, Utku Evci et autres

Large neural networks achieve state-of-the-art performance on many tasks, yet their sheer size hinders deployment on resource-constrained devices. Among existing compression approaches, cross-layer parameter sharing remains relatively unexplored for transformer models. In this paper, we introduce Fine-grained Parameter Sharing (FiPS), a unified …

1 citation arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

Towards Optimal Adapter Placement for Efficient Transfer Learning

Aleksandra Nowak, Otniel-Bogdan Mercea, Anurag Arnab, Jonas Pfeiffer et autres

Parameter-efficient transfer learning (PETL) aims to adapt pre-trained models to new downstream tasks while minimizing the number of fine-tuned parameters. Adapters, a popular approach in PETL, inject additional capacity into existing networks by incorporating low-rank projections, achieving performance comparable to full fine-tuning …

0 citations arXiv (Cornell University)
Accès ouvert 2024 preprint OpenAlex

Progressive Gradient Flow for Robust N:M Sparsity Training in Transformers

Abhimanyu Rajeshkumar Bambhaniya, Amir Yazdanbakhsh, Suvinay Subramanian, Sheng-Chun Kao et autres

N:M Structured sparsity has garnered significant interest as a result of relatively modest overhead and improved efficiency. Additionally, this form of sparsity holds considerable appeal for reducing the memory footprint owing to their modest representation overhead. There have been efforts to develop …

2 citations arXiv (Cornell University)
Accès ouvert 2023 preprint OpenAlex

Dynamic Sparse Training with Structured Sparsity

Mike Lasby, А. В. Голубева, Utku Evci, Mihai Nica et autres

Dynamic Sparse Training (DST) methods achieve state-of-the-art results in sparse neural network training, matching the generalization of dense models while enabling sparse training and inference. Although the resulting models are highly sparse and theoretically less computationally expensive, achieving speedups with unstructured sparsity …

4 citations arXiv (Cornell University)
Accès ouvert 2023 preprint OpenAlex

JaxPruner: A concise library for sparsity research

Joo Hyung Lee, Wonpyo Park, Nicole Mitchell, Jonathan Pilault et autres

This paper introduces JaxPruner, an open-source JAX-based pruning and sparse training library for machine learning research. JaxPruner aims to accelerate research on sparse neural networks by providing concise implementations of popular pruning and sparse training algorithms with minimal memory and latency overhead. …

1 citation arXiv (Cornell University)

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.