Aller au contenu principal
Accès ouvert déclaré 2026 article

A unified lexicon of predictive DNA sequence motifs from ENCODE transcription factor binding and chromatin accessibility assays

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : us. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Abstract: Regulatory DNA contains sequence elements that guide transcription factor (TF) binding and chromatin accessibility, that in turn control the expression of nearby genes. Most current descriptions of such regulatory elements are based on classical, statistical enrichment-based motif discovery methods applied to TF binding signals. Here, we present ENCODE GRAMMAR (Genomic Regulatory Atlas of sequence Models, Motifs, Annotations and Rules): one of the largest collections of regulatory DNA deep learning models trained on TF binding and chromatin accessibility to date, and ENCODE MotifCompendium: the first comprehensive lexicon of predictive regulatory motifs derived from the models. For ENCODE GRAMMAR, we trained BPNet and ChromBPNet models across 2,339 TF ChIP-seq datasets in 788 TFs, 1,143 DNase-seq datasets in 287 samples, and 369 ATAC-seq datasets in 204 samples. From the models, we extracted 286,836 sequence motifs that quantitatively predict regulatory signal in the cellular context of each dataset. To consolidate the discovered motifs into a single, non-redundant union across all cell contexts, we developed MotifCompendium—a GPU-accelerated motif management package, that can perform an accelerated calculation of motif-optimized pairwise similarity between 10,000 motifs in just 6 seconds on a 12GB GPU, flag for noisy, undesirable motifs, cluster the motifs, and provide other utilities to facilitate motif analyses. Using MotifCompendium, we consolidated the discovered motifs into a single, unified lexicon of 3,384 motifs that are predicted to drive TF binding and chromatin accessibility across cell contexts (ENCODE MotifCompendium). From ENCODE MotifCompendium, we observed that motifs from chromatin accessibility can be highly orthogonal to those from TF ChIP-seq, especially in different cell contexts (primary cells vs. cell lines). The lexicon also shows substantial overlap with existing TF binding motif databases, recapitulating, on average, ~93% of TF motifs from previous curated databases, while 10% of the lexicon represent entirely novel binding modes not captured by existing databases. We identify previously uncharacterized zinc finger-like motifs, composite motifs with two or more sub-motifs in preferential spacing, and cell-type-specific variants of motifs. This work provides a foundational resource of predictive models, a scalable computational framework for extracting sequence features, and the first unified lexicon of model-derived regulatory motifs.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Genomics and Chromatin DynamicsMachine Learning in BioinformaticsGene Regulatory Network Analysis

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.