Transcription factor regulation and epigenetic modifiers
Le résumé fourni par la source
Transcription factors bind DNA in specific sequence contexts. Their binding to particular motifs, leads to downstream control of genic expression and cell fate specification. Understanding the particular motifs they bind is therefore of great importance, but is complicated by experimental artifacts, different factors involved in their binding, and the presence of epigenetic modifications. Both ChIP-sequencing (ChIP-seq) and more recently cleavage under targets and release using nuclease (CUT&RUN) are used to find transcription factor binding sites, but both suffer from various sources of noise. This can be mitigated by judicious use of controls, careful bioinformatic processing, and replicated analyses and comparisons. One way to reduce noise is to analyze a factor in which the target has been knocked-out, as a negative control. Paired wild-type and knockout experiments can generate improved motifs but require optimal differential analysis. I introduce peaKO—a computational method to automatically optimize motif analyses with knockout controls, which we compare to two other methods. Going beyond conventional transcription factor analyses, in addition to distinguishing one nucleobase from another, some transcription factors can distinguish between unmodified and modified bases. Current models of transcription factor binding tend not to take DNA modifications into account, while the recent few that do often have limitations. This makes a comprehensive and accurate profiling of transcription factor affinities difficult. Here, I develop methods to identify transcription factor binding sites in modified DNA. My models expand the standard A/C/G/T DNA alphabet to include cytosine modifications. I develop Cytomod to create modified genomic sequences and we also enhance the MEME Suite, adding the capacity to handle custom alphabets. I adapt the well-established position weight matrix (PWM) model of transcription factor binding affinity to this expanded DNA alphabet. Using these methods, I identify modification-sensitive transcription factor binding motifs. I validate my binding preference predictions for OCT4 using CUT&RUN experiments across conventional, methylation-enriched, and hydroxymethylation-enriched sequences. My approach extends to other increasingly investigated DNA modifications, across diverse organisms. In eukaryotes, both 5-methylcytosine (5mC) and 5-hydroxymethylcytosine (5hmC) are recognized as stable epigenetic marks, with diverse functions. Bacteria, archaea, and viruses contain other modified DNA nucleobases. Numerous databases describe RNA, but none specifically catalogue DNA modifications, despite their broad regulatory importance. To address this need, I developed DNAmod: the DNA modification database, providing a useful publicly-available resource. This work serves to enhance our understanding of transcriptional regulation, by improving the ability to elucidate and analyze transcription factor binding sites in a variety of contexts. Through improved methods for the analysis of transcription factor binding data, creation of methods to explore the impact of epigenetic modifications, and a database to curate these various modifications, I improve our ability to investigate transcription factor binding. As more genome-wide single-base resolution modification data becomes available, I expect that my methods will allow researchers to select from improved methods to assess transcription factor binding, and decide how to integrate diverse modified nucleobase datasets into their analyses. This will in turn yield greater insights into altered binding affinities across many different modifications, and potentially with diverse sequence methods, across a multitude of different organisms.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.