Difference between revisions of "IC4R003-Epigenomic-2011-21984925"

From RiceWiki
Jump to: navigation, search
(Corresponding Author)
(The Background of This Project)
Line 2: Line 2:
 
* '''Comparison of Four ChIP-Seq Analytical Algorithms Using Rice Endosperm H3K27 Trimethylation Profiling Data'''
 
* '''Comparison of Four ChIP-Seq Analytical Algorithms Using Rice Endosperm H3K27 Trimethylation Profiling Data'''
 
==The Background of This Project==
 
==The Background of This Project==
* Roughly 150 million y ago, flowering plants diverged to form the two dominant extant lineages, monocots and dicots (1). Arabidopsis thaliana, the preeminent plant genetic system, is a dicot, whereas cereal crops, such as rice, wheat, and maize, that feed much of the world are monocots. In both plant groups, pollen grains contain two sperm nuclei, one of which fertilizes a diploid central cell to give rise to triploid endosperm (2). A. thaliana endosperm is consumed by the developing embryo, whereas cereal endosperm persists and makes up the bulk of the mature seed— a developmental difference of particular practical importance (3). Developing seeds are genetic battlegrounds on multiple fronts: parents are proposed to be in conflict over resource allocation (2), whereas the embryo must repress parasitic transposable elements (TEs) to prevent damage to the genome.
+
* Chromatin immunoprecipitation (ChIP) coupled with high throughput sequencing (ChIP-Seq) has emerged as one of the most promising tools for profiling protein-DNA binding sites and chromatin modifications on a genome-wide scale [1]. The goal of ChIP-Seq studies is to find those genomic DNA fragments that are enriched in immunoprecipitation fractions using antibodies specific for DNA associated proteins of interest. Enriched regions, those with a high density of short DNA reads after immunopre- cipitation and DNA sequencing, are referred to as peaks. Many programs for identification of peaks with ChIP-Seq data have been developed in recent years [2,3,4,5,6,7,8,9,10,11]. The reported algorithms differ in their approaches for identifying potential enriched regions of the genome. Some algorithms, for example MACS [10] and PeakSeq [8], use a simple sliding window and group all reads within each window together. Others use a finer resolution method, either considering each base pair singly as in FindPeaks [4] or defining the windows based on the read locations as represented by USeq [7]. After identifying windows, the algorithms must then determine which windows are the true enriched regions. Methods without a control (FindPeaks) either simply report the number of reads in the windows or make an assumption about the background distribution, such as assuming the reads follow a Poisson distribution (FindPeaks), and calculate significance based on the assumed distribution.
* Most of our knowledge about DNA methylation in plant seeds is derived from A. thaliana. Processes involving genetic conflict tend to evolve rapidly (9), and therefore, methylation dynamics in cereal seeds may be quite different. In this project, the researchers use deep bisulfite sequencing to examine DNA methylation in rice seeds. Wild-type rice endosperm methylation patterns—globally reduced non-CG methylation and local CG hypomethylation—resemble those of DME-deficient A. thaliana endosperm, a finding consistent with lack of DME in monocots. Reduced endosperm methylation is common in genes with preferential endosperm expression, in- dicating that demethylation is a major mechanism for gene activation in rice endosperm. Short TEs are hypermethylated at CHH sites in embryo, suggesting that endosperm demethylation func- tions to immunize the embryo against TEs through small RNAs.
+
* Those including a control sample (MACS, PeakSeq) use the control to more accurately model the background distribution of the reads and calculate an empirical False Discovery Rate (FDR) via, for example, a sample swap technique. Distinguishing between multiple small peaks or a single large peak is also challenging. While some algorithms merge overlapping peaks (MACS) or peaks within a user-supplied threshold (USeq, PeakSeq), others (Find-Peaks) compare the height of peaks to the depth of the separating valley to differentiate multiple small peaks from one large peak. Pepke et al. [12] discussed a number of additional peak identification algorithms in a review article. They made distinc- tions among the algorithms, including how the algorithms aggregated the reads, the criteria for significant peak identification, read shifting to account for reading the end of the reads, use of control, and input parameters. Similarly, Barski and Zhao [13] also reviewed a number of algorithms for peak identification. Thus far, however, no program has emerged as the consensus best approach for identifying peaks in histone modification and DNA binding studies. Therefore, it is important to compare these available algorithms and to suggest essential parameters to assist molecular biology laboratories in selecting the best program for their data analysis.
 +
* In this project, the researchers identified H3K27me3 modification sites within rice (Oryza sativa) young endosperm using the ChIP-Seq approach. Four different peak identification algorithms (PeakSeq, USeq, MACS, and FindPeaks) were used to locate H3K27me3 enrichment sites. ChIP-PCR was used to evaluate the quality of the peaks identified by these algorithms. We also analyzed the relative location of the peaks with respect to gene expression. Finally, we examined the Gene Ontology (GO) annotations [27] of the ChIP enriched genes.
  
 
==Labs working on this Project==
 
==Labs working on this Project==

Revision as of 04:27, 22 June 2016

Project Title

  • Comparison of Four ChIP-Seq Analytical Algorithms Using Rice Endosperm H3K27 Trimethylation Profiling Data

The Background of This Project

  • Chromatin immunoprecipitation (ChIP) coupled with high throughput sequencing (ChIP-Seq) has emerged as one of the most promising tools for profiling protein-DNA binding sites and chromatin modifications on a genome-wide scale [1]. The goal of ChIP-Seq studies is to find those genomic DNA fragments that are enriched in immunoprecipitation fractions using antibodies specific for DNA associated proteins of interest. Enriched regions, those with a high density of short DNA reads after immunopre- cipitation and DNA sequencing, are referred to as peaks. Many programs for identification of peaks with ChIP-Seq data have been developed in recent years [2,3,4,5,6,7,8,9,10,11]. The reported algorithms differ in their approaches for identifying potential enriched regions of the genome. Some algorithms, for example MACS [10] and PeakSeq [8], use a simple sliding window and group all reads within each window together. Others use a finer resolution method, either considering each base pair singly as in FindPeaks [4] or defining the windows based on the read locations as represented by USeq [7]. After identifying windows, the algorithms must then determine which windows are the true enriched regions. Methods without a control (FindPeaks) either simply report the number of reads in the windows or make an assumption about the background distribution, such as assuming the reads follow a Poisson distribution (FindPeaks), and calculate significance based on the assumed distribution.
  • Those including a control sample (MACS, PeakSeq) use the control to more accurately model the background distribution of the reads and calculate an empirical False Discovery Rate (FDR) via, for example, a sample swap technique. Distinguishing between multiple small peaks or a single large peak is also challenging. While some algorithms merge overlapping peaks (MACS) or peaks within a user-supplied threshold (USeq, PeakSeq), others (Find-Peaks) compare the height of peaks to the depth of the separating valley to differentiate multiple small peaks from one large peak. Pepke et al. [12] discussed a number of additional peak identification algorithms in a review article. They made distinc- tions among the algorithms, including how the algorithms aggregated the reads, the criteria for significant peak identification, read shifting to account for reading the end of the reads, use of control, and input parameters. Similarly, Barski and Zhao [13] also reviewed a number of algorithms for peak identification. Thus far, however, no program has emerged as the consensus best approach for identifying peaks in histone modification and DNA binding studies. Therefore, it is important to compare these available algorithms and to suggest essential parameters to assist molecular biology laboratories in selecting the best program for their data analysis.
  • In this project, the researchers identified H3K27me3 modification sites within rice (Oryza sativa) young endosperm using the ChIP-Seq approach. Four different peak identification algorithms (PeakSeq, USeq, MACS, and FindPeaks) were used to locate H3K27me3 enrichment sites. ChIP-PCR was used to evaluate the quality of the peaks identified by these algorithms. We also analyzed the relative location of the peaks with respect to gene expression. Finally, we examined the Gene Ontology (GO) annotations [27] of the ChIP enriched genes.

Labs working on this Project

  • Department of Computer Science and Engineering, Mississippi State University, Mississippi, United States of America,
  • Institute for Genomics, Biocomputing, and Biotechnology, Mississippi State University, Mississippi, United States of America
  • Department of Biochemistry and Molecular Biology, Mississippi State University, Mississippi, United States of America

Corresponding Author

  • Zhaohua Peng (E-mail: zp7@BCH.msstate.edu) & Susan M. Bridges (E-mail: bridges@cse.msstate.edu)