IC4R004-Epigenomic-2012-22778444

From RiceWiki
Revision as of 04:30, 22 June 2016 by Xysj1990 (talk | contribs) (Created page with "==Project Title== * '''Transcriptome and methylome interactions in rice hybrids''' ==The Background of This Project== * Chromatin immunoprecipitation (ChIP) coupled with high ...")
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to: navigation, search

Project Title

  • Transcriptome and methylome interactions in rice hybrids

The Background of This Project

  • Chromatin immunoprecipitation (ChIP) coupled with high throughput sequencing (ChIP-Seq) has emerged as one of the most promising tools for profiling protein-DNA binding sites and chromatin modifications on a genome-wide scale [1]. The goal of ChIP-Seq studies is to find those genomic DNA fragments that are enriched in immunoprecipitation fractions using antibodies specific for DNA associated proteins of interest. Enriched regions, those with a high density of short DNA reads after immunopre- cipitation and DNA sequencing, are referred to as peaks. Many programs for identification of peaks with ChIP-Seq data have been developed in recent years [2,3,4,5,6,7,8,9,10,11]. The reported algorithms differ in their approaches for identifying potential enriched regions of the genome. Some algorithms, for example MACS [10] and PeakSeq [8], use a simple sliding window and group all reads within each window together. Others use a finer resolution method, either considering each base pair singly as in FindPeaks [4] or defining the windows based on the read locations as represented by USeq [7]. After identifying windows, the algorithms must then determine which windows are the true enriched regions. Methods without a control (FindPeaks) either simply report the number of reads in the windows or make an assumption about the background distribution, such as assuming the reads follow a Poisson distribution (FindPeaks), and calculate significance based on the assumed distribution.
  • Those including a control sample (MACS, PeakSeq) use the control to more accurately model the background distribution of the reads and calculate an empirical False Discovery Rate (FDR) via, for example, a sample swap technique. Distinguishing between multiple small peaks or a single large peak is also challenging. While some algorithms merge overlapping peaks (MACS) or peaks within a user-supplied threshold (USeq, PeakSeq), others (Find-Peaks) compare the height of peaks to the depth of the separating valley to differentiate multiple small peaks from one large peak. Pepke et al. [12] discussed a number of additional peak identification algorithms in a review article. They made distinc- tions among the algorithms, including how the algorithms aggregated the reads, the criteria for significant peak identification, read shifting to account for reading the end of the reads, use of control, and input parameters. Similarly, Barski and Zhao [13] also reviewed a number of algorithms for peak identification. Thus far, however, no program has emerged as the consensus best approach for identifying peaks in histone modification and DNA binding studies. Therefore, it is important to compare these available algorithms and to suggest essential parameters to assist molecular biology laboratories in selecting the best program for their data analysis.
  • In this project, the researchers identified H3K27me3 modification sites within rice (Oryza sativa) young endosperm using the ChIP-Seq approach. Four different peak identification algorithms (PeakSeq, USeq, MACS, and FindPeaks) were used to locate H3K27me3 enrichment sites. ChIP-PCR was used to evaluate the quality of the peaks identified by these algorithms. We also analyzed the relative location of the peaks with respect to gene expression. Finally, we examined the Gene Ontology (GO) annotations [27] of the ChIP enriched genes.

Labs working on this Project

  • Department of Molecular, Cell and Developmental Biology, University of California, Los Angeles, CA 90095; b Howard Hughes Medical Institute,
  • University of California, Los Angeles, CA 90095; f Molecular Biology Institute, University of California, Los Angeles, CA 90095; c Department of Plant
  • Pathology, Ohio State University, Columbus, OH 43210; d Department of Plant and Soil Sciences, Delaware Biotechnology Institute, University of Delaware,
  • Newark, DE 19711; e US Department of Agriculture—Agricultural Research Service Dale Bumpers National Rice Research Center, Stuttgart, AR 72160;
  • Eli and Edythe Broad Center of Regenerative Medicine and Stem Cell Research, University of California, Los Angeles, CA 90095

Corresponding Author

  • Steven E. Jacobsen (E-mail:jacobsen@ucla.edu) & Matteo Pellegrini (E-mail: matteop@mcdb.ucla.edu)