Difference between revisions of "IC4R002-GWAS-2011-21915109"

From RiceWiki
Jump to: navigation, search
(Created page with "== Project Title == ''' Genetic Architecture of Aluminum Tolerance in Rice( Oryza sativa ) Determined through Genome-Wide Association Analysis and QTL Mapping ''' ==The Back...")
 
(The Background of This Project)
 
(13 intermediate revisions by the same user not shown)
Line 1: Line 1:
 
== Project Title ==
 
== Project Title ==
''' Genetic Architecture of Aluminum Tolerance in Rice( Oryza sativa ) Determined through Genome-Wide Association Analysis and QTL Mapping '''
+
''' Genome-wide association mapping reveals a rich genetic architecture of complex traits in Oryza sativa  '''
  
 
==The Background of This Project==
 
==The Background of This Project==
[[File:gwas-1.PNG|700px|thumb|right|'''Figure 1.''' '' GWA Analysis of Al Tolerance within and across Rice Subpopulations.'']]
+
* Asian rice, Oryza sativa is a cultivated, inbreeding species that feeds over half of the world's population. Understanding the genetic basis of diverse physiological, developmental, and morphological traits provides the basis for improving yield, quality and sustainability of rice. Here we show the results of a genome-wide association study based on genotyping 44,100 SNP variants across 413 diverse accessions of O. sativa collected from 82 countries that were systematically phenotyped for 34 traits. Using cross-population-based mapping strategies, we identifi ed dozens of common variants infl uencing numerous complex traits. Signifi cant heterogeneity was observed in the genetic architecture associated with subpopulation structure and response to environment. This work establishes an open-source translational research platform for genome-wide association studies in rice that directly links molecular variation in genes and metabolic pathways with the germplasm resources needed to accelerate varietal development and crop improvement.
* While rice (Oryza sativa) is significantly more Al tolerant than other cereals, no genes underlying Al tolerance in rice have been reported. Using genome-wide association(GWA) and bi-parental QTL mapping, we investigated the genetic architecture of Al tolerance in rice. Japonica varieties were twice as Al tolerant as indica and aus varieties. Overall, 57% of the phenotypic variation was correlated with subpopulation, consistent with observations that different genes and genomic regions were associated with Al tolerance in different subpopulations. Four regions identified by GWA co-localized with a priori candidate genes, and two highly significant regions co-localized with previously identified quantitative trait loci(QTL). Haplotype and sequence analysis around the candidate gene, Nrat1, identified a susceptible haplotype explaining 40% of the Al tolerance variation within the aus subpopulation and three non-synonymous mutations within Nrat1 that were predictive of Al sensitivity. Using Indica 6 Japonica mapping populations, we identified QTLs associated with transgressive variation where alleles from a susceptible indica or aus parent enhanced Al tolerance in a tolerant japonica background. This work demonstrates the importance of subpopulation in interpreting and manipulating complex traits in rice and provides a roadmap for breeders aiming to capture genetic value from phenotypically inferior lines.
 
  
 
==Plant Culture & Treatment==
 
==Plant Culture & Treatment==
* Plants were grown hydroponically in a growth chamber as described by Famoso et al. Al tolerance was determined based on relative root growth (RRG) after three days in Al (160 m M Al 3+ ) or control solution. The hydroponic solution used in this study was chemically designed and optimized for rice Al tolerance screening; for a detailed comparison of the phenotypic procedures employed in this work compared to previously published rice Al tolerance work see Famoso et al. (2010). To obtain uniform seedlings, 80 seeds were germinated and the 30 most uniform seedlings were visually selected and transferred to a control hydroponic solution for a 24 hour adjustment period. After the 24 hour adjustment period, root length was measured with a ruler and the 20 most uniform seedlings were selected and distributed to fresh control solution (0 uM Al 3+ ) or Al treatment solution (160 uM Al 3+ ). Plants were grown in their respective treatments for ,72 hours and the total root system growth was quantified using an imaging and root quantification system as described by Famoso et al.(2010). The mean total root growth was calculated for Al treated and control plants and RRG was calculated as mean growth (Al)/mean growth (control). The 373 genotypes screened for Al tolerance and used in the association analysis are part of a set of 400 O. sativa genotypes that have been genotyped with 44,000 SNPs as described by Zhao et al.
+
* The Rice Diversity Panel consists of 413 Asian rice ( O. sativa ) cultivars, including many landraces, which originated from 82 countries, representing all the major rice-growing regions of the world. The panel contains 87 indica , 57 aus , 96 temperate japonica , 97 tropical japonica , 14 groupV/aromatic, and 62 highly admixed accessions. All accessions were purifi ed for two generations(single seed descent) before DNA extraction. In all, 20 of these 413 accessions were purifi ed as part of the Oryza SNP project. Six cultivars (Azucena, Moroberekan, Nipponbare, Dom-Sofi d, IR64, M-202) were purifi ed separately, once by Ali et al. and once as part of the Oryza SNP panel. Further information for each accession(accession name, accession number, country of origin and subpopulation ancestry based on PCA) is given in Supplementary Data 1 .
  
 
==Research Findings==
 
==Research Findings==
* Two immortalized QTL mapping populations were analyzed for Al tolerance. One consisted of 134 recombinant inbred lines (RIL) derived from the cross IR64/Azucena , and the other was comprised of 78 backcross inbred lines (BIL) derived from the cross Nipponbare/Kasalath//Nipponbare. These populations were used to evaluate Al tolerance using three different indices of relative root growth (RRG), (1) longest root growth (LRG-RRG), (2) primary root growth (PGR-RRG) and total root growth (TRG-RRG) (see Materials and Methods for details). The phenotypic distribution was approximately normal for each population, no matter which root screening index was used. The QTL mapping populations allowed us to determine which of the three root evaluation methods would be most useful for evaluating the diversity panel as a whole.<br><br>
+
* A rice diversity panel consisting of 413 inbred accessions of O. sativa collected from 82 countries was genotyped using an Aff ymetrix single nucleotide polymorphism (SNP) array containing 44,100 SNPs (hereaft er referred to as the 44 K chip). With a genome size of ~ 380 Mb (ref. 13), this custom-designed genotyping chip provides high quality data (less than 4.5 % missing data), with ~ 1 SNP per 10 kb across the 12 chromosomes of rice. The diversity panel was evaluated for 34 traits related to plant morphology, grain quality, plant development and agronomic performance using fi eld-grown plants with replications within and between years.<br><br>
* To identify Al tolerance loci based on genome-wide association(GWA) mapping, we used an existing genotypic dataset consisting of 36,901 SNPs, and the total root growth (TRG-RRG) Al tolerance phenotype generated on 373 O. sativa accessions over the course of this study. GWA mapping was conducted, using SNPs with a MAF.0.05, across all 373 genotypes as well as independently within the indica, aus, temperate japonica, and tropical japonica subpopulations '''(Figure 1)'''. The Efficient Mixed-Model Association (EMMA) model was used in each analysis (both within and across subpopulations) to correct for confounding effects due to subpopulation structure and relatedness between individuals. As the subpopulation structure was highly correlated with Al tolerance, it was observed that analyzing all samples (373) together with the EMMA model resulted in an overcorrection (causing type 2 error) and a corresponding reduction in SNP significance. To address this problem, a PCA approach was also employed when analyzing all (373) samples together. However, the PCA approach resulted in a slight under-correction for population structure, demonstrating that results from each GWA method has limitations when used across all germplasm in this highly structured diversity panel.
+
* Using principle component analysis (PCA) 14 to summarize global genetic variation in the diversity panel, we observed clear, deep subpopulation structure in this collection of germplasm( Figure. 1 ). The top four principal components (PCs) explained almost half of the genetic variation ( Fig. 1b ). Th e fi ve subpopulations indica , aus , temperate japonica , tropical japonica and aromatic formed clear clusters based on the top four PCs, and were well diff erentiated from each other, with pairwise Fst (F-statistic) values ranging from 0.23 – 0.53. Th is is in agreement with previous fi ndings where global germplasm collections have been used in combination with much smaller numbers of SNP or simple sequence repeat(SSR) genotypes 8,15 – 17 . Because the array was designed to assay variation in all O. sativa groups, most SNPs are shared or polymorphic across subpopulations.
* We chose to further investigate the variation in and around the Nrat1 gene on chromosome 2 because multiple independent lines of evidence supported the existence of a gene(s) in this region responsible for a significant portion of the variation for Al tolerance in rice. Evidence included a strong GWA peak in the aus subpopulation, a previously reported QTL, and the localization of the Nrat1 Al transporter gene. Using the 44 K SNP data, LD in this region was calculated to be ,150 kb in the aus subpopulation and 11 distinct haplotypes were observed in the entire diversity panel across a 139 kb region around the Nrat1 gene(1.536 Mb–1.675 Mb on chr. 2) (Figure 2). Haplotype 1 (Hap.1), which was unique to the aus subpopulation, was found in 8 Al sensitive aus accessions and one Al sensitive aus/indica admixed line. These 9 genotypes were among the least Al tolerant (7 th percentile, mean RRG=0.16) of the 373 accessions screened. Haplotype 1 explained 40% of the phenotypic variation for Al tolerance within the aus subpopulation. In addition, four aus accessions that were highly or moderately Al tolerant were found to contain a tropical japonica introgression across this region (described in the section on Introgression analysis below).
+
[[File:IC4R002-GWAS-2011-21915109-1.PNG|1000px|thumb|center|'''Figure 1''' ''Principal component analysis was used to provide a statistical summary of the genetic data, and the top four principle components are illustrated in the bottom panels.'']]
[[File:IC4R001-GWAS-2011-21829395-2.PNG|700px|thumb|right|'''Figure 2.''' '' Haplotype analysis of the Nrat1 gene region.'']]
+
* The phenotypes we examined in our GWAS can be classifi ed broadly into six categories: plant morphology-related traits; yield-related traits; seed and grain morphology-related traits; stress-related phenotypes; cooking, eating and nutritional-quality-related traits; and plant development, represented by fl owering time, which we measured in three geographic locations that diff ered in day-length and ambient temperature. Canonical correlation analysis demonstrated that phenotypes within a category are oft en correlated, ranging from a low of − 0.41 between brown rice seed width and brown rice seed length, to a high of 0.9 between hulled and dehulled seed morphology.
 +
* The results of our genome-wide association scans are summarized in Supplementary Figures S3 – S36 where we show SNP-trait associations discovered in the diversity panel as a whole, as well as in each subpopulation individually. As can be seen in the quantile–quantile plots ( Figure. 3 ), the distribution of observed − log10 P -values from the na ï ve analysis (no population structure adjustment) departed quite far from the expected distribution under a model of no association (that is, the P -values should lie on the diagonal line), with signifi cant infl ation of nominal P -values leading to a high level of false positive signals. Use of a modifi ed mixed model strategy 22 – 24 allowed us to consider diff erent levels of population structure and relatedness in our diversity panel. Th is eff ectively eliminated the excess of low P -values for most traits, but it also likely eliminated true positives. Th is is a common problem seen in other systems as well; for example, geographic coordinates correlate closely with fl owering time in plants 24 . For this reason, we believe a combination of na ï ve and population structure-adjusted hits, coupled with subpopulation-specifi c analyses in rice, is the most thoughtful way to identify potential variants for follow up.
 +
[[File:IC4R002-GWAS-2011-21915109-2.PNG|500px|thumb|right|'''Figure 2''' ''Quantile – Quantile plots for both na ï ve and mixed model for plant height in all samples. '']]
  
 
== Labs working on this Project ==
 
== Labs working on this Project ==
* Department of Plant Breeding and Genetics, Cornell University, Ithaca, New York, United States of America
+
* Department of Biological Statistics and Computational Biology, Cornell University, Ithaca, New York 14850, USA
* Department of Biological Statistics and Computational Biology, Cornell University, Ithaca, New York, United States of America
+
* Department of Genetics, Stanford University, Stanford, California 94305, USA
* Robert W. Holley Center for Agriculture and Health, Agricultural Research Service, US Department of Agriculture, Cornell University, Ithaca, New York, United States of America
+
* Department of Plant Breeding and Genetics, Cornell University, Ithaca, New York 14850, USA
 +
* USDA ARS,Dale Bumpers National Rice Research Center, Stuttgart, Arkansas 72160, USA
 +
* Rice Research and Extension Center, University of Arkansas, Stuttgart, Arkansas 72160, USA
 +
* Institute of Biological and Environmental Sciences, University of Aberdeen, Aberdeen AB24 3UU, UK
 +
* Department of Soil Science,Bangladesh Agricultural University, Mymensingh 2202, Bangladesh. Correspondence and requests for materials should be addressed to S.R.M.
  
 
==Corresponding Author==
 
==Corresponding Author==
''' Susan R. McCouch'''(srm4@cornell.edu)
+
* '''Susan R. McCouch'''(srm4@cornell.edu)
 +
* '''Carlos D. Bustamante'''(cdbustam@stanford.edu)

Latest revision as of 10:16, 21 June 2016

Project Title

Genome-wide association mapping reveals a rich genetic architecture of complex traits in Oryza sativa

The Background of This Project

  • Asian rice, Oryza sativa is a cultivated, inbreeding species that feeds over half of the world's population. Understanding the genetic basis of diverse physiological, developmental, and morphological traits provides the basis for improving yield, quality and sustainability of rice. Here we show the results of a genome-wide association study based on genotyping 44,100 SNP variants across 413 diverse accessions of O. sativa collected from 82 countries that were systematically phenotyped for 34 traits. Using cross-population-based mapping strategies, we identifi ed dozens of common variants infl uencing numerous complex traits. Signifi cant heterogeneity was observed in the genetic architecture associated with subpopulation structure and response to environment. This work establishes an open-source translational research platform for genome-wide association studies in rice that directly links molecular variation in genes and metabolic pathways with the germplasm resources needed to accelerate varietal development and crop improvement.

Plant Culture & Treatment

  • The Rice Diversity Panel consists of 413 Asian rice ( O. sativa ) cultivars, including many landraces, which originated from 82 countries, representing all the major rice-growing regions of the world. The panel contains 87 indica , 57 aus , 96 temperate japonica , 97 tropical japonica , 14 groupV/aromatic, and 62 highly admixed accessions. All accessions were purifi ed for two generations(single seed descent) before DNA extraction. In all, 20 of these 413 accessions were purifi ed as part of the Oryza SNP project. Six cultivars (Azucena, Moroberekan, Nipponbare, Dom-Sofi d, IR64, M-202) were purifi ed separately, once by Ali et al. and once as part of the Oryza SNP panel. Further information for each accession(accession name, accession number, country of origin and subpopulation ancestry based on PCA) is given in Supplementary Data 1 .

Research Findings

  • A rice diversity panel consisting of 413 inbred accessions of O. sativa collected from 82 countries was genotyped using an Aff ymetrix single nucleotide polymorphism (SNP) array containing 44,100 SNPs (hereaft er referred to as the 44 K chip). With a genome size of ~ 380 Mb (ref. 13), this custom-designed genotyping chip provides high quality data (less than 4.5 % missing data), with ~ 1 SNP per 10 kb across the 12 chromosomes of rice. The diversity panel was evaluated for 34 traits related to plant morphology, grain quality, plant development and agronomic performance using fi eld-grown plants with replications within and between years.

  • Using principle component analysis (PCA) 14 to summarize global genetic variation in the diversity panel, we observed clear, deep subpopulation structure in this collection of germplasm( Figure. 1 ). The top four principal components (PCs) explained almost half of the genetic variation ( Fig. 1b ). Th e fi ve subpopulations indica , aus , temperate japonica , tropical japonica and aromatic formed clear clusters based on the top four PCs, and were well diff erentiated from each other, with pairwise Fst (F-statistic) values ranging from 0.23 – 0.53. Th is is in agreement with previous fi ndings where global germplasm collections have been used in combination with much smaller numbers of SNP or simple sequence repeat(SSR) genotypes 8,15 – 17 . Because the array was designed to assay variation in all O. sativa groups, most SNPs are shared or polymorphic across subpopulations.
Figure 1 Principal component analysis was used to provide a statistical summary of the genetic data, and the top four principle components are illustrated in the bottom panels.
  • The phenotypes we examined in our GWAS can be classifi ed broadly into six categories: plant morphology-related traits; yield-related traits; seed and grain morphology-related traits; stress-related phenotypes; cooking, eating and nutritional-quality-related traits; and plant development, represented by fl owering time, which we measured in three geographic locations that diff ered in day-length and ambient temperature. Canonical correlation analysis demonstrated that phenotypes within a category are oft en correlated, ranging from a low of − 0.41 between brown rice seed width and brown rice seed length, to a high of 0.9 between hulled and dehulled seed morphology.
  • The results of our genome-wide association scans are summarized in Supplementary Figures S3 – S36 where we show SNP-trait associations discovered in the diversity panel as a whole, as well as in each subpopulation individually. As can be seen in the quantile–quantile plots ( Figure. 3 ), the distribution of observed − log10 P -values from the na ï ve analysis (no population structure adjustment) departed quite far from the expected distribution under a model of no association (that is, the P -values should lie on the diagonal line), with signifi cant infl ation of nominal P -values leading to a high level of false positive signals. Use of a modifi ed mixed model strategy 22 – 24 allowed us to consider diff erent levels of population structure and relatedness in our diversity panel. Th is eff ectively eliminated the excess of low P -values for most traits, but it also likely eliminated true positives. Th is is a common problem seen in other systems as well; for example, geographic coordinates correlate closely with fl owering time in plants 24 . For this reason, we believe a combination of na ï ve and population structure-adjusted hits, coupled with subpopulation-specifi c analyses in rice, is the most thoughtful way to identify potential variants for follow up.
Figure 2 Quantile – Quantile plots for both na ï ve and mixed model for plant height in all samples.

Labs working on this Project

  • Department of Biological Statistics and Computational Biology, Cornell University, Ithaca, New York 14850, USA
  • Department of Genetics, Stanford University, Stanford, California 94305, USA
  • Department of Plant Breeding and Genetics, Cornell University, Ithaca, New York 14850, USA
  • USDA ARS,Dale Bumpers National Rice Research Center, Stuttgart, Arkansas 72160, USA
  • Rice Research and Extension Center, University of Arkansas, Stuttgart, Arkansas 72160, USA
  • Institute of Biological and Environmental Sciences, University of Aberdeen, Aberdeen AB24 3UU, UK
  • Department of Soil Science,Bangladesh Agricultural University, Mymensingh 2202, Bangladesh. Correspondence and requests for materials should be addressed to S.R.M.

Corresponding Author

  • Susan R. McCouch(srm4@cornell.edu)
  • Carlos D. Bustamante(cdbustam@stanford.edu)