概要: Single cell RNA sequencing (scRNA-seq) can be used to characterize variation in gene expression levels at high resolution. However; the sources of experimental noise in scRNA-seq are not yet well understood. We investigated the technical variation associated with sample processing using the single cell Fluidigm C1 platform. To do so; we processed three C1 replicates from three human induced pluripotent stem cell (iPSC) lines. We added unique molecular identifiers (UMIs) to all samples; to account for amplification bias. We found that the major source of variation in the gene expression data was driven by genotype; but we also observed substantial variation between the technical replicates. We observed that the conversion of reads to molecules using the UMIs was impacted by both biological and technical variation; indicating that UMI counts are not an unbiased estimator of gene expression levels. Based on our results; we suggest a framework for effective scRNA-seq studies.
项目整体设计: We combined the 96 single cell samples from each C1 chip into their own master mix and sequenced across three lanes of a HiSeq 2500 (3 individuals x 3 replicates x 96 wells x 3 lanes = 2592 files). We prepared two separate library preparations for each bulk sample; combined them all into one master mix; and sequenced across four lanes (3 individuals x 3 replicates x 2 library preparations x 4 lanes = 72 files).
Undifferentiated feeder-free iPSCs generated from Yoruba LCLs were grown in E8 medium (Life Tech) (G. Chen et al. 2011) on Matrigel-coated tissue culture plates with daily media feeding at 37 °C with 5% (vol/col) CO2. For standard maintenance, cells were split every 3-4 days using cell release solution (0.5 mM EDTA and NaCl in PBS) at the confluence of roughly 80%. For the single cell suspension, iPSCs were individualized by Accutase Cell Detachment Solution (BD) for 5-7 minutes at 37 °C and washed twice with E8 media immediately before each experiment. Cell viability and cell counts were then measured by the Automated Cell Counter (Bio-Rad) to generate resuspension densities of 2.5 X 105 cells/mL in E8 medium for C1 cell capture.; Undifferentiated feeder-free iPSCs generated from Yoruba LCLs were grown in E8 medium (Life Tech) (G. Chen et al. 2011) on Matrigel-coated tissue culture plates with daily media feeding at 37 °C with 5% (vol/col) CO2. For standard maintenance, cell were split every 3-4 days using cell release solution (0.5 mM EDTA and NaCl in PBS) at the confluence of roughly 80%. For the single cell suspension, iPSCs were individualized by Accutase Cell Detachment Solution (BD) for 5-7 minutes at 37 °C and washed twice with E8 media immediately before each experiment. Cell viability and cell counts were then measured by the Automated Cell Counter (Bio-Rad) to generate resuspension densities of 2.5 X 105 cell/mL in E8 medium for C1 cell capture.
处理方案:
-
提取方案:
Single cell loading and capture was performed following the Fluidigm manual; A bulk sample, a 40 ul aliquot of ~10,000 cell, was collected in parallel with each C1 chip using the same reaction mixes following the C1 protocol of ""Tube Controls with Purified RNA
建库方案:
For sequencing library preparation, fragmentation and isolation of 5^ fragments were performed according to the UMI protocol (Islam et al. 2014). Instead of using commercial available Tn5 transposase, Tn5 protein stock was freshly purified in house using the IMPACT system (pTXB1, NEB) following the protocol previously described (Picelli et al. 2014). The activity of Tn5 was tested and shown to be comparable with the EZ-Tn5-Transposase (Epicentre). Importantly, all the libraries in this study were generated using the same batch of Tn5 protein purification. For each of the bulk samples, two libraries were generated using two different indices in order to get sufficient material.
测序信息
分子类型:
poly(A)+ RNA
库的片段类型:
SINGLE
库的链类型:
Reverse
测序平台:
ILLUMINA
测序仪型号:
Illumina HiSeq 2500
链特异性:
Specific
样本
基本信息:
样本描述:
生物条件:
实验变量:
方案:
测序信息:
质量评估:
数据来源
GEN样本编号
GEN数据集编号
系列编号
项目编号
样本编号
样本名称
生物样本编号
样本访问号
实验访问号
释放时间
提交时间
最后更新时间
物种
种族
族裔
年龄
年龄单位
性别
来源名称
组织
细胞类型
细胞亚型
细胞系
疾病
疾病状态
发育阶段
突变/变异
表型
Condition Detail
生长方案
处理方案
提取方案
建库方案
分子类型
库的片段类型
链特异性
库的链类型
加标(Spike-In)
测序方法
测序平台
测序仪型号
细胞数
测序片段数
碱基数
平均测序片段长度_1
平均测序片段长度_2
唯一比对率
多重比对率
覆盖度
文章
Batch effects and the effective design of single-cell gene expression studies.