Scaled RNA sequencing counts accompaining the publication Singh and Gollapalli 2023
Authors/Creators
- 1. Walter and Eliza Hall Institute of Medical Research, Melbourne Australia
- 2. Metabolic Research Laboratories, National Institute of Immunology; New Delhi, India
- 3. Systems Biology of Aging Laboratory, Columbia University; New York, USA
- 4. Metabolic Research Laboratories, National Institute of Immunology; New Delhi, India; Systems Biology of Aging Laboratory, Columbia University; New York, USA
Description
RNA-Seq analysis of WT and taurine-deficient osteoblasts.
RNA was extracted from the mouse primary calvarial osteoblasts prepared from day 3–5 neonates of WT or Slc6a6-/- genotypes using the RNeasy extraction kit (Qiagen). RNA sequencing libraries were prepared with Illumina-compatible NEBNext UltraTM II Directional RNA library kit at Genotypic Technology, Inc. The raw data were trimmed for adapter sequences and low-quality bases (<q30) using Cutadapt with default parameters [DOI:10.14806/ej.17.1.200] and checked for quality using FastQC [Andrews, S. (2010). FastQC: A Quality Control Tool for High Throughput Sequence Data]. Mouse raw reads were aligned to the mouse reference genome (mm9) using Hisat2(24) with default parameters. HTSeq [G Putri, S Anders, PT Pyl, JE Pimanda, F Zanini Analyzing high-throughput sequencing data with HTSeq 2.0 arXiv:2112.00939 (2021)] was used to estimate gene abundance. Downstream analyses were done using the analysis framework tidybulk (25, 26). The filterByExpr functionality of edgeR (27) was used to identify the abundant gene-transcripts included in downstream analyses using default parameters and the knock-out phenotype as the factor of interest. The algorithm trimmed mean of M valued (TMM) values identified sample-wise scaling to compensate for differences in sequencing depth across samples (28). For aiding data exploration, we reduced the data dimensionality using principal component analysis (PCA). Differential expression analyses were performed using the robust likelihood ratio implementation of edgeR, testing for differences in gene transcript abundance greater than 1 log-fold (29, 30). The ppcseq method was used to investigate the presence of outlier observation among the significant results to avoid biases in the statistics (31). The ggplot2 software produced most of the data visualizations, whereas tidyHeatmap was used for heatmap visualization (32-34). We performed gene-set enrichment analyses using the MSigDB C2 experimentally derived gene set, the clusterProfiler algorithm, and the visualization tool enrichplot (cran.r-project.org/) (https://rdrr.io/cran/msigdbr/) (35). An aging signature was composed of a union of gene sets from MSigDB and a set of genes derived from aging datasets identified in this study through gene set enrichment analysis and published literature (36). The p-values of gene set analyses were corrected for multiple testing using the Benjamini Hochberg correction (37).
Files
counts_scaled.csv
Files
(4.4 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:8f0a63506d23af8813fc968a34c5f522
|
4.4 MB | Preview Download |