PMC7557106	quality_control,normalization,integration,variable_genes,dimensionality_reduction,clustering,visualization	The resultant files were imported into R version 3.6.3 and processed using Seurat version 3.1.5 for quality control, normalisation, integration, variable gene selection, dimensionality reduction, clustering and visualisation.
PMC7557106	clustering	Subsequent graph‐based clustering of cells according to their gene expression profile was achieved using the FindNeighbors and FindClusters functions of Seurat with resolution set to 0.5.
PMC7057921	alignment	Gene fusion prediction For each cell, the clean reads were mapped to the human genome reference sequences (hg19 version) using the STAR aligner (v2.4.1) (12), and fusion gene detection was performed using STAR-Fusion (https : //github.com / STAR-Fusion / STAR-Fusion; v1.3.2), which compared with other methods, was sufficient for fusion RNA prediction (13).
PMC7184133	variable_genes	First, we followed the “ Seurat ” pipeline [ 36 ] to identify highly variable genes and demonstrated that the eight cell types were as shown in the previous report [ 35 ] (Fig.
PMC7184133		The Seurat R package [ 36 ] was used to perform the single-cell RNA-seq dataset analysis.
PMC6468969	quality_control	In total, 24 120 genes in 21 750 cells passed the Seurat quality control filtering (see the Experimental Section) and were used for downstream analysis (Table S1, Supporting Information).
PMC6468969	trajectory,clustering	b) Single‐cell pseudotime lineage trajectory obtained by semisupervised clustering of subtype‐specific gene panel using Monocle.
PMC6468969		c) Representative images of tumor cells in microvasculature and the matched Monocle plots for top three and bottom three colocalization coefficient samples (Day 4–7).
PMC6468969	marker_genes	We then reconstructed the pseudotime lineage relationships with three subtype markers (555 genes) using semi‐supervised analysis in Monocle package.64, 65 The resulting plot had three major branches and three major nodes connecting them with two minor branches (Figure 5b).
PMC6468969	alignment,quantification	Raw reads were preprocessed for cell barcodes and UMIs, and then aligned to the human genome (hg19) using STAR v2.5.2b as described in Dropseq method.84 Digital expression matrix was generated for the cells with over 10 000 reads per cell.
PMC6468969	differential_expression	The Seurat package (V2.3.0) in R (V3.4.1) was applied to identify differentially expressed genes among 26 027 single cells from nine different GBM patients and one GBM cell reference (GS5) .85 Cells were considered in the analysis only if they met the following quality control criteria : 1) expression of more than 1000 genes and fewer than 5000 genes; 2) low expression of mitochondrial genes (< 10 % of total counts in a cell).
PMC6468969	differential_expression,marker_genes	Genes that were differentially expressed in each cluster were identified using the Seurat function FindMarkers, which returned the gene names, average log fold‐change, and adjusted p‐value for genes enriched in each cluster.
PMC6468969	trajectory,clustering	The Monocle package (V2.6.4) was used to plot single cell pseudotime trajectories to discover the behavioral similarity and transitions.64, 65 We use the proneural, mesenchymal, and classical subtype genes identified before to perform the semi‐supervised analysis.58 Monocle looked for variable genes and augmented the markers when construct the clustering and ordering of the cells.
PMC7610308	alignment	Data were analyzed and mapped to the mouse genome (mm10) using CellRanger software (10× Genomics).
PMC8181206	trajectory	By ordering monocytes and macrophages to reconstruct pseudo‐time trajectories using Monocle 2 algorithm, we observed a clear directional flow that monocytes bifurcated to two branches of SPP1 TAMs and C1QC Mφ (Figure 5D), suggesting distinct cellular differentiation paths of these two subpopulations.
PMC8181206		The generated outputs were processed using R package Seurat (version 3.1).
PMC8181206	normalization	The filtered gene expression matrix was normalized using Seurat 's NormalizeData function so that the expression of each gene was multiplied 10,000 and log transformed.
PMC8181206	dimensionality_reduction	Afterwards we ran PCA analysis using top 2000 genes identified by FindVariableFeatures function of Seurat.
PMC8181206	clustering,marker_genes,classification	Identification of cellular clusters and differentially expressed genes We used Seurat 's FindAllMarkers function for identification of major canonical cellular clusters, by which we identified marker genes and designated cell cluster labels accordingly.
PMC8181206	trajectory	Construction of single‐cell trajectories To construct cell development trajectories, we used the Monocle 2 package to align the cells in pseudo‐time order.
PMC8181206	dimensionality_reduction	The DDRTree approach implemented in the reduceDimension function of Monocle 2 was used to map cells.
PMC8558147	alignment	Index of the reference genome was built using STAR (v2.5.1b) and paired-end clean reads were aligned to the reference genome.
PMC8558147	quality_control,dimensionality_reduction,clustering	Seurat v.3 was used for quality control, dimensionality reduction, and cell clustering.
PMC8558147	normalization	The filtered gene-barcode matrix was first normalized using ‘ LogNormalize ’ methods in Seurat v.3 with default parameters.
PMC8558147	variable_genes	The top 2000 variable genes were then identified using the ‘ vst ’ method in Seurat FindVariableFeatures function.
PMC8558147	clustering	Meanwhile, graph-based clustering was performed on the PCA-reduced data for clustering analysis with Seurat v.3.
PMC8558147	differential_expression	2015) in Seurat v.3 (FindAllMarkers function) was used to perform differential gene expression analysis.
PMC8558147	visualization	Violin plots were performed by the function VlnPlot in Seurat with the default parameters.
PMC8558147	marker_genes	Expression heatmap for marker genes was performed by the function DoHeatmap in Seurat with the default parameters.
PMC8558147	classification	Based on this scoring system, each cell was classified in either G2 / M, S, or G1 phase using the CellCycleScoring function in Seurat.
PMC8558147		Cell type specific expression of disease-associated genes To calculate the average expression level for each cluster in scRNA-seq data, the function AverageExpression in Seurat was used with the default parameters.
PMC8558147	dimensionality_reduction,trajectory	UMAP embeddings and cell clusters generated from Seurat were used as input, and trajectory graph learning and pseudo-time measurement through reversed graph embedding was performed with Monocle3.
PMC8187228	integration,clustering	The resultant datasets were analyzed, both singly and merged, using the Seurat R package v2.4.3 as per the clustering workflow.
PMC8187228	variable_genes	Highly variable genes were identified using Seurat 's FindVariableGenes function with default parameters.
PMC8187228	visualization	Heat maps, t-distributed stochastic neighbur embedding (t-SNE) visualizations, and violin plots were produced using Seurat functions.
PMC8187228	trajectory	Pseudotime between clusters was assessed using the Monocle workflow in R. Canonical correlation analysis (CCA) was performed on the datasets to look for response to Adam10 knockout.
PMC7671315	clustering	Phenograph and Seurat are based on the shared nearest neighbor graph and they use the Louvain algorithm to detect cell community (21,22).
PMC7671315	clustering	SAFE is another consensus clustering method that integrates clustering results from t-SNE, CIDR, Seurat and SC3 (25).
PMC7671315		Other methods, such as Seurat (22), SC3 (24) and so on, are not in consideration of comparison.
PMC7671315		Seurat is a useful and efficient package for biologists to analyze single-cell data.
PMC7671315	clustering	However, due to a lot of parameters, especially ‘ resolution ’ parameter, being elaborately selected by users, Seurat usually tends to overestimate the cluster number that can not perform well under default parameters.
PMC7671315	clustering	For the sake of fairness, we do not show the clustering performance of Seurat here.
PMC7671315		We will show the additional comparison with Seurat and SC3 in the ‘ Discussion and Conclusion ’ section.
PMC7671315		Additional comparison with Seurat and SC3 in real datasets As we mentioned in the ‘ Results ’ section, we did not compare with Seurat and SC3 in whole simulation and real datasets for fairness.
PMC7671315		We indeed compare our method with Seurat and SC3 on 10 real datasets.
PMC7671315	clustering	Since clustering performance of Seurat is highly dependent on model parameters, especially ‘ resolution ’, here we evaluate its performance under default parameters of pre-process and PCA selection procedure, and then take the resolution parameter in the clustering step to range from 0.5 to 1.5 by 0.1, which is recommended by the author of Seurat.
PMC7671315	clustering	Figure 10A shows the ARI values of Seurat, SC3 and scziDesk in five small datasets and five large datasets.
PMC7671315	clustering	Despite selecting the optimal resolution parameter from the candidate range, Seurat still achieves unsatisfactory performance in ARI values, which is mostly due to its linear dimension reduction and overestimation of the cluster numbers (shown in Figure 10B).
PMC7671315	clustering	(A) Comparison of ARI values between Seurat, SC3 and scziDesk in 10 real datasets.
PMC7671315	clustering	(B) Comparison of optimal cluster number estimated by Seurat with the true cluster number in 10 real datasets.
PMC7671315	differential_expression	In addition, we applied function ‘ FindAllMarkers ’ with default parameters in Seurat (22) package to find out the DEGs of each cluster.
PMC7671315	differential_expression	Moreover, we used the completely same way to obtain the DEGs under the golden standard clusters with cell type annotation, and also under the Seurat standard analysis pipeline with resolution parameter equal to 0.3 to get 14 clusters.
PMC7671315	clustering	However, for Seurat clustering in Figure 11B, although most clusters can match to unique cell type, there is no cluster match with the 14th cell type, Fraction A pre-pro B cell.
PMC7671315	differential_expression	(B) Comparison of similarity of DEGs in 14 Seurat clusters with golden standard 14 cell types.
PMC5428745	trajectory	Monocle equivalently allows re-fitting of the pseudotimes with the constraint that one of the inferred “ states ” is the initial or root state.
PMC5428745	trajectory	We compared the Pearson correlation of the estimated pseudotimes to the true pseudotimes (Figure 2C) for both MFA, PC1 (the first principal component of the data), Monocle and Diffusion Pseudotime, giving values of 0.98, 0.98, 0.98 and 0.99 (to 2 s.f.
PMC5428745	trajectory	Specifically, we generated synthetic datasets with 0 %, 20 %, …, 80 % of genes exhibiting transient expression, and inferred the pseudotimes using DPT, MFA and Monocle 2.
PMC5428745		The performance of MFA remains competitive up to around 40 % of genes exhibiting transient expression, after which DPT and Monocle 2 perform significantly better.
PMC5428745	trajectory	However, MFA is highly consistent with DPT and Monocle 2 on the two real datasets examined (Figure 4 and Figure 5), implying the occurrence of transient expression is limited enough in practice for the linearity assumption to be feasible.
PMC5428745	clustering	We next sought to compare the performance of MFA to existing bifurcation inference algorithms, in particular Wishbone, DPT and Monocle (v2), along with the second principal component of the data (PC2), which we noted from exploratory analyses was highly correlated with the existing Wishbone values.
PMC5428745		We sub-sampled down to 1,000 cells for Monocle comparisons for computational convenience and used the previously published results for Wishbone (from 3).
PMC5428745	trajectory	The root cell for DPT was selected as the cell with the minimum value for the second principal component and similarly the root state for Monocle was chosen such that it contained that cell.
PMC5428745		However, there is virtually no correlation with Monocle (ρ = 0.01), though as this low correlation only occurs with Monocle we assume it is not an issue with MFA.
PMC5428745	visualization	Figure 4F shows a tSNE representation of the cells colored by branch allocation for each of Wishbone, Monocle and DPT.
PMC5428745	trajectory	We see that MFA is largely consistent with Wishbone and DPT, detecting a bifurcation at the “ pinch ” in the tSNE plot, but as with the pseudotimes there is barely any correspondence in branch allocations with Monocle (which, as of version 2, does not allow pre-specification of the number of branches to model).
PMC5428745		For computational convenience with all algorithms, we sub-sampled the data down to 2,000 randomly chosen cells, with the exception of Monocle, which we subsequently sub-sampled further down to 1,000 cells.
PMC5428745		We found good correspondence to all other methods (Figure 5C), with Pearson correlations of 0.84, 0.86, 0.80 and 0.69 for PC2, Wishbone, Monocle, and DPT, respectively.
PMC5428745	trajectory	As of version 2, Monocle does not allow for the number of branches to be selected a priori and typically returns a large number.
PMC5428745	trajectory	We find good agreement between MFA and Monocle and DPT, and similarities with the Monocle assignments (MFA branch 2 loosely corresponds to Monocle branch 17).
PMC5428745	trajectory	MFA is compared with several popular existing algorithms that are not generative probabilistic models : Wishbone, Diffusion Pseudotime, and Monocle 2.
PMC5428745		* Running Monocle 2 on a smaller set of sub-sampled cells than the other methods could put it at a disadvantage.
PMC8187165	clustering,classification	Methods comparison SC3 (version 1.14.0), t-SNE (package Rtsne version 0.15), Seurat (version 2.3.4) and SparseDC (version 0.1.17) were also applied to the small data sets, where all four are to benchmark cell clustering, and three for classification except t-SNE.
PMC8187165	clustering,dimensionality_reduction	For Seurat, we chose the improved PCA-based clustering pipeline as representative method from several packaged dimensionality reduction models.
PMC8187165	clusertering	The number of cell clusters is inferred by the default functions for SC3 and Seurat, or directly set as original group number for SparseDC and t-SNE.
PMC8187165	marker_genes,classification	The whole sets of the identified marker genes are used when built corresponding KNN classifiers for SC3, Seurat and SparseDC.
PMC8187165		Here, we considered four established methods : SC3, t-SNE, Seurat [ 29 ] and SparseDC for fair comparisons.
PMC6945013	alignment	Trimmed reads were mapped to the GRCh38 assembly of the human genome using STAR [ – outSAMmapqUnique 60 – outSAMtype BAM SortedByCoordinate ] (Dobin).
PMC6945013	alignment	Reads were mapped to the ERCC92.fa sequence file downloaded from Thermofisher.com using STAR (Dobin).
PMC7644053	integration,clustering	For scRNAseq, we used Seurat to integrate data and implement shared nearest neighbors (SNN) clustering (see STAR Methods).
PMC7644053	alignment	scRNAseq Processing and Analysis scRNAseq data was aligned to the hg19 genome and processed with CellRanger version 2.0.2.
PMC7644053	integration,clustering,visualization	Dataset integration, SNN clustering, and UMAP visualization of scRNAseq data were performed with Seurat version 3.0.0.9000.
PMC7644053	clustering	Clusters were identified by SNN clustering with Seurat FindNeighbors and FindClusters functions using 50 dimensions and a resolution parameter of 2.5.
PMC7644053	marker_genes,differential_expression	Cell cluster marker genes (padj < 0.05) and smoking DEGs (padj < 0.05) were identified using Seurat version 2.3.4 implementing the MAST algorithm with UMI included as a latent variable.
PMC7644053	alignment	Bulk RNAseq Analysis For antibody-purified CD8 T cell fraction bulk RNAseq, reads were aligned to the hg19 genome with STAR.
PMC6687332	alignment,quality_control,quantification	Briefly, using cellranger mkfastq and cellranger count, FASTQ files were generated and aligned to the mm10 genome, sequencing reads were filtered by base-calling quality scores, and then cell barcodes and UMIs were assigned to each read in the FASTQ files.
PMC6687332	variable_genes	Briefly, we identified the top 1000 variable genes in each dataset, used the intersection of the variable genes to perform canonical correlation analysis (CCA), and then aligned the canonical correlation vectors (CCs) using the R package Seurat [ 27 ].
PMC6687332	dimensionality_reduction	In our analysis, we chose to align the first 25 CCs after examining the shared correlation strength as a function of the number of CCs for both sample (the Seurat function MetageneBicorPlot).
PMC6687332	dimensionality_reduction,marker_genes,clustering	t-Stochastic Neighbor Embedding (t-SNE), Clustering Analysis, and Definition of Marker Genes We used the first 25 aligned CCs to run t-SNE [ 28 ] dimensionality reduction and find Louvain clusters, using the default resolution of 0.6, using the R package Seurat [ 27 ].
PMC6687332	quality_control	Data Filtering Cells from all 3 conditions were merged using the RunMultiCCA function of Seurat [ 27 ], as described in Butler et al.
PMC6687332	differential_expression	Differential Expression Analysis by Sample We used Seurat [ 27 ] to identify differentially expressed genes by sample for each cluster, using the Wilcoxon test to generate p-values.
PMC6687332	trajectory	Pseudotime Analysis and Branched Gene Expression Analysis We used the R package Monocle [ 30 ] to reconstruct the divergence of cell lineages / trajectories in the cells identified as macrophages in our analysis.
PMC6687332	differential_expression	Briefly, we first used Monocle to estimate size factors, dispersion, and differential gene expression of the subset of macrophage cells, and then used the top 1000 most differentially expressed genes between the macrophage clusters to order cells in pseudotime.
PMC6687332	trajectory	We used the BEAM feature of Monocle to define genes that show significant divergent expression across each branch point in the pseudotime analysis, using default parameters.
PMC7202009	alignment	Evaluation of STAR and Kallisto on Single Cell RNA-Seq Data Alignment Alignment of scRNA-Seq data are the first and one of the most critical steps of the scRNA-Seq analysis workflow, and thus the choice of proper aligners is of paramount importance.
PMC7202009	alignment	Recently, STAR an alignment method and Kallisto a pseudoalignment method have both gained a vast amount of popularity in the single cell sequencing field.
PMC7202009	alignment	We observe that STAR globally produces more genes and higher gene-expression values, compared to Kallisto, as well as Bowtie2, another popular alignment method for bulk RNA-Seq.
PMC7202009		STAR also yields higher correlations of the Gini index for the genes with RNA-FISH validation results.
PMC7202009	gene_markers	Using 10x genomics PBMC 3 K scRNA-Seq and mouse cortex single nuclei RNA-Seq data, STAR shows similar or better cell-type annotation results, by detecting a larger subset of known gene markers.
PMC7202009	alignment	However, the gain of accuracy and gene abundance of STAR alignment comes with the price of significantly slower computation time (4 folds) and more memory (7.7 folds), compared to Kallisto.
PMC7202009	alignment	Recently, two methods have gained popularity in the single cell field : STAR (Dobin) and Kallisto (Bray).
PMC7202009	alternative_splicing	STAR detects the splice junctions and aligns the sequence to the reference genome non-contiguously.
PMC7202009	alignment	Previous benchmarking studies showed STAR was one of the most reliable reference genome based aligners in RNA-seq analysis (Baruzzo); (Engström); (Yang).
PMC7202009	alignment	Both STAR and Kallisto can quantify expression, but the direct comparison between these two specific methods was lacking, especially in the scRNA-seq field.
PMC7202009	alignment	STAR alignment : STAR alignment was performed for Drop-seq data using STAR version 2.5.2a provided by University of Michigan HPC.
PMC7202009	alignment	STAR alignment : STAR alignment was performed for trimmed data using STAR version 2.5.2a, and aligned to reference genome GRCh38.
PMC7202009	quality_control	Data processing and analysis on 10x Genomics PBMC 3 K data Cell Ranger filtered genes and cell matrix (available on 10x website) were processed using Seurat version 3.2.0 (Butler; Stuart), following the configuration from https : //satijalab.org / seurat / v3.0 / pbmc3k_tutorial.html.
PMC7202009	marker_genes	Clusters were annotated using provided labels by the Seurat group based on the pre-defined gene markers.
PMC7202009	alignment	Fastq files were preprocessed to fit the input scheme for both STAR and Kallisto.
PMC7202009		STAR version 2.7.1a with – solo command was used.
PMC7202009	alignment	The STAR index was built with a read length of 98.
PMC7202009	clustering	Downstream clustering analysis : Seurat version 3.2.0 was used for downstream analysis.
PMC7202009	normalization	Data were then log normalized with a scale factor of 10000 in Seurat.
PMC7202009		STARsolo : STAR version 2.7.3a with – solo command was used.
PMC7202009	alignment	The STAR index was built with a read length of 50.
PMC7202009	clustering	Downstream clustering analysis : Seurat version 3.2.0 was used for downstream analysis.
PMC7202009	normalization	Data were then log normalized with a scale factor of 10000 in Seurat.
PMC7202009	alignment	Comparisons of STAR vs. Kallisto alignment results on Drop-Seq and Fluidigm data STAR and Kallisto are based on different concepts.
PMC7202009	alignment	STAR is a conventional aligner that aligns to the reference genome, whereas Kallisto uses transcriptome quantification for pseudoalignment.
PMC7202009	alignment	We used GRCh38 as the reference genome for STAR and GRCh38 as the reference transcriptome for Kallisto, per recommendation of the authors.
PMC7202009	alignment	For the scRNA-seq reads from Drop-seq platform, STAR has 62.40 % alignment rate, compared to 35.11 % pseudoalignment rate from Kallisto; for the reads from Fluidigm platform, STAR has 66.57 % alignment rate, compared to 34.03 % from Kallisto (Table S1).
PMC7202009	quantification	To generate the count matrix, we used STAR and Kallisto genombam command (Yi) followed by featureCount.
PMC7202009	alignment	Comparison of single-cell gene expression from STAR and Kallisto Alignment.
PMC7202009	quantification	The x-axis represents the expression level with STAR protocol, and the y-axis represents the expression level with Kallisto protocol.
PMC7202009		The intersect regions are the genes that are detected in common by STAR and Kallisto.
PMC7202009	alignment	Comparisons of STAR vs. Kallisto alignment results on Drop-Seq and Fluidigm data Specifically, we first checked the overall correlation of alignments from STAR and Kallisto workflows.
PMC7202009	quantification	We added a pseudo-count of 1 to all gene counts before log transformation, then calculated the Pearson ’ s correlation of all genes across all cells between STAR and Kallisto.
PMC7202009	quantification	As shown in Figure 1A-B, the correlation coefficient between Kallisto and STAR aligned gene counts is 0.836 and 0.862 for Drop-seq and Fluidigm data, respectively, demonstrating a strong concordance between them.
PMC7202009	quantification	However, further examination shows that STAR yields more uniquely expressed genes for both Drop-seq and Fluidigm platforms (Figure 1C-D).
PMC7202009	quantification	For Drop-seq data, STAR and Kallisto detect 16892 common genes, but 7116 and 1906 unique genes, respectively.
PMC7202009	quantification	For Fluidigm data, STAR and Kallisto detect 23193 common genes, but 13710 and 645 unique genes, respectively.
PMC7202009		The modes of the density distribution of gene numbers in each cell (Figure 1E-F) shift to higher values for STAR alignment, confirming that indeed STAR systematically detects more genes in each cell.
PMC7202009	alignment	Overall Kallisto pseudoaligned to more genes (proportion-wise) with shorter length (< 3000bp), whereas STAR can handle longer gene alignment better, as shown in Figure 1G-H.
PMC7202009	alignment	Overall, genes detected using STAR have cumulatively significantly lower dropout rates for both Drop-seq (K-S test p-value < 2.2 e-16) and Fluidigm (K-S test p-value < 2.2e-16) platforms.
PMC7202009	quantification,alignment	In conclusion, despite high correlations between STAR and Kallisto, STAR detected more genes and also yielded more abundant gene expression counts across cells compared to Kallisto.
PMC7202009	alignment	Validation of STAR and Kallisto results using RNA FISH data To address the issue of lack of absolute truth of gene expression in the comparisons above, we next assessed STAR and Kallisto performance using smRNA-FISH data as the ground truth measurement.
PMC7202009	alignment	For Drop-seq data, as shown in Figure 2A-B, both STAR and Kallisto missed detecting some of the 26 genes.
PMC7202009		STAR missed three genes : VEGFC, AXL and WNT5A, whereas Kallisto missed four genes : VEGFC, WNT5A, NGFR, PDGFRB.
PMC7202009		After removing these outliers, the correlation between STAR Gini and smRNA-FISH Gini was 0.50, and the correlation between Kallisto Gini and smRNA-FISH Gini was 0.53.
PMC7202009		For Fluidigm data, as shown in Figure 2C-D, STAR detected all 26 genes, whereas Kallisto missed to detect 3 genes : VEGFC, WNT5A, and NGFR.
PMC7202009		The correlation between STAR Gini and smRNA-FISH Gini is 0.55, whereas the correlation between Kallisto Gini and smRNA-FISH was 0.47 (after removing undetected genes).
PMC7202009		In summary, the comparisons with smRNA-FISH data show that STAR tends to miss fewer genes.
PMC7202009		Moreover, STAR has on par with (Drop-seq) or slightly better (Fluidigm) correlations with smRNA-FISH based truth measure, despite the fact that the Gini coefficients between STAR and Kallisto are highly correlated for commonly detected genes (Figure S2A-B).
PMC7202009		Validation of STAR and Kallisto results using RNA FISH data.
PMC7202009		FISH data with STAR and Kallisto protocols for selected genes.
PMC7202009	alignment	The x-axis represents the Gini coefficient obtained from RNA FISH study, and the y-axis represents the Gini coefficient from the Drop-seq experiment using STAR and Kallisto respectively.
PMC7202009	alignment	FISH data with STAR and Kallisto protocols.
PMC7202009	alignment	Comparison of Bowtie2 vs. STAR and Kallisto alignment We further compared Bowtie2, a popular alignment method on bulk RNA-Seq data, to STAR and Kallisto, on the above mentioned single cell datasets (Figure S3).
PMC7202009	quantification	Among all three methods, STAR still has the most abundant genes detected (24008).
PMC7202009	quantification	However, Bowtie2 has slightly lower detected gene count per cell compared to both STAR and Kallisto (Figure S3 G and H).
PMC7202009	alignment	For Gini coefficient comparison, Bowtie2 has closer coordination with STAR (Figure S3 F), compared to Kallisto (Figure S3 E).
PMC7202009	alignment	STAR failed to detect two of 26 genes, whereas Bowtie2 and Kallisto both missed four genes (Figure S3 D-F).
PMC7202009	alignment	For efficiency, with premade reference, STAR took 15 min and 28 G memory to finish the alignment, whereas Bowtie2 used 56 min and 14.2 G memory.
PMC7202009	alignment	Therefore, by comparing with another traditional aligner Bowtie2, we have reached the consistent conclusion with other aligner benchmarking studies on the real and simulated datasets (Baruzzo); (Teissandier) that STAR is currently one of the top performing methods on alignment rates and speed comparing to other traditional aligners.
PMC7202009	alignment	For Fluidigm data, the experiments were conducted using the same pipelines as in Figures 1 and 2, with the notation that STAR was aligned to reference genome GRCh38 whereas Kallisto used GRCh38 cDNA+intron index.
PMC7202009	alignment	Both Kallisto and STAR were assigned with one processor and 60 GB memory and ran in one thread mode to minimize the effect of parallel processing.
PMC7202009	alignment	Overall, Kallisto pseudoalignment takes 1/4 amount of time as STAR (Figure 3A).
PMC7202009	alignment	Moreover, the maximum memory usage of Kallisto was ∼3.6 GB, which is about 1/7 of the memory usage (∼28 GB) of STAR (Figure 3B).
PMC7202009		(A) Computing time of Fluidigm data (800 cells) on the computer cluster (configurations in Methods section), using STAR and Kallisto.
PMC7202009		(B) Memory usage plot for Fluidigm data using STAR and Kallisto.
PMC7202009	alignment	STARsolo (available in STAR after version 2.7.0) and Kallisto bustools are pipelines developed based on each method to analyze the UMI-based data, such as 10x genomics data.
PMC7202009	clustering	Both pipelines produced the same output format from 10x ’ s Cell Ranger tool, so we next used the Seurat package (Butler; Stuart) to find cell clusters using unsupervised clustering method.
PMC7202009	clustering	STARsolo output resulted in nine clusters (Figure 4A), each assigned to a cell type using the predefined markers in the Seurat vignette.
PMC7202009		90.73 % of cell labels from STARsolo matched with those from Cell Ranger (based on STAR), confirming the similarity between the two workflows (Table S2).
PMC7202009	alignment	In this report, we compared the performance of the two most popular alignment / pseudoalignment methods : STAR and Kallisto.
PMC7202009	alignment	Through the comparison, it appeared that in this dataset, despite high correlations between STAR and Kallisto, STAR tends to yield more genes overall, as well as more abundance of the genes, compared to Kallisto; whereas Kallisto detects more shorter genes compared to STAR.
PMC7202009		Based on the 26 genes that have smRNA FISH results, STAR appeared to detect more of them, compared to Kallisto; also STAR had on-par (Drop-seq) or slightly better (Fluidigm) correlations with the “ reference measure ” of smRNA-FISH.
PMC7202009	marker_genes,classification	We chose this PBMC 3 K dataset because the Seurat group provided marker genes and cell types, so we could use such knowledge as the “ reference ”.
PMC7202009	alignment	The result showed that STAR alignment harvested cell clusters that could be well identified by predefined cell-specific markers; the cell clusters generated from Kallisto with cDNA plus intron (not just cDNA) index information were very similar, but still missed one cell type.
PMC7202009	alignment	in their single cell pipeline evaluation studie : Kallisto with cDNA index had a low fraction of assigned reads, and for UMI-based methods STAR performs better (Vieth).
PMC7202009	alignment	, we also found that Kallisto (with only cDNAs as the reference) is 4 times faster than STARsolo, the recent version of STAR aligner adapter for scRNA-Seq, and the memory usage of Kallisto is 7.7 times less than STARsolo.
PMC7202009	alignment	In summary, based on the datasets used in this study, we conclude that Kallisto ’ s use of computing resources is much less demanding than STAR when only cDNA sequences are used as the reference; however, such efficiency gain is at the cost of loss of information.
PMC7737787	alignment	After removing UMIs and low‐quality bases, the filtered reads were aligned to the human reference genome (hg19) by STAR, and BAM files were prepared by SAMtools.
PMC7737787	quality_control	Individual cells with fewer than 600 covered genes and over 20 % mitochondrial reads were filtered out, and 1986 single cells remained (401 immune cells and 1585 CTCs) for subsequent analysis using the Seurat 3.0 software package (Table 1).
PMC7737787	marker_genes	The cluster‐specific marker genes were identified by the FindAllMarkers function in Seurat 3.0.
PMC7737787	classification,marker_genes	To infer the cell type identity, Seurat 3.0 was used to generate expression heatmaps of selected gene markers of known cell types, including T cells (CD2, CD3D, CD3E, and CD3 G), B cells (CD19, MS4A1, CD79A, and CD79B), monocytes (CD14, CD68, and CD163), lung cells (SFTPA1, SFTPA2, SFTPB, and NAPSA), epithelial cells (EPCAM, CDH1, KRT7, KRT8, KRT18, and MUC1), and proliferation cells (CCND1 and TOP2A).
PMC7737787		Cell cycle analysis Cell cycle assignment was performed in R version 3.6.0 using the CellCycleScoring function included in Seurat 3.0 package.
PMC6417818	visualization	To address the aforementioned challenges, many computational tools have been developed to analyze and visualize high-dimensional scRNA-seq data, including Monocle, Wishbone, SMILE and FVFC.
PMC6417818	visualization	We also compared Mapper with one of the state-of-the-art scRNA-seq visualization methods, Monocle 2.
PMC6417818	visualization,trajectory	Monocle 2 also provides an interesting visualization, where non-malignant cells branch out into two clusters of malignant cells.
PMC7200215		10× Cell Ranger [ 53 ], STAR [ 63 ], Seurat [ 64 ], and SCran [ 65 ]).
PMC8356963		Seurat pipeline Seurat (3.1.0) workflow was performed on E16.5 hippocampal dataset (25) following the Guided Clustering Tutorial www.satijalab.org/seurat/v3.1/pbmc3k_tutorial.html (accessed 20 February 2020), with modifications.
PMC8356963	normalization	The correlation was then calculated on the whole Seurat normalized data matrix and the heatmap was plotted subsetting this (Figure 2B).
PMC8356963	normalization	(B) Pearson correlation matrix of the same selected genes as in (A), using Seurat (34) normalized expression levels (obtained following the website vignettes – Guided Clustering Tutorial).
PMC8356963	variable_genes	(C, D) Comparison between COTAN global differentiation index (GDI, C) and Seurat highly variable features (D) analysis.
PMC8356963		We compared COEX to correlations coefficients computed on gene expression levels, obtained by Seurat (34).
PMC8356963	variable_genes	We then compared GDI to the highly variable feature analysis of Seurat (34).
PMC8356963	variable_genes,classification	While, highly variable features analysis of Seurat (Figure 2D) was much less precise in discriminating between CGs and cell identity genes (compare Figure 2D to C) with, for example, the two neuronal markers Dcx and Map2 close to Gadph and Sub1.
PMC8356963	differential_expression	The GDI is a useful tool to detect differentially expressed genes, similarly to Seurat ’ s variable features, but with constitutive genes and not-constitutive genes more separated (as shown in Figure 2).
PMC5465230	quantification	Popular tools for assessing relative expression include Cufflinks [ 53–55 ], and STAR [ 56 ].
PMC5465230		Monocle, developed by Trapnell et al.
PMC5465230	trajectory	Additionally, they performed pseudo-time ordering using Monocle with their consensus-ordering genes and found similar dynamic gene expression related to quiescence and activation of NSCs.
