The North Pacific Eukaryotic Gene Catalog: clustered nucleotide metatranscripts and read counts
Authors/Creators
Description
This data continues with the development of the NPEGC Trinity de novo metatranscriptome assemblies from the protein data repository of The North Pacific Eukaryotic Gene Catalog. The nucleotide sequences corresponding to the NPEGC cluster representatives are collected together in these repository files:
NPac.G1PA.bf100.id99.nt.fasta.gz
NPac.G2PA.bf100.id99.nt.fasta.gz
NPac.G3PA.bf100.id99.nt.fasta.gz
NPac.G3PA_diel.bf100.id99.nt.fasta.gz
NPac.D1PA.bf100.id99.nt.fasta.gz
These nucleotide sequences have been sourced from the Zenodo repository for raw assemblies: The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3
Key processing steps are sampled below with links to the detailed code on the main github code repository: https://github.com/armbrustlab/NPac_euk_gene_catalog
Code used to build the kallisto indices and map the short reads against indices with kallisto are online in the code repository here: NPEGC.nt_kallisto_counts.sh
There are two main steps:
1. Generate the kallisto index on the sets of clustered nucleotide metatranscripts
2. Map the short reads from environmental samples back to the assembly index
As generated above, kallisto generates separate results files for each of the sample files. Even after compression, the total size of the tarballed kallisto output results directories are prohibitively large (>50GB). We use the code in this template R script to join together the 'est_count' estimated count values for the tens of millions of protein sequences in each project metatranscriptome, along with length.
The code in this template script was used for each project: aggregate_kallisto_counts.R
The output count files for each project are Gzip-compressed and uploaded to the NPEGC nucleotide data repository here:
G1PA.raw.est_counts.csv.gz
G2PA.raw.est_counts.csv.gz
G3PA.raw.est_counts.csv.gz
G3PA_diel.raw.est_counts.csv.gz
D1PA.raw.est_counts.csv.gz
Files
Files
(29.1 GB)
| Name | Size | |
|---|---|---|
|
md5:dc5f1768813ded1972d9df9bf23ea894
|
1.0 GB | Download |
|
md5:13e11b78c3a97b995b02afcc430c0414
|
712.6 MB | Download |
|
md5:48d1326b68e2c8294628c38bc67eecc0
|
1.2 GB | Download |
|
md5:740e115d789023a57c1a0338dba062fc
|
590.0 MB | Download |
|
md5:6020a4eb78c50324867b1b7d678fd2a1
|
517.0 MB | Download |
|
md5:9fb0eadbd3f44f508f3b1e238cedcddc
|
6.1 GB | Download |
|
md5:cc3959a5cc6479620c971688d644ce17
|
4.4 GB | Download |
|
md5:a2d2c64ff5c06a88e85a9cc8b4ab4298
|
5.9 GB | Download |
|
md5:501b5beb24a30a8c5e6bad378ec4f497
|
5.2 GB | Download |
|
md5:e509add218dbc2a4a0d9c8f3b04910bc
|
3.5 GB | Download |
Additional details
Related works
- Is supplement to
- Dataset: 10.5281/zenodo.10472590 (DOI)