There is a newer version of the record available.

Published January 26, 2024 | Version 0.9

The North Pacific Eukaryotic Gene Catalog: clustered nucleotide metatranscripts and read counts

Description

This data continues with the development of the NPEGC Trinity de novo metatranscriptome assemblies from the protein data repository of The North Pacific Eukaryotic Gene Catalog. The nucleotide sequences corresponding to the NPEGC cluster representatives are collected together in these repository files:

NPac.G1PA.bf100.id99.nt.fasta.gz
NPac.G2PA.bf100.id99.nt.fasta.gz
NPac.G3PA.bf100.id99.nt.fasta.gz
NPac.G3PA_diel.bf100.id99.nt.fasta.gz
NPac.D1PA.bf100.id99.nt.fasta.gz

These nucleotide sequences have been sourced from the Zenodo repository for raw assemblies: The North Pacific Eukaryotic Gene Catalog: Raw assemblies from Gradients 1, 2 and 3

Key processing steps are sampled below with links to the detailed code on the main github code repository: https://github.com/armbrustlab/NPac_euk_gene_catalog


Code used to build the kallisto indices and map the short reads against indices with kallisto are online in the code repository here: NPEGC.nt_kallisto_counts.sh

There are two main steps:
1. Generate the kallisto index on the sets of clustered nucleotide metatranscripts
2. Map the short reads from environmental samples back to the assembly index

As generated above, kallisto generates separate results files for each of the sample files. Even after compression, the total size of the tarballed kallisto output results directories are prohibitively large (>50GB). We use the code in this template R script to join together the 'est_count' estimated count values for the tens of millions of protein sequences in each project metatranscriptome, along with length.

The code in this template script was used for each project: aggregate_kallisto_counts.R
The output count files for each project are Gzip-compressed and uploaded to the NPEGC nucleotide data repository here: 

G1PA.raw.est_counts.csv.gz
G2PA.raw.est_counts.csv.gz
G3PA.raw.est_counts.csv.gz
G3PA_diel.raw.est_counts.csv.gz
D1PA.raw.est_counts.csv.gz

Files

Files (29.1 GB)

Name Size
md5:dc5f1768813ded1972d9df9bf23ea894
1.0 GB Download
md5:13e11b78c3a97b995b02afcc430c0414
712.6 MB Download
md5:48d1326b68e2c8294628c38bc67eecc0
1.2 GB Download
md5:740e115d789023a57c1a0338dba062fc
590.0 MB Download
md5:6020a4eb78c50324867b1b7d678fd2a1
517.0 MB Download
md5:9fb0eadbd3f44f508f3b1e238cedcddc
6.1 GB Download
md5:cc3959a5cc6479620c971688d644ce17
4.4 GB Download
md5:a2d2c64ff5c06a88e85a9cc8b4ab4298
5.9 GB Download
md5:501b5beb24a30a8c5e6bad378ec4f497
5.2 GB Download
md5:e509add218dbc2a4a0d9c8f3b04910bc
3.5 GB Download

Additional details

Related works

Is supplement to
Dataset: 10.5281/zenodo.10472590 (DOI)