Published August 31, 2015 | Version v1

Codon similarity data in ATTED-II ver 8.0 (Ptr, Zma)

  • 1. Tohoku University

Description

Codon similarity data in ATTED-II ver 8.0

The gene-to-gene codon similarity data is organized in the form of tables, each named according to the Entrez Gene ID of a particular query gene. Each table encompasses three columns, specifying: the Entrez Gene ID of a corresponding gene, an MR (Mutual Rank) value (where a smaller number signifies a stronger relationship), and a Pearson correlation coefficient (where a larger number suggests a stronger association).

Protein-coding sequences utilized in this study were retrieved from NCBI's RefSeq database. For each gene, a 61-dimensional vector was derived from the count of codons in the protein-coding sequence. In instances where multiple RefSeq sequences were associated with a single gene, the longest sequence was selected for the codon usage calculation. Pearson correlation coefficients (PCCs) were determined between the vectors of any two given genes. These PCCs were subsequently converted into MRs, employed as an index to evaluate the similarity in codon usage between the genes.

Files

codon_Ptr.v15-08.G41366-S61.codon.mrgeo.d.zip

Files (37.9 GB)

Name Size
md5:a970789a1c030163473ae6b24a2c7bf5
15.3 GB Preview Download
md5:bfeffd283df0b76e7fb8c1c12cf0da32
22.6 GB Preview Download

Additional details

Related works

Is supplement to
Journal article: 10.1093/pcp/pcv165 (DOI)