There is a newer version of the record available.

Published April 20, 2023 | Version 0.1.1

Proboscidean Palaeoproteomic Reference Dataset

Authors/Creators

  • 1. Section for Molecular Ecology and Evolution , Globe Institute, University of Copenhagen

Description

This entry contains the 'Proboscidean Palaeoproteomic Reference Dataset'.

We used PaleoProPhyler ( https://github.com/johnpatramanis/Proteomic_Pipeline ) to generate a palaeoproteomic reference dataset of protein sequences from ancient and present-day Proboscidae. Using the first two modules of PaleoProPhyler, we translated more than 35 publicly available whole genomes from extant and extinct species. Details on the processing of the sequences can be found below.

 

Which individuals / species are included?

The full list of individuals, the original fastq repository location and the species included in the dataset are contained within the tab seperated file 'METATADATA.txt', that also contains headers. Most individuals of the dataset belong to one of these 3 species: Loxodonta africana, Elephas maximusMammuthus  primigenius. 

 

Which Proteins are included?

We compiled a small initial list of 262 proteins that had been indentified in either teeth, bone or items made out of ivory. For each protein, both the canonical and all alternative protein coding isoforms (based on the Loxodonta africana reference proteome of Ensembl) were translated, leading to more than 350 unique protein sequences for each individual in the dataset. The protein list is available in the file 'proteins.txt'

Other files included in the zip folder:

- ALL_PROT_REFERENCE.fa contains all of the sequences generated as part of the Proboscidean Palaeoproteomic Reference Dataset described above

- PER_PROTEIN is a folder containing one fasta file for each protein within the Proboscidean Palaeoproteomic Reference Dataset, each protein fasta file has the sequences of all individuals for that particular protein

- PER_SAMPLE is a folder containing one fasta file for each sample/individual within the Proboscidean Palaeoproteomic Reference Dataset, each sample fasta file has the sequences of all proteins for that particular sample.

Files

PRD.zip

Files (5.2 MB)

Name Size Download all
md5:155078a2603861245365737e626dd0f6
5.2 MB Preview Download

Additional details

Funding

European Commission
PUSHH - Palaeoproteomics to Unleash Studies on Human History 861389