Phylogenomic analyses reveal that Panguiarchaeum is a clade of genome-reduced Asgard archaea within the Njordarchaeia
Authors/Creators
Description
Abstract
The Asgard archaea are a diverse archaeal phylum important for our understanding of cellular evolution because they include the lineage that gave rise to eukaryotes. Recent phylogenomic work has focused on characterising the diversity of Asgard archaea in an effort to identify the closest extant relatives of eukaryotes. However, resolving archaeal phylogeny is challenging, and the positions of two recently-described lineages - Njordarchaeales and Panguiarchaeales - are uncertain, in ways that directly bear on hypotheses of early evolution. In initial phylogenetic analyses, these lineages branched either with Asgards or with the distantly-related Korarchaeota, and it has been suggested that their genomes may be affected by metagenomic contamination. Resolving this debate is important because these clades include genome-reduced lineages that may help inform our understanding of the evolution of symbiosis within Asgard archaea. Here, we performed phylogenetic analyses revealing that the Njordarchaeales and Pangiuarchaeales constitute the new class Njordarchaeia within Asgard archaea. We found no evidence of metagenomic contamination affecting phylogenetic analyses. Njordarchaeia exhibit hallmarks of adaptations to (hyper-)thermophilic lifestyles, including biased sequence compositions that can induce phylogenetic artifacts unless adequately modelled. Panguiarchaeum is metabolically distinct from its relatives, with reduced metabolic potential and various auxotrophies. Phylogenetic reconciliation recovers a complex common ancestor of Asgard archaea that encoded the Wood-Ljungdahl pathway. The subsequent loss of this pathway during the reductive evolution of Panguiarchaeum may have been associated with the switch to a symbiotic lifestyle based on H2-syntrophy. Thus, Panguiarchaeum may contain the first obligate symbionts within Asgard archaea.
Table of contents
Repository Contents
1_Genome_files.tar.gz:
- Folder 'faa': this folder contains all protein sequence files for 966 archaeal, 1325 bacteria, and 137 eukaryotic MAGs/genomes/largely complete transcriptomes used in this study.
- Folder 'annotations': this folder contains the protein annotation results for the protein sequence described above.
2_phylogenies.tar.gz
- 1_species_trees: this folder contains all alignments and tree files for the species tree inference using different datasets and single gene tree files used for gene tree inspection and ranking. Files are organised as follows and are associated with the corresponding parts of the manuscript: Figure 1, Supplementary Figures 1, 5-7, 10-13.
-
- Folder '1_single_gene_tree_inspection' includes all treefiles and pdf for single gene tree inspection.
- Folder '2_ranking' includes all treefiles and split scores from marker ranking.
- Folder '3_concatenation' includes, unaligned sequences (under subfolder: .faa unaligned) untrimmed alignment files (under subfolder mafft), trimmed alignment files (under subfolder bmge), site likelihood files (.sitelh, under subfolder trees_sitelh) and treefiles (constrained and unconstrained trees_sitelh/treefiles) inferred under different models (See Methods) based on different datasets (subdirectories: 966 taxa, 303 taxa and 71 taxa)
- 2_single_gene_trees_ESP_and_other: this folder contains unaligned sequences, untrimmed alignments, trimmed alignment files and treefiles for single gene tree inference of ESP proteins, and other single gene trees. Files are organised as individual folders for one single gene tree phylogeny and are associated with the corresponding parts of the manuscript: Supplementary Figures 2-4, 15-20 and 28-33. For detail commands, refers to 3_workflow_scripts/2_phylogeneitc_analyses.
3_workflow_scripts.tar.gz
- 1_workflows: this folder includes workflows (.qmd) used in this study.
- 1_marker_inspection.qmd: this document includes code to generate files for marker gene inspection.
- 2_phylogenetic_analyses.qmd: this document includes commands and examples for phylogenetic inference.
- 3_annotations_workflow.qmd: this document includes the workflow for annotating protein sequences against different databases, analysis of gene presence and absence profiles, and amino acid composition.
- 4_ALE_workflow.qmd: this document includes the workflow for reconciliation analyses.
- *2_scripts: this folder contains Python and bash scripts used in 1_workflows, and example scripts for phylogenetic tree visualisation.
4_ALE.tar.gz:
- reconcilliations-TableEvents_clean.7z: This file summarises the reconciliation results using default parameters.
- OR_recon-TableEvents_clean.7z: This file summarises the reconciliation results based on per-arCOG-category optimised origination rate (See Methods).
- SpeciesTreeRef.newick: species tree used in the reconciliation, with internal node ids corresponding in the files described above.
5_contamination_assessment.tar.gz:
- anvio8_workflow.sh: this bash file documents examples of how contaminiation assessment was performed following anvi'o's tutorial
- contigs_db: anvi'o contig databases for all 4 examined MAGs
- profiles_db: anvi'o profile databases for all 4 examined MAGs
Notes
Files
Additional details
Identifiers
Dates
- Created
-
2025-01-31