AMPSphere pre-computed resources and auxiliary files for the manuscript codes
Authors/Creators
- 1. Institute of Science and Technology for Brain-inspired Intelligence, Fudan University, Shanghai, China
- 2. Departments of Bioengineering and Chemical and Biomolecular Engineering, School of Engineering and Applied Science, University of Pennsylvania, Philadelphia, PA, USA
- 3. Centro de Biotecnología y Genómica de Plantas, Universidad Politécnica de Madrid (UPM) - Instituto Nacional de Investigación y Tecnología Agraria y Alimentaria (INIA-CSIC), Madrid, Spain
- 4. Structural and Computational Biology Unit, European Molecular Biology Laboratory, Heidelberg, Germany
Description
AMPSphere is a comprehensive catalog of antimicrobial peptides predicted using Macrel (DOI: 10.7717/peerj.10555) from 63,410 public metagenomes, ProGenomes v2.2 database (82,400 high-quality microbial genomes), and c.a. 4k non-whitelisted microbial genomes from NCBI. Currently, AMPSphere is available as a web resource at https://ampsphere.big-data-biology.org/. AMPSphere v.2022-03 contains 863,498 sequences (avg length: 36 amino acids, range 8-98). DRAMP database was used to find confirmed sequences with strict homology to reference. This approach showed that 2,488 peptides were previously confirmed in our dataset. The present repository is a data dump for the precomputed resources and files needed for its generation and analysis as a complement to the GitHub repository. The complementary documentation is also available for each one of the files. To use the files just download them and apply the command `untar` to decompress the folders.