Published November 24, 2016 | Version v1

Leveraging the CEDAR Workbench for Ontology-linked Submission of Adaptive Immune Receptor Repertoire Data to the Sequence Read Archive (SRA)

  • 1. Department of Pathology, Yale School of Medicine, New Haven, CT, USA
  • 2. Center for Expanded Data Annotation and Retrieval, Stanford Center for Biomedical Informatics Research
  • 3. Department of Emergency Medicine; Yale Center for Medical Informatics, Yale University School of Medicine, New Haven, CT, USA.
  • 4. Department of Pathology, Yale School of Medicine, New Haven, CT, USA.

Description

Next-generation sequencing technologies have led to a rapid production of high-throughput sequence data characterizing adaptive immune-receptor repertoires (AIRRs). As part of the AIRR community (http://airr-community.org) data standards working group, we have developed an initial set of metadata recommendations for publishing AIRR sequencing studies. These recommendations will be implemented in several public repositories, including the NCBI sequence read archive (SRA). Submissions to SRA typically use a flat-file template and include only a minimal amount of term validation. In order to ease the metadata authoring and to implement the ontological terms validation of repertoire sequence data, we are developing an interactive template through CEDAR workbench that will allow for ontological validation, and subsequent deposition in SRA. CEDAR workbench also allows the user to populate the template with metadata for data submission to various data repositories. The incorporation of template-element level ontology mapping not only facilitates validation of data submission, but also enables intelligent queries within and across repositories.

Files

Leveraging the CEDAR Workbench for Ontology-linked Submission of Adaptive Immune Receptor Repertoire Data to the Sequence Read ArchiveĀ (SRA) .pdf