820878
doi
10.1109/BigData.2016.7840618
oai:zenodo.org:820878
user-linkeddata
user-eu
user-bioexcel
Jung, Segun
The University of Chicago
D'Arcy, Mike
University of Southern California, Los Angeles, CA, USA
Heavner, Ben
Institute for Systems Biology, Seattle, WA, USA
Foster, Ian
The University of Chicago and Argonne National Laboratory, Chicago IL, USA
Kesselman, Carl
University of Southern California, Los Angeles, CA, USA
Madduri, Ravi
The University of Chicago and Argonne National Laboratory, Chicago IL, USA
Rodriguez, Alexis
The University of Chicago and Argonne National Laboratory, Chicago IL, USA
Soiland-Reyes, Stian
The University of Manchester, Manchester, UK
Goble, Carole
The University of Manchester, Manchester, UK
Clark, Kristi
University of Southern California, Los Angeles, CA, USA
Deutsch, Eric W.
Institute for Systems Biology, Seattle, WA, USA
Dinov, Ivo
The University of Michigan, Ann Arbor, MI, USA
Price, Nathan
Institute for Systems Biology, Seattle, WA, USA
Toga, Arthur
University of Southern California, Los Angeles, CA, USA
I'll take that to go: Big data bags and minimal identifiers for exchange of large, complex datasets
Chard, Kyle
The University of Chicago and Argonne National Laboratory, Chicago IL, USA
url:https://static.aminer.org/pdf/fa/bigdata2016/BigD418.pdf
url:https://www.research.manchester.ac.uk/portal/files/45989205/bagminid.pdf
url:http://bd2k.ini.usc.edu/tools/
url:https://github.com/ini-bdds/bdbag
url:https://www.research.manchester.ac.uk/portal/en/publications/ill-take-that-to-go(8335e672-1d85-4649-a245-56fbdb1bd423).html
url:https://w3id.org/ro/bagit
info:eu-repo/semantics/openAccess
Creative Commons Attribution 4.0 International
https://creativecommons.org/licenses/by/4.0/legalcode
Big Data
data analysis
BDBags
Big Data analysis
Big Data bags
Big Data sharing
Minid
data assembling
data collections
data descriptions
datasets
identifiers
research objects
Encoding
Metadata
Payloads
Robustness
Software
Uniform resource locators
bdbag
<p><em>Big data workflows</em> often require the assembly and exchange of complex, multi-element datasets. For example, in biomedical applications, the input to an analytic pipeline can be a dataset consisting thousands of images and genome sequences assembled from diverse repositories, requiring a description of the contents of the dataset in a concise and unambiguous form. Typical approaches to creating datasets for big data workflows assume that all data reside in a single location, requiring costly data marshaling and permitting errors of omission and commission because dataset members are not explicitly specified.</p>
<p>We address these issues by proposing simple methods and tools for assembling, sharing, and analyzing large and complex datasets that scientists can easily integrate into their daily workflows. These tools combine a simple and robust method for describing data collections (<strong>BDBags</strong>), data descriptions (<strong>Research Objects</strong>), and simple persistent identifiers (<strong>Minids</strong>) to create a powerful ecosystem of tools and services for big data analysis and sharing.</p>
<p>We present these tools and use biomedical case studies to illustrate their use for the rapid assembly, sharing, and analysis of large datasets.</p>
IEEE
2016-12-05
info:eu-repo/semantics/conferencePaper
820877
user-linkeddata
user-eu
user-bioexcel
award_title=Centre of Excellence for Biomolecular Research; award_number=675728; award_identifiers_scheme=url; award_identifiers_identifier=https://cordis.europa.eu/projects/675728; funder_id=00k4n6c32; funder_name=European Commission;
1579541816.085578
713184
md5:91195ab648922564b86d629e83ea88d8
https://zenodo.org/records/820878/files/bagminid.pdf
public
https://static.aminer.org/pdf/fa/bigdata2016/BigD418.pdf
Is identical to
url
https://www.research.manchester.ac.uk/portal/files/45989205/bagminid.pdf
Is identical to
url
http://bd2k.ini.usc.edu/tools/
Is supplemented by
url
https://github.com/ini-bdds/bdbag
Is supplemented by
url
https://www.research.manchester.ac.uk/portal/en/publications/ill-take-that-to-go(8335e672-1d85-4649-a245-56fbdb1bd423).html
Is part of
url
https://w3id.org/ro/bagit
Cites
url
2016 IEEE International Conference on Big Data (Big Data)
978-1-4673-9005-7
319-328
2016-12-05