Published April 10, 2019 | Version v1

Research Discipline Classification with Supervised Learning

  • 1. Leibniz Supercomputing Centre
  • 2. CAU Kiel
  • 3. LMU München

Description

One of the biggest technical challenges for services providers is to scale their quality of service with the amount of research data. Complete information and the usage of standardized vocabularies improve the quality of metadata and the service as a whole. Studies have shown that services often offer incomplete metadata or do not use well-established standards, such as the Dewey decimal classification of research disciplines. This poster presents one approach to meet this challenge, namely the usage of machine learning for metadata enrichment. Classification of the discipline of research data is taken as an example how trained models can help in improving metadata quality. Over 16 million metadata records were filtered and processed. The resulting 100.000 records had sufficient quality to choose a training set and an evaluation set. The poster will show the experiences with model configurations and their precision in classifying research disciplines. The model, its configuration and the sampled data will be available for research data repositories to build enrichment services that will help to boost discoverability of research data products with discipline classification.

Files

2019_04_01_weber_print_w-out_cm.pdf

Files (764.4 kB)

Name Size Download all
md5:02e781bbdc1356b387e1efa606c24517
764.4 kB Preview Download