Research Discipline Classification with Supervised Learning
Authors/Creators
- 1. Leibniz Supercomputing Centre
- 2. CAU Kiel
- 3. LMU München
Description
One of the biggest technical challenges for services providers is to scale their quality of service with the amount of research data. Complete information and the usage of standardized vocabularies improve the quality of metadata and the service as a whole. Studies have shown that services often offer incomplete metadata or do not use well-established standards, such as the Dewey decimal classification of research disciplines. This poster presents one approach to meet this challenge, namely the usage of machine learning for metadata enrichment. Classification of the discipline of research data is taken as an example how trained models can help in improving metadata quality. Over 16 million metadata records were filtered and processed. The resulting 100.000 records had sufficient quality to choose a training set and an evaluation set. The poster will show the experiences with model configurations and their precision in classifying research disciplines. The model, its configuration and the sampled data will be available for research data repositories to build enrichment services that will help to boost discoverability of research data products with discipline classification.
Files
2019_04_01_weber_print_w-out_cm.pdf
Files
(764.4 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:02e781bbdc1356b387e1efa606c24517
|
764.4 kB | Preview Download |