Data-Driven Characterization of MEDLINE Indexing Practices Using Open-Source and Publicly Available Tools
Description
In 2024 the National Library of Medicine (NLM) transitioned to using the Medical Text Indexer – NeXt Generation (MTI-X) to index all new records in its public-facing MEDLINE database. Record metadata indicates if MTI-X indexing is retained, replaced entirely by human indexing or modified by human curators; the latter applies the status of “curated” to the record. User requests for more information on curation, and whether certain subjects are prioritized for human review, have not yet been formally answered. This new, more opaque system of indexing has resulted in uncertainty around the use of indexing terms in the regular workflows of health researchers, including librarians. This project used publicly accessible data and tools to illuminate the indexing and curation practices of the NLM. Metadata for the 25,439 records created between September 1st, 2025-September 5th, 2025, was collected and re-collected at regular intervals using the NLM E-Utilities API. To date, researchers have collected a total of 2,913,807 data points. The collection of data at different intervals allowed researchers to create the first per-record change log from MEDLINE data. The development and analysis of these changes was carried out in R version 4.5.1. With these results, researchers are now able to identify trends in curated indexing, with a view towards reverse engineering curation practices at the NLM. This presentation will provide a detailed walkthrough of the methodology and tools used to plan, organize and execute this project. It will highlight the utility of big data within an information science context and showcase resources that can be used to apply these same principles to other data-driven projects.
Files
Garlock_IASSIST_June2026-dist.pdf
Files
(402.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:db27b74b25f36d9078cfee9c9f1faa68
|
402.0 kB | Preview Download |