Documents first indexed in the SLUB catalog in 2022
Description
The data set is available as a gzip-compressed, line-delimited JSON file and contains the 7,127,497 documents first indexed in the SLUB catalog in 2022 with the fields id (document identifier) and first_indexed (timestamp). The id consists of a creator id (optional), source id and a record id according to the scheme [{creator_id}-]{source_id}-{record_id}. If the record id contains characters that are unsuitable for URLs, it is base64-encoded without any padding. The first_indexed field, which is part of VuFind's Solr index schema, was determined using a Django application that regularly monitors the Solr cores of the SLUB catalog. The respective document sets from different points in time are compared with each other in order to determine new documents, i.e. document identifiers. If a new document is found, first_indexed is assigned the timestamp from the last_indexed field. The field obtained in this way serves as the basis for creating a list of new titles. However, it should be noted that neither all first indexed documents necessarily represent new titles, nor do the documents contained in this data set still have to be in the catalog. To check whether a document is currently in the SLUB catalog, its detailed view can be retrieved using the following URL scheme: https://katalog.slub-dresden.de/id/{id}. Example: https://katalog.slub-dresden.de/id/0-173837243X.
Files
Files
(91.8 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:6c179033a8f79f1ee92dabe0cd6e4aac
|
91.8 MB | Download |
Additional details
Dates
- Collected
-
2023-05-25monitor.solradd