Published May 4, 2025
| Version 1.0
Software
Open
hibernator11/workshop-notebooks-wac-dhnb2025: Version 1
Authors/Creators
Description
This project has been developed as part of the workshop "Web Archives Collections as data" for the DHNB 2025 conference. The coordinators of the workshop are:
- Olga Holownia , International Internet Preservation Consortium
- Gustavo Candela, University of Alicante, Spain
- Helena Byrne, British Library, UK
- Jon Carlstedt Tønnessen, National Library of Norway, Norway
- Anders Klindt Myrvoll, Royal Danish Library, Denmark
- Sophie Ham, National Library of the Netherlands, Netherlands
- Steven Claeyssens, National Library of the Netherlands, Netherlands
- Mahendra Mahey, Tallinn University, Estonia
This project provides examples of code based on Jupyter Notebooks to extract Web Archive Content to create collections as data. This examples facilitate the extraction and reuse of datasets of text extracted from all available captures of archived web pages. The datasets can be employed to analyse changes over time and to identify the most relevant words.
Files
hibernator11/workshop-notebooks-wac-dhnb2025-1.0.zip
Files
(8.5 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:d269e98dddedafa5d6d751a3acee9c6d
|
8.5 MB | Preview Download |
Additional details
Related works
- Is supplement to
- Software: https://github.com/hibernator11/workshop-notebooks-wac-dhnb2025/tree/1.0 (URL)