Published May 4, 2025 | Version 1.0

hibernator11/workshop-notebooks-wac-dhnb2025: Version 1

  • 1. University of Alicante
  • 2. National Library of Norway
  • 3. ROR icon Tallinn University

Description

This project has been developed as part of the workshop "Web Archives Collections as data" for the DHNB 2025 conference. The coordinators of the workshop are:

  • Olga Holownia , International Internet Preservation Consortium
  • Gustavo Candela, University of Alicante, Spain
  • Helena Byrne, British Library, UK
  • Jon Carlstedt Tønnessen, National Library of Norway, Norway
  • Anders Klindt Myrvoll, Royal Danish Library, Denmark
  • Sophie Ham, National Library of the Netherlands, Netherlands
  • Steven Claeyssens, National Library of the Netherlands, Netherlands
  • Mahendra Mahey, Tallinn University, Estonia

This project provides examples of code based on Jupyter Notebooks to extract Web Archive Content to create collections as data. This examples facilitate the extraction and reuse of datasets of text extracted from all available captures of archived web pages. The datasets can be employed to analyse changes over time and to identify the most relevant words.

Files

hibernator11/workshop-notebooks-wac-dhnb2025-1.0.zip

Files (8.5 MB)

Additional details