10.5281/zenodo.3687262
https://zenodo.org/records/3687262
oai:zenodo.org:3687262
Ruest, Nick
Nick
Ruest
0000-0003-1891-1112
York University
Gagné, Carole
Carole
Gagné
Bibliothèque et Archives nationales du Québec
Mitchell, Dave
Dave
Mitchell
Bibliothèque et Archives nationales du Québec
Coalition Avenir Québec (CAQ) web archive collection derivatives
Zenodo
2020
web archives
parquet
dataframes
2020-02-25
10.5281/zenodo.3687261
https://zenodo.org/communities/wahr
Creative Commons Attribution 4.0 International
Web archive derivatives of the Coalition Avenir Québec (CAQ) collection from the Bibliothèque et Archives nationales du Québec. The derivatives were created with the Archives Unleashed Toolkit. Merci beaucoup BAnQ!
These derivatives are in the Apache Parquet format, which is a columnar storage format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See this notebook for examples.
Domains
.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)
Produces a DataFrame with the following columns:
domain
count
Web Pages
.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))
Produces a DataFrame with the following columns:
crawl_date
url
mime_type_web_server
mime_type_tika
content
Web Graph
.webgraph()
Produces a DataFrame with the following columns:
crawl_date
src
dest
anchor
Image Links
.imageLinks()
Produces a DataFrame with the following columns:
src
image_url
Binary Analysis
Audio
Images
PDFs
Presentation program files
Spreadsheets
Text files
Videos
Word processor files