Stonewall 50 Commemoration web archive collection derivatives

Ruest, Nick; Dolkart, Andrew S.; Thurman, Alex

doi:10.5281/zenodo.3631347

Published January 30, 2020 | Version v1

Dataset Open

Stonewall 50 Commemoration web archive collection derivatives

1. York University
2. Columbia University

Web archive derivatives of the Stonewall 50 Commemoration collection from Columbia University Libraries. The derivatives were created with the Archives Unleashed Toolkit and Archives Unleashed Cloud.

The cul-12143-parquet.tar.gz derivatives are in the Apache Parquet format, which is a columnar storage format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See this notebook for examples.

Domains

.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)

Produces a DataFrame with the following columns:

domain
count

Web Pages

.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))

Produces a DataFrame with the following columns:

crawl_date
url
mime_type_web_server
mime_type_tika
content

Web Graph

.webgraph()

Produces a DataFrame with the following columns:

crawl_date
src
dest
anchor

Image Links

.imageLinks()

Produces a DataFrame with the following columns:

src
image_url

Binary Analysis

Audio
Images
PDFs
Presentation program files
Spreadsheets
Text files
Word processor files

The cul-12143-auk.tar.gz derivatives are the standard set of web archive derivatives produced by the Archives Unleashed Cloud.

Gephi file, which can be loaded into Gephi. It will have basic characteristics already computed and a basic layout.
Raw Network file, which can also be loaded into Gephi. You will have to use that network program to lay it out yourself.
Full text file. In it, each website within the web archive collection will have its full text presented on one line, along with information around when it was crawled, the name of the domain, and the full URL of the content.
Domains count file. A text file containing the frequency count of domains captured within your web archive.

Files

Files (2.8 GB)

Name	Size
cul-12143-auk.tar.gz md5:55cdbda3b11e7229c5dd7929bf40141d	1.1 GB	Download
cul-12143-parquet.tar.gz md5:2f58d9e884dded855d1222052e218288	1.7 GB	Download

	All versions	This version
Views	626	625
Downloads	192	192
Data volume	325.3 GB	325.3 GB

Stonewall 50 Commemoration web archive collection derivatives

Authors/Creators

Description

Files

Files (2.8 GB)