Dataset Open Access

Burke Library New York City Religions web archive collection derivatives

Ruest, Nick; Baker, Matthew C.; Thurman, Alex

Web archive derivatives of the Burke Library New York City Religions collection from Columbia University Libraries. The derivatives were created with the Archives Unleashed Toolkit and Archives Unleashed Cloud.

The cul-1945-parquet.tar.gz derivatives are in the Apache Parquet format, which is a columnar storage format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See this notebook for examples.

Domains

.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)

Produces a DataFrame with the following columns:

  • domain
  • count

Web Pages

.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))

Produces a DataFrame with the following columns:

  • crawl_date
  • url
  • mime_type_web_server
  • mime_type_tika
  • content

Web Graph

.webgraph()

Produces a DataFrame with the following columns:

  • crawl_date
  • src
  • dest
  • anchor

Image Links

.imageLinks()

Produces a DataFrame with the following columns:

  • src
  • image_url

Binary Analysis

  • Images
  • PDFs
  • Presentation program files
  • Spreadsheets
  • Text files
  • Word processor files
     

The cul-1945-auk.tar.gz derivatives are the standard set of web archive derivatives produced by the Archives Unleashed Cloud.

  • Gephi file, which can be loaded into Gephi. It will have basic characteristics already computed and a basic layout.
  • Raw Network file, which can also be loaded into Gephi. You will have to use that network program to lay it out yourself.
  • Full text file. In it, each website within the web archive collection will have its full text presented on one line, along with information around when it was crawled, the name of the domain, and the full URL of the content.
  • Domains count file. A text file containing the frequency count of domains captured within your web archive.

Files (36.3 GB)
Name Size
cul-1945-auk.tar.gz
md5:75a4a021e2bd261a971fae1ef4bd4092
15.9 GB Download
cul-1945-parquet.tar.gz
md5:d36351202c8fc7e32f8e6a3428b39c86
20.4 GB Download
36
5
views
downloads
All versions This version
Views 3636
Downloads 55
Data volume 79.7 GB79.7 GB
Unique views 3535
Unique downloads 55

Share

Cite as