Dataset Open Access

Queer Japan Web Archive collection derivatives

Ruest, Nick; Yanagihara, Yoshie; Shida, Tetsuyuki; Nakamura, Haruko; Abrams, Samantha


Dublin Core Export

<?xml version='1.0' encoding='utf-8'?>
<oai_dc:dc xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:creator>Ruest, Nick</dc:creator>
  <dc:creator>Yanagihara, Yoshie</dc:creator>
  <dc:creator>Shida, Tetsuyuki</dc:creator>
  <dc:creator>Nakamura, Haruko</dc:creator>
  <dc:creator>Abrams, Samantha</dc:creator>
  <dc:date>2020-01-31</dc:date>
  <dc:description>Web archive derivatives of the Queer Japan Web Archive collection from the Ivy Plus Libraries Confederation. The derivatives were created with the Archives Unleashed Toolkit and Archives Unleashed Cloud.

The ivy-12172-parquet.tar.gz derivatives are in the Apache Parquet format, which is a columnar storage format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See this notebook for examples.

Domains

.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)

Produces a DataFrame with the following columns:


	domain
	count


Web Pages

.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))

Produces a DataFrame with the following columns:


	crawl_date
	url
	mime_type_web_server
	mime_type_tika
	content


Web Graph

.webgraph()

Produces a DataFrame with the following columns:


	crawl_date
	src
	dest
	anchor


Image Links

.imageLinks()

Produces a DataFrame with the following columns:


	src
	image_url


Binary Analysis


	Audio
	Images
	PDFs
	Presentation program files
	Spreadsheets
	Text files
	Videos
	Word processor files
	 


The ivy-11854-auk.tar.gz derivatives are the standard set of web archive derivatives produced by the Archives Unleashed Cloud.


	Gephi file, which can be loaded into Gephi. It will have basic characteristics already computed and a basic layout.
	Raw Network file, which can also be loaded into Gephi. You will have to use that network program to lay it out yourself.
	Full text file. In it, each website within the web archive collection will have its full text presented on one line, along with information around when it was crawled, the name of the domain, and the full URL of the content.
	Domains count file. A text file containing the frequency count of domains captured within your web archive.
</dc:description>
  <dc:identifier>https://zenodo.org/record/3633284</dc:identifier>
  <dc:identifier>10.5281/zenodo.3633284</dc:identifier>
  <dc:identifier>oai:zenodo.org:3633284</dc:identifier>
  <dc:relation>doi:10.5281/zenodo.3633283</dc:relation>
  <dc:relation>url:https://zenodo.org/communities/wahr</dc:relation>
  <dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
  <dc:rights>https://creativecommons.org/licenses/by/4.0/legalcode</dc:rights>
  <dc:subject>web archives</dc:subject>
  <dc:subject>parquet</dc:subject>
  <dc:subject>dataframes</dc:subject>
  <dc:subject>Society &amp; Culture</dc:subject>
  <dc:subject>Gay community</dc:subject>
  <dc:subject>Self-help groups</dc:subject>
  <dc:subject>Helplines</dc:subject>
  <dc:subject>Japan</dc:subject>
  <dc:title>Queer Japan Web Archive collection derivatives</dc:title>
  <dc:type>info:eu-repo/semantics/other</dc:type>
  <dc:type>dataset</dc:type>
</oai_dc:dc>
30
5
views
downloads
All versions This version
Views 3030
Downloads 55
Data volume 231.3 MB231.3 MB
Unique views 2626
Unique downloads 22

Share

Cite as