3635634
doi
10.5281/zenodo.3635634
oai:zenodo.org:3635634
user-wahr
Williams, Kristina
Ivy Plus Libraries Confederation
Wilczek, JKeely
Ivy Plus Libraries Confederation
Darrington, Jeremy
Ivy Plus Libraries Confederation
Denniston, Ryan
Ivy Plus Libraries Confederation
Abrams, Samantha
Ivy Plus Libraries Confederation
State Elections Web Archive collection derivatives
Ruest, Nick
York University
info:eu-repo/semantics/openAccess
Creative Commons Attribution 4.0 International
https://creativecommons.org/licenses/by/4.0/legalcode
web archives
parquet
dataframes
Politics & Elections
Local elections
Political candidates
Political campaigns
<p>Web archive derivatives of the <a href="https://archive-it.org/collections/10793">State Elections Web Archive</a> collection from the <a href="https://archive-it.org/home/IvyPlus">Ivy Plus Libraries Confederation</a>. The derivatives were created with the <a href="https://github.com/archivesunleashed/aut/">Archives Unleashed Toolkit</a> and <a href="https://cloud.archivesunleashed.org/">Archives Unleashed Cloud</a>.</p>
<p>The <strong>ivy-10793-parquet.tar.gz</strong> derivatives are in the <a href="https://parquet.apache.org/">Apache Parquet format</a>, which is a <a href="http://en.wikipedia.org/wiki/Column-oriented_DBMS">columnar storage</a> format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See <a href="https://github.com/archivesunleashed/notebooks/blob/master/datathon-nyc/parquet_pandas_stonewall.ipynb">this</a> notebook for examples.</p>
<p><strong>Domains</strong></p>
<pre><code class="language-java">.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)</code></pre>
<p>Produces a DataFrame with the following columns:</p>
<ul>
<li>domain</li>
<li>count</li>
</ul>
<p><strong>Web Pages</strong></p>
<pre><code class="language-java">.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))</code></pre>
<p>Produces a DataFrame with the following columns:</p>
<ul>
<li>crawl_date</li>
<li>url</li>
<li>mime_type_web_server</li>
<li>mime_type_tika</li>
<li>content</li>
</ul>
<p><strong>Web Graph</strong></p>
<pre><code class="language-java">.webgraph()</code></pre>
<p>Produces a DataFrame with the following columns:</p>
<ul>
<li>crawl_date</li>
<li>src</li>
<li>dest</li>
<li>anchor</li>
</ul>
<p><strong>Image Links</strong></p>
<pre><code class="language-java">.imageLinks()</code></pre>
<p>Produces a DataFrame with the following columns:</p>
<ul>
<li>src</li>
<li>image_url</li>
</ul>
<p><a href="https://github.com/archivesunleashed/aut-docs/blob/master/current/binary-analysis.md#binary-analysis"><strong>Binary Analysis</strong></a></p>
<ul>
<li>Audio</li>
<li>Images</li>
<li>PDFs</li>
<li>Presentation program files</li>
<li>Spreadsheets</li>
<li>Text files</li>
<li>Word processor files<br>
</li>
</ul>
<p>The <strong>ivy-10793-auk.tar.gz </strong>derivatives<strong> </strong>are the <a href="https://cloud.archivesunleashed.org/derivatives">standard set of web archive derivatives</a> produced by the Archives Unleashed Cloud.</p>
<ul>
<li><strong>Gephi </strong>file, which can be loaded into <a href="https://gephi.org/">Gephi</a>. It will have basic characteristics already computed and a basic layout.</li>
<li><strong>Raw Network</strong> file, which can also be loaded into <a href="https://gephi.org/">Gephi</a>. You will have to use that network program to lay it out yourself.</li>
<li><strong>Full text</strong> file. In it, each website within the web archive collection will have its full text presented on one line, along with information around when it was crawled, the name of the domain, and the full URL of the content.</li>
<li><strong>Domains count</strong> file. A text file containing the frequency count of domains captured within your web archive.</li>
</ul>
Zenodo
2020-02-04
info:eu-repo/semantics/other
3635633
user-wahr
1580887252.555348
2089202845
md5:e0db18b6ef254099a8edc3f844ab2545
https://zenodo.org/records/3635634/files/ivy-10793-auk.tar.gz
3224743218
md5:3b0db047a7fcf1f744cb1a65bb7c1568
https://zenodo.org/records/3635634/files/ivy-10793-parquet.tar.gz
public
10.5281/zenodo.3635633
isVersionOf
doi