Dataset Open Access

National Statistical Offices and Central Banks Web Archive collection derivatives

Ruest, Nick; Adams, James; Barnhart, Marcella; Bordelon, Bobray; Crowley, Gwyneth; Donatiello, Joann; Abrams, Samantha


MARC21 XML Export

<?xml version='1.0' encoding='UTF-8'?>
<record xmlns="http://www.loc.gov/MARC21/slim">
  <leader>00000nmm##2200000uu#4500</leader>
  <datafield tag="653" ind1=" " ind2=" ">
    <subfield code="a">web archives</subfield>
  </datafield>
  <datafield tag="653" ind1=" " ind2=" ">
    <subfield code="a">parquet</subfield>
  </datafield>
  <datafield tag="653" ind1=" " ind2=" ">
    <subfield code="a">dataframes</subfield>
  </datafield>
  <datafield tag="653" ind1=" " ind2=" ">
    <subfield code="a">Government</subfield>
  </datafield>
  <datafield tag="653" ind1=" " ind2=" ">
    <subfield code="a">Statistics</subfield>
  </datafield>
  <datafield tag="653" ind1=" " ind2=" ">
    <subfield code="a">Banks and banking</subfield>
  </datafield>
  <controlfield tag="005">20200202072048.0</controlfield>
  <controlfield tag="001">3633683</controlfield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="u">Ivy Plus Libraries Confederation</subfield>
    <subfield code="a">Adams, James</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="u">Ivy Plus Libraries Confederation</subfield>
    <subfield code="a">Barnhart, Marcella</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="u">Ivy Plus Libraries Confederation</subfield>
    <subfield code="a">Bordelon, Bobray</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="u">Ivy Plus Libraries Confederation</subfield>
    <subfield code="a">Crowley, Gwyneth</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="u">Ivy Plus Libraries Confederation</subfield>
    <subfield code="a">Donatiello, Joann</subfield>
  </datafield>
  <datafield tag="700" ind1=" " ind2=" ">
    <subfield code="u">Ivy Plus Libraries Confederation</subfield>
    <subfield code="a">Abrams, Samantha</subfield>
  </datafield>
  <datafield tag="856" ind1="4" ind2=" ">
    <subfield code="s">3632737322</subfield>
    <subfield code="z">md5:318ae86eec428b35f72969dd23255212</subfield>
    <subfield code="u">https://zenodo.org/record/3633683/files/ivy-10637-auk.tar.gz</subfield>
  </datafield>
  <datafield tag="856" ind1="4" ind2=" ">
    <subfield code="s">4819597169</subfield>
    <subfield code="z">md5:cd49dbb809924513bd9a0b9543766c9f</subfield>
    <subfield code="u">https://zenodo.org/record/3633683/files/ivy-10637-parquet.tar.gz</subfield>
  </datafield>
  <datafield tag="542" ind1=" " ind2=" ">
    <subfield code="l">open</subfield>
  </datafield>
  <datafield tag="260" ind1=" " ind2=" ">
    <subfield code="c">2020-02-01</subfield>
  </datafield>
  <datafield tag="909" ind1="C" ind2="O">
    <subfield code="p">openaire_data</subfield>
    <subfield code="p">user-wahr</subfield>
    <subfield code="o">oai:zenodo.org:3633683</subfield>
  </datafield>
  <datafield tag="100" ind1=" " ind2=" ">
    <subfield code="u">York University</subfield>
    <subfield code="0">(orcid)0000-0003-1891-1112</subfield>
    <subfield code="a">Ruest, Nick</subfield>
  </datafield>
  <datafield tag="245" ind1=" " ind2=" ">
    <subfield code="a">National Statistical Offices and Central Banks Web Archive collection derivatives</subfield>
  </datafield>
  <datafield tag="980" ind1=" " ind2=" ">
    <subfield code="a">user-wahr</subfield>
  </datafield>
  <datafield tag="540" ind1=" " ind2=" ">
    <subfield code="u">https://creativecommons.org/licenses/by/4.0/legalcode</subfield>
    <subfield code="a">Creative Commons Attribution 4.0 International</subfield>
  </datafield>
  <datafield tag="650" ind1="1" ind2="7">
    <subfield code="a">cc-by</subfield>
    <subfield code="2">opendefinition.org</subfield>
  </datafield>
  <datafield tag="520" ind1=" " ind2=" ">
    <subfield code="a">&lt;p&gt;Web archive derivatives of the&amp;nbsp;&lt;a href="https://archive-it.org/collections/10637"&gt;National Statistical Offices and Central Banks Web Archive&lt;/a&gt; collection from the &lt;a href="https://archive-it.org/home/IvyPlus"&gt;Ivy Plus Libraries Confederation&lt;/a&gt;. The derivatives were created with the &lt;a href="https://github.com/archivesunleashed/aut/"&gt;Archives Unleashed Toolkit&lt;/a&gt; and &lt;a href="https://cloud.archivesunleashed.org/"&gt;Archives Unleashed Cloud&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;ivy-10637-parquet.tar.gz&lt;/strong&gt; derivatives&amp;nbsp;are&amp;nbsp;in&amp;nbsp;the &lt;a href="https://parquet.apache.org/"&gt;Apache&amp;nbsp;Parquet format&lt;/a&gt;,&amp;nbsp;which&amp;nbsp;is&amp;nbsp;a &lt;a href="http://en.wikipedia.org/wiki/Column-oriented_DBMS"&gt;columnar&amp;nbsp;storage&lt;/a&gt; format. These derivatives are generally small enough to work with on your local machine, and can be easily converted to Pandas DataFrames. See &lt;a href="https://github.com/archivesunleashed/notebooks/blob/master/datathon-nyc/parquet_pandas_stonewall.ipynb"&gt;this&lt;/a&gt; notebook for examples.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Domains&lt;/strong&gt;&lt;/p&gt;

&lt;pre&gt;&lt;code class="language-java"&gt;.webpages().groupBy(ExtractDomainDF($"url").alias("url")).count().sort($"count".desc)&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Produces&amp;nbsp;a&amp;nbsp;DataFrame&amp;nbsp;with&amp;nbsp;the&amp;nbsp;following&amp;nbsp;columns:&lt;/p&gt;

&lt;ul&gt;
	&lt;li&gt;domain&lt;/li&gt;
	&lt;li&gt;count&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Web&amp;nbsp;Pages&lt;/strong&gt;&lt;/p&gt;

&lt;pre&gt;&lt;code class="language-java"&gt;.webpages().select($"crawl_date", $"url", $"mime_type_web_server", $"mime_type_tika", RemoveHTMLDF(RemoveHTTPHeaderDF(($"content"))).alias("content"))&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Produces&amp;nbsp;a&amp;nbsp;DataFrame&amp;nbsp;with&amp;nbsp;the&amp;nbsp;following&amp;nbsp;columns:&lt;/p&gt;

&lt;ul&gt;
	&lt;li&gt;crawl_date&lt;/li&gt;
	&lt;li&gt;url&lt;/li&gt;
	&lt;li&gt;mime_type_web_server&lt;/li&gt;
	&lt;li&gt;mime_type_tika&lt;/li&gt;
	&lt;li&gt;content&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Web&amp;nbsp;Graph&lt;/strong&gt;&lt;/p&gt;

&lt;pre&gt;&lt;code class="language-java"&gt;.webgraph()&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Produces&amp;nbsp;a&amp;nbsp;DataFrame&amp;nbsp;with&amp;nbsp;the&amp;nbsp;following&amp;nbsp;columns:&lt;/p&gt;

&lt;ul&gt;
	&lt;li&gt;crawl_date&lt;/li&gt;
	&lt;li&gt;src&lt;/li&gt;
	&lt;li&gt;dest&lt;/li&gt;
	&lt;li&gt;anchor&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Image&amp;nbsp;Links&lt;/strong&gt;&lt;/p&gt;

&lt;pre&gt;&lt;code class="language-java"&gt;.imageLinks()&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Produces&amp;nbsp;a&amp;nbsp;DataFrame&amp;nbsp;with&amp;nbsp;the&amp;nbsp;following&amp;nbsp;columns:&lt;/p&gt;

&lt;ul&gt;
	&lt;li&gt;src&lt;/li&gt;
	&lt;li&gt;image_url&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://github.com/archivesunleashed/aut-docs/blob/master/current/binary-analysis.md#binary-analysis"&gt;&lt;strong&gt;Binary&amp;nbsp;Analysis&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
	&lt;li&gt;Audio&lt;/li&gt;
	&lt;li&gt;Images&lt;/li&gt;
	&lt;li&gt;PDFs&lt;/li&gt;
	&lt;li&gt;Presentation&amp;nbsp;program&amp;nbsp;files&lt;/li&gt;
	&lt;li&gt;Spreadsheets&lt;/li&gt;
	&lt;li&gt;Text&amp;nbsp;files&lt;/li&gt;
	&lt;li&gt;Word&amp;nbsp;processor&amp;nbsp;files&lt;br&gt;
	&amp;nbsp;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;strong&gt;ivy-10637-auk.tar.gz &lt;/strong&gt;derivatives&lt;strong&gt; &lt;/strong&gt;are the &lt;a href="https://cloud.archivesunleashed.org/derivatives"&gt;standard set of web archive derivatives&lt;/a&gt; produced by the Archives Unleashed Cloud.&lt;/p&gt;

&lt;ul&gt;
	&lt;li&gt;&lt;strong&gt;Gephi &lt;/strong&gt;file, which can be loaded into &lt;a href="https://gephi.org/"&gt;Gephi&lt;/a&gt;. It will have basic characteristics already computed and a basic layout.&lt;/li&gt;
	&lt;li&gt;&lt;strong&gt;Raw Network&lt;/strong&gt; file, which can also be loaded into &lt;a href="https://gephi.org/"&gt;Gephi&lt;/a&gt;. You will have to use that network program to lay it out yourself.&lt;/li&gt;
	&lt;li&gt;&lt;strong&gt;Full text&lt;/strong&gt; file. In it, each website within the web archive collection will have its full text presented on one line, along with information around when it was crawled, the name of the domain, and the full URL of the content.&lt;/li&gt;
	&lt;li&gt;&lt;strong&gt;Domains count&lt;/strong&gt; file. A text file containing the frequency count of domains captured within your web archive.&lt;/li&gt;
&lt;/ul&gt;</subfield>
  </datafield>
  <datafield tag="773" ind1=" " ind2=" ">
    <subfield code="n">doi</subfield>
    <subfield code="i">isVersionOf</subfield>
    <subfield code="a">10.5281/zenodo.3633682</subfield>
  </datafield>
  <datafield tag="024" ind1=" " ind2=" ">
    <subfield code="a">10.5281/zenodo.3633683</subfield>
    <subfield code="2">doi</subfield>
  </datafield>
  <datafield tag="980" ind1=" " ind2=" ">
    <subfield code="a">dataset</subfield>
  </datafield>
</record>
34
2
views
downloads
All versions This version
Views 3434
Downloads 22
Data volume 8.5 GB8.5 GB
Unique views 3131
Unique downloads 22

Share

Cite as