Published October 22, 2022 | Version v1

Reading HaunthiTrust, or Spooky Season fun in the HTDL

Authors/Creators

  • 1. University of Notre Dame

Description

Just for fun, let's read a HathiTrust collection called HaunthiTrust.

Notes

This one was tricky. First, I used the HathiTrust Data API to download and cache both the PDF and plain text versions of the given HathiTrust collection file. Then, because the downloaded PDF files had zero OCR, I used the plain text as input for the Reader's build process. Once the carrel was build, I replaced the plain text in the cache with their corresponding PDF files. This makes for a big study carrel, but the PDF files intended to read complete with their pictures.

Methods

All Distant Reader data sets ("study carrels") use the same method of creation. First, a set of narrative files of just about any type and any number are saved in a folder/directory. Second, the plain text is pulled from each file and saved. Third, feature extraction is done against the plain text to create tab-delimited indexes of bibliographics, email addresses, URLs, parts-of-speech, named-entities, and computed keywords. Fourth, all of the indexes are reduced to an SQLite database file. Finally, everything (the original files, the plain text files, the indexes, and SQLite database) is compressed into a zip file for distribution. The result is a platform- and network-independent data set that can be read and processed by any number of GUI applications, programming languages, or a Python module called the Distant Reader Toolbox.

Files

index.zip

Files (524.8 MB)

Name Size
md5:07ebfd599e7b9072d4f07379bb1a85b0
524.8 MB Preview Download

Additional details

Software

Repository URL
https://github.com/ericleasemorgan/reader-toolbox
Programming language
Python
Development Status
Active