Published July 24, 2019 | Version v1

Documents to Data: the evolution of approaches to a library archive

  • 1. Princeton University

Description

In Digital Humanities we speak of moving from “documents to data.” In many projects, this is literal, a process of extracting information or turning text into tokens suitable for computational analysis. For the Shakespeare and Company Project, it entailed a conceptual shift from thinking of archival materials as texts to be encoded and described, to thinking of them as data to be managed in a relational database.

This project is based on the Sylvia Beach papers, held at Princeton University, which document the privately owned lending library in Paris frequented by notable writers of the Lost Generation. Materials include logbooks with membership information and lending cards for a subset of members with addresses and borrowing histories.

This poster will present the history of a multi-year project in three phases, each with benefits, difficulties, and stakes. The evolution of the project demonstrates the development of our thinking as a team as we moved toward a public-facing site designed for a broad audience. In the first phase, we encoded content from the library using TEI/XML, an approach commonly employed for documentary editing. The choice of TEI/XML fit the initial aims of the project, but even rich transcription did not offer the opportunity to fully connect the people, places, and books referenced. Consequently, the second phase was dedicated to designing a custom relational database to model the world of the library by explicitly surfacing different types of connections. The third phase required migrating data from the TEI/XML to the relational database, a lengthy process that exposed inconsistencies in the encoding, but also gave us an opportunity to eliminate redundant and unsynchronized information. The conversion process highlighted the benefits and the difficulties of both systems in pursuing similar research questions. A TEI corpus and a relational database both support querying and making connections, but a database is designed for explicit connections, which makes it easier to identify and group member activities with individual people across multiple data sources. Both approaches require technical expertise, resulting in barriers to non-technical team members working with the data.  We found the relational database to be more inclusive for project members: we built and progressively refined a web-based interface that was easier to use than oXygen XML editor, and provided on-the-fly data exports in familiar formats such as CSV. Team members could then do their own analysis (without learning query languages such as SQL or XQuery) and, as a result, had more meaningful engagements with the data. 

To illustrate the history of the project, this poster will include sample images of the archival material. It will provide a diagram that maps the transition of the data from its location across multiple XML documents to a relational database. The poster will present examples of the data work enabled by and insights gained since conversion to a relational database. Finally, it will include visuals from the public-facing web application now in development which will eventually provide researchers and the public access to the world of this library.

 

Files

Koeser-Munson_ACH2019_documents-to-data.pdf

Files (7.5 MB)

Name Size Download all
md5:de78cdbbeaf6f8d45f73bdad5dbce4c2
7.5 MB Preview Download

Additional details