Published May 22, 2024
| Version v1
Poster
Open
I-Analyzer: A Digital Text Corpus Tool for All
Authors/Creators
Description
Distant reading in a corpus of text resources is a recurring need across various fields within digital humanities. It allows for quick extraction of key words and phrases from a text, seeing patterns in texts over time, or finding a relevant subset of documents for close reading. With the availability of newspaper corpora with sufficient OCR quality and high-quality metadata, Utrecht University’s Research Software Lab started development on the text corpus tool I- Analyzer in 2016. From the start, the goal of I-analyzer has been to provide researchers and students with the stepping stones that distant reading and other text mining techniques can supply, without them having to be proficient coders. Built to be compatible with different data formats (currently xml, csv, xlsx, and html) and configurable to collect and index text and most types of metadata, I-Analyzer has since been extended with many digital humanities corpora - newspapers, literature, parliamentary debates, judicial rulings, online reviews, and funerary inscriptions to name but a few - as well as more visualization and search functionalities.
Other text mining applications have also been built for this goal, such as Delpher (Der Weduwen 2015) or Voyant Tools (Sinclair / Rockwell 2012), but they either do not provide the option to work with multiple corpora, or they do not offer full-text search, or filtering and visualization through metadata. The Gale Digital Scholar tool does offer all these functionalities to some extent, but, unlike I-Analyzer, it is a commercial product that researchers (or their affiliated institutions) need to pay to use. I-Analyzer is open-source, and several of the corpora hosted on the platform are as well, which means that anyone can use and even adapt I-Analyzer and the open-source corpora hosted there.
New corpora can be added to I-Analyzer relatively quickly by defining the format of the source data. Data are indexed using Elasticsearch, which enables fast full-text search and filtering. The user is presented with an interface resembling a search engine: using a search bar and a side bar of metadata filters, they can inspect documents matching their search for close reading or download a subset of the corpus to process further offline. Currently, visualization options include document- and term frequency of the search term, as well as a representation of the most frequent n-grams including the search term. Moreover, for selected corpora, diachronic word models were trained using Word2Vec (Mikolov et al. 2013) to inspect which words co-occur most frequently with a given search term over time, comparable to capabilities of ShiCo (Martinez-Ortiz et al. 2016) or Hansard-shiny. Through interactive visualization’s, I-Analyzer attempts to give tools to researchers who might otherwise never be able to engage with digital methods such as word modeling.
I-Analyzer has been used in various research projects across different fields, ranging from the impact of translations on reader experience through book reviews (Kotze et al. 2021), the occurrence of concepts and names on Jewish funerary inscriptions (Saar 2021) to the conceptual history of words surrounding democracy in parliamentary speeches (Ihalainen et al. 2022). Currently, we are working on extending the infrastructure such that it is going to be possible for corpora to be added by researchers themselves, without the need for one of our developers to directly oversee this process. To this end, wecreated this poster.You can find I-Analyzer here, with its public datasets accessible to anyone after making a free account: https://ianalyzer.hum.uu.nl
Files
Poster I-Analyzer DHBenelux 2024.pdf
Files
(474.6 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:51c61871187761461ff9b6a7fbb9787c
|
474.6 kB | Preview Download |
Additional details
Software
- Repository URL
- https://ianalyzer.hum.uu.nl/