Empowering Text Encoding with Large Language Models: Benefits and Challenges
Authors/Creators
Description
Abstract
This contribution will discuss how Large Language Models (LLMs) can be used to support and enhance text encoding with the standard of the Text Encoding Initiative, demonstrating an exemplary workflow – from data model creation to data extraction, analysis, and presentation – in the context of editing letters.
With the introduction of GPT-4 in November 2022 (OpenAI 2023), LLMs experienced a sudden surge in interest that quickly caught the attention of the Digital Humanities (Pollin 2024). LLMs, advanced AI systems that learn from large amounts of text, can be used for various tasks such as text analysis, classification, interpretation of (historical) data, or language translation (Chen et al. 2023), which recommends LLMs for the application within the context of the Text Encoding Initiative.
To maximize the effectiveness of the Large Language Models (LLMs) used, we utilize two techniques: Prompt Engineering and RAG (Retrieval Augmented Generation). Prompt Engineering involves formulating precise, context-specific user queries. This method benefits significantly from the Human-in-the-Loop principle, where human feedback refines interactions between humans and models (Bsharat 2023). RAG, on the other hand, introduces external data during response generation, broadening the model’s factual accuracy and knowledge base (Saravia 2022). Together, these strategies optimize input quality and expand informational reach.
As a case study, we look at the Hammer-Purgstall Letter Edition, which is currently available in a PDF version (Höflechner 2021). It includes the correspondence of the Austrian orientalist and diplomat Joseph von Hammer-Purgstall (1774-1856), with over 8,500 letters primarily in German, but also in English, Italian, and French. The correspondences are currently recorded in several Word files that can be up to 800 pages long. Unsurprisingly, this method of recording is extremely inefficient and cumbersome. Therefore, the project is currently being revised to create a sustainable digital edition that reflects the current state of data storage and presentation.
We conduct a comparative study of two generative AI models (proprietary and open source) and present an experiment where we assign specific tasks to the Large Language Models to test their performance and efficiency. These tasks are:
-
Creation of a TEI letter model, which creates the letter structure (opener, dateline, closer, paragraph, salutation, signature, and postscript) based on sample data.
-
Annotation of (ambiguous) editorial interventions, e. g. abbreviations and expansions, uncertainties, and foreign words.
-
Collection of letter metadata in the TEI Header.
-
Named Entity Recognition including linking with authority data.
-
Data control through schemas.
-
Visualization of the network of correspondents.
At the time of submission of the paper, experiments with GPT-4, one of the leading proprietary models, were conducted. However, open-source models are catching up rapidly, and their quality is continuously improving. Therefore, a comparison with the open-source model Mixtral (MistralAI 2024) will also be presented.
The results obtained through this study will be transferred to future encoding workflows and tasks. The examination is conducted with a critical perspective on issues such as bias in the data, hallucination, non-reproducibility of results, and the largely unknown training sets. (Sharma et al. 2023)
Bsharat, Sondos Mahmoud, Aidar Myrzakhan, and Zhiqiang Shen. 2023. “Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4.” ArXiv. December 26, 2023. https://doi.org/10.48550/arXiv.2312.16171.
Chen, Zhutian, Chenyang Zhang, Qianwen Wang, Jakob Troidl, Simon Warchol, Johanna Beyer, Nils Gehlenborg und Hanspeter Pfister. 2023. “Beyond Generating Code: Evaluating GPT on a Data Visualization Course.” ArXiv. June 5, 2023. https://doi.org/10.48550/arXiv.2306.02914.
Höflechner, Walter unter Mitarbeit von Alexandra Wagner, Gerit Koitz-Arko und Sylvia Kowatsch (eds.). 2021. Joseph von Hammer-Purgstall. Briefe, Erinnerungen, Materialien. http://gams.uni-graz.at/hp.
MistralAI 2024. “Mixtral-8x22B-v0.1-AWQ.” Hugging Face. https://huggingface.co/mistral-community/Mixtral-8x22B-v0.1-AWQ.
OpenAI et al. 2023. “GPT-4 Technical Report.” ArXiv. March 15, 2023. https://doi.org/10.48550/arXiv.2303.08774.
Pollin, Christopher. 2024. “Workshopreihe Angewandte Generative KI in den (digitalen) Geisteswissenschaften” (v1.1.0). Zenodo. https://doi.org/10.5281/zenodo.10647754.
Saravia, Elvis. 2022. “Prompt Engineering Guide.” GitHub. December 2022. https://github.com/dair-ai/Prompt-Engineering-Guide.
Sharma, Mrinank; Tong, Meg; Korbak, Tomasz; Duvenaud, David; Askell, Amanda; Bowman, Samuel R.; Cheng, Newton; Durmus, Esin Durmus; Hatfield-Dodds, Zac; Johnston, Scott R.; Kravec, Shauna; Maxwell, Timothy; McCandlish, Sam; Ndousse, Kamal; Rausch, Oliver; Schiefer, Nicholas; Yan, Da; Zhang, Miranda; Perez, Ethan. 2023. “Towards Understanding Sycophancy in Language Models.” ArXiv. October 20, 2023. https://doi.org/10.48550/arXiv.2310.13548.
Files
2024-10_ScholgerStrutzPollin_TEI-LLM_English.pdf
Files
(5.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:a7d06625399cbcaf11d147bd4f243365
|
2.5 MB | Preview Download |
|
md5:b3c1d9a1fd56b2e14232b38d27a8051a
|
2.7 MB | Preview Download |