Published February 20, 2026
| Version v1
Conference paper
Open
Detecting Literary Evaluations: Can Large Language Models Compete with Human Annotators?
Authors/Creators
- 1. Trier Center for Digital Humanities, Deutschland
Contributors
Data manager (6):
- 1. Universität Bielefeld
- 2. Universität Wien
- 3. Digital Humanities im deutschsprachigen Raum
- 4. Universität zu Köln
- 5. Universität Trier
Description
This study examines to which extent and in which settings Large Language Models (LLMs) can be used to annotate the complex and multi-layered phenomenon of evaluations within literary texts. It uses a gold-standard annotation of German-language fictional narratives published between 1800 and 2015 and compares human annotator agreement to the agreement of LLMs with the gold-standard annotation. The study focuses on ChatGPT, Deepseek, and Llama Sauerkraut-LM using three different prompts and the major vote method. Our results indicate that although LLMs can identify literary evaluations to some degree, their reliability still falls short compared to human annotators. LLMs' performance varies widely across texts, linguistic modernity not being the decisive factor. Clause-level evaluations were more reliably detected by LLMs than noun phrase-level evaluations. The study advances our knowledge of the potential and limitations of LLMs for very complex tasks in the literary domain.
Files
ILYAS_Salmoon_Detecting_Literary_Evaluations__Can_Large_Lang.pdf
Files
(602.9 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:70d2cecb4468d1dc18732dbba2effc0e
|
573.6 kB | Preview Download |
|
md5:7e7bfee29b277e8be24b94b95d213ce7
|
29.3 kB | Preview Download |
Additional details
Related works
- Is part of
- Book: 10.5281/zenodo.18591948 (DOI)