Published February 20, 2026 | Version v1

Detecting Literary Evaluations: Can Large Language Models Compete with Human Annotators?

  • 1. Trier Center for Digital Humanities, Deutschland

Description

This study examines to which extent and in which settings Large Language Models (LLMs) can be used to annotate the complex and multi-layered phenomenon of evaluations within literary texts. It uses a gold-standard annotation of German-language fictional narratives published between 1800 and 2015 and compares human annotator agreement to the agreement of LLMs with the gold-standard annotation. The study focuses on ChatGPT, Deepseek, and Llama Sauerkraut-LM using three different prompts and the major vote method. Our results indicate that although LLMs can identify literary evaluations to some degree, their reliability still falls short compared to human annotators. LLMs' performance varies widely across texts, linguistic modernity not being the decisive factor. Clause-level evaluations were more reliably detected by LLMs than noun phrase-level evaluations. The study advances our knowledge of the potential and limitations of LLMs for very complex tasks in the literary domain.

Files

ILYAS_Salmoon_Detecting_Literary_Evaluations__Can_Large_Lang.pdf

Additional details

Related works

Is part of
Book: 10.5281/zenodo.18591948 (DOI)