Published May 6, 2026 | Version v1

AI-in-the-Loop: Scaling Brand Visibility through LLM-Augmented Annotation

  • 1. Yext, Inc.

Description

Our study reports a proof-of-concept evaluation of Large Language Models (LLMs) as additional annotators within a production data annotation workflow. We analyze 12,745 social media posts across two taxonomies: Funnel stage and Intent. We evaluate two models (GPT-5, GPT-5 Mini) under two prompting strategies, zero-shot and few-shot. We benchmark human–LLM (HL) and gold–LLM (GL) disagreement against empirical human–human (HH) disagreement baselines (50.5% overall). Our quantitative analysis goes beyond aggregate rates by giving an in-depth insight into disagreement patterns in the form of confusion matrices, revealing different systematic biases in humans and AI. Qualitative assessment further indicates that while LLMs effectively surface human omissions, their disagreements are more often “unhelpful” than those between humans. Ultimately, this work provides a methodological foundation for integrating LLMs as collaborative agents that enhance data reliability, rather than simply automating annotation at scale.

Files

2026_ACM_CHI-poster_paper_camera_ready_zenodo.pdf

Files (20.2 MB)

Name Size Download all
md5:d5826dca0ba210564a59213a4d0c79da
6.8 MB Preview Download
md5:50a3ca4581d378aa841335775053926b
13.4 MB Preview Download