Published July 27, 2026 | Version v2

Replication Package of How Developers Experience Debugging Unfamiliar Codebases with Code Tours Generated and Evaluated by Local LLMs

Description

This replication package contains the code, data, and instructions needed to replicate the experiments described in the paper "How Developers Experience Debugging Unfamiliar Codebases with Code Tours Generated and Evaluated by Local LLMs".

Context. Code tours are interactive, onboarding documentation that guide developers through a codebase. Large Language Models (LLMs) can automatically synthesize code tours. Prior work on code tour generation has not examined developer experience or trust calibration when debugging unfamiliar codebases with code tours generated and evaluated by open-weight LLMs.
Objectives. This study surveys how the properties of components in open-weight LLM-authored code tours influence developers' experiences when debugging unfamiliar codebases.
Method. We built a pipeline that generated and evaluated code tours from real reproducible bugs. Twenty-six developers with varying backgrounds participated in a user study. In total, 26 code tours were authored from real Java bugs mined from 2025 GitHub commits, with each tour independently judged by two different LLMs, resulting in 52 evaluated configurations. Participants thought aloud as they explored each tour. Three authors qualitatively coded the interviews to identify recurring themes.
Results. Developers generally preferred tours that scaled detail with the code length, avoided merely restating code, were easily scannable, and adopted a guiding tone. However, some preferences were mutually exclusive, such as the use of imperative mood. Stack traces were often insufficient to identify all steps developers found relevant. Developers also trusted descriptions they perceived as human-written more than those they believed were AI-generated. Finally, LLM-generated annotations of tour quality were unreliable: sycophancy, confabulation, and incoherence were pervasive.
Conclusion. This work lays a basis for future research on fine-tuning open-weight models for code tour generation, personalizing generation to accommodate diverging preferences, selecting relevant steps beyond stack traces, calibrating users' trust to avoid both disuse and misuse, and improving open-weight LLMs' ability to be more trustworthy evaluators

Files

balfroim/HumanFactorsCodeTour-zenodo.zip

Files (51.7 GB)

Name Size
md5:5c42be7f8a002cc7467735245c0b4ef8
16.3 MB Preview Download
md5:c450a1c505dfe956fa058457762a5a73
51.7 GB Download

Additional details

Related works

Funding

Service Public de Wallonie
ARIAC project 2010235