Published October 2, 2026 | Version v0.2.0

trajectory-judge: measuring what LLM judges miss when an agent reaches the right answer the wrong way

Authors/Creators

  • 1. Utrecht University

Description

A controlled study of what LLM judges of tool-using agents actually detect, under conditions where the ground truth is known by construction. A synthetic support-desk environment with a scripted oracle gives known-correct runs; a six-type injector breaks one step at a known index, keeps each fault's clean parent, and records whether the environment outcome survived and whether the final reply changed. Five judges (a rule checker, an outcome-only judge, a step-rubric judge with two models, and a self-consistency ensemble) are scored on detection, localisation, typing, calibration and cost over 400 trajectories. Comparing each fault with its clean parent shows that recall can credit a judge with detection it does not have. Every raw verdict is committed, so all tables and figures rebuild offline with no model calls.

Files

mohammadi-hadi/trajectory-judge-v0.2.0.zip

Files (2.1 MB)

Name Size Download all
md5:5e8d8a3d58c095f7a9c81cd272f1db8f
2.1 MB Preview Download

Additional details

Related works

Is documented by
arXiv:2609.00038 (arXiv)
Is supplement to
https://github.com/mohammadi-hadi/trajectory-judge (URL)