Published April 2, 2026
| Version v12
Preprint
Open
PDR in Production: Empirical Evidence for Cross-Session Behavioral Reliability Scoring in Autonomous AI Agents
Description
PDR in Production v2.8: Empirical Validation of Behavioral Trust Scoring in Multi-Agent Systems. v2.8 adds §7.6.8 (The Evaluator's Blind Spot: Cross-Run Gap in Agent Evaluation Frameworks). 15-repository overnight survey (00:00–06:10 UTC April 2, 2026) confirms the cross-session measurement gap is present universally in agent evaluation frameworks — not only in deployment/observability infrastructure. Frameworks spanning Python, TypeScript, Go, and Shell; 0–55 stars; including CI/CD eval harnesses, LLM judge pipelines, agent audit trails, benchmark runners, and observability platforms. All 15 exhibit identical omission: rich per-run metric capture with no cross-run slope analysis layer. Introduces 'evaluator's blind spot' terminology for measurement systems that capture per-instance temporal data but omit trend computation over that data. Meta-level significance: the tools built to audit agent reliability share the same architectural blind spot that PDR addresses. v2.7 added §7.6.7 documenting cross-domain convergence (attestation, enterprise SLO, audit gates). Prior versions established 65+ independent repositories confirming the gap. Concept DOI: 10.5281/zenodo.19154458
Files
pdr-in-production-v2.8.pdf
Files
(245.2 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:8c469f7d2fe08f222f35c5d8b84d9db0
|
245.2 kB | Preview Download |
Additional details
Related works
- Is new version of
- Preprint: 10.5281/zenodo.19378492 (DOI)
- Is version of
- Preprint: 10.5281/zenodo.19154458 (DOI)