Published July 12, 2026 | Version v1

A BENCHMARK FRAMEWORK FOR EVALUATING AUTONOMOUS AI AGENTS IN ACCOUNTING

  • 1. student of ''Tashkent Institute of Irrigation and Agricultural Mechanization Engineers'' National Research University
  • 2. Worldly Knowledge Publishing Centre

Description

The rapid proliferation of large language model (LLM)-based autonomous agents has introduced new possibilities for automating complex, multi-step accounting tasks, ranging from journal entry generation and bank reconciliation to tax computation and financial statement drafting. Despite growing industry adoption, the accounting profession currently lacks a standardized, discipline-specific benchmark for evaluating the reliability, regulatory compliance, and auditability of these agents. Existing AI benchmarks are largely borrowed from general natural-language-processing or software-engineering domains and fail to capture the procedural rigor, traceability, and normative constraints that define professional accounting practice.

This paper proposes a Benchmark Framework for Evaluating Autonomous AI Agents in Accounting (BEAAA), a structured evaluation protocol comprising six dimensions: task accuracy, regulatory compliance, auditability and traceability, robustness to adversarial or ambiguous inputs, cost-efficiency, and explainability. The framework operationalizes each dimension through a task taxonomy spanning four accounting sub-domains — bookkeeping, reconciliation, financial reporting, and tax preparation — and a graded scoring rubric that captures partial task completion rather than binary success/failure outcomes.

We illustrate the framework through a pilot evaluation of three representative agent architectures across a synthetic task suite of 120 accounting scenarios. Results indicate substantial variance across agents in compliance and auditability scores despite comparable raw task-accuracy scores, suggesting that accuracy alone is an insufficient basis for deployment decisions in regulated accounting environments. The proposed framework offers accounting researchers, software vendors, and regulators a common vocabulary and methodology for comparing autonomous agents, and it identifies auditability and explainability as the most under-addressed dimensions in current agent design. Implications for standard-setting bodies and directions for future benchmark development are discussed.

Files

540-552.pdf

Files (329.0 kB)

Name Size Download all
md5:49f554987ffbe281438dff4b90cd36b2
329.0 kB Preview Download

Additional details

References

  • 1.Bao, Y., Ke, B., Li, B., Yu, Y. J., & Zhang, J. (2020). Detecting accounting fraud in publicly traded U.S. firms using a machine learning approach. Journal of Accounting Research, 58(1), 199-235.
  • 2.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877-1901.
  • 3.Cao, S., Jiang, W., Yang, B., & Zhang, A. L. (2021). How to talk when a machine is listening: Corporate disclosure in the age of AI. Review of Financial Studies, 34(8), 3852-3897.
  • 4.Deloitte. (2025). Agentic AI in finance and accounting: 2025 industry outlook. Deloitte Insights.
  • 5.EY. (2024). The future of the finance function: Autonomous agents and the modern controller. EY Global.