A BENCHMARK FRAMEWORK FOR EVALUATING AUTONOMOUS AI AGENTS IN ACCOUNTING
Authors/Creators
- 1. student of ''Tashkent Institute of Irrigation and Agricultural Mechanization Engineers'' National Research University
- 2. Worldly Knowledge Publishing Centre
Description
The rapid proliferation of large language model (LLM)-based autonomous agents has introduced new possibilities for automating complex, multi-step accounting tasks, ranging from journal entry generation and bank reconciliation to tax computation and financial statement drafting. Despite growing industry adoption, the accounting profession currently lacks a standardized, discipline-specific benchmark for evaluating the reliability, regulatory compliance, and auditability of these agents. Existing AI benchmarks are largely borrowed from general natural-language-processing or software-engineering domains and fail to capture the procedural rigor, traceability, and normative constraints that define professional accounting practice.
This paper proposes a Benchmark Framework for Evaluating Autonomous AI Agents in Accounting (BEAAA), a structured evaluation protocol comprising six dimensions: task accuracy, regulatory compliance, auditability and traceability, robustness to adversarial or ambiguous inputs, cost-efficiency, and explainability. The framework operationalizes each dimension through a task taxonomy spanning four accounting sub-domains — bookkeeping, reconciliation, financial reporting, and tax preparation — and a graded scoring rubric that captures partial task completion rather than binary success/failure outcomes.
We illustrate the framework through a pilot evaluation of three representative agent architectures across a synthetic task suite of 120 accounting scenarios. Results indicate substantial variance across agents in compliance and auditability scores despite comparable raw task-accuracy scores, suggesting that accuracy alone is an insufficient basis for deployment decisions in regulated accounting environments. The proposed framework offers accounting researchers, software vendors, and regulators a common vocabulary and methodology for comparing autonomous agents, and it identifies auditability and explainability as the most under-addressed dimensions in current agent design. Implications for standard-setting bodies and directions for future benchmark development are discussed.
Files
540-552.pdf
Files
(329.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:49f554987ffbe281438dff4b90cd36b2
|
329.0 kB | Preview Download |
Additional details
References
- 1.Bao, Y., Ke, B., Li, B., Yu, Y. J., & Zhang, J. (2020). Detecting accounting fraud in publicly traded U.S. firms using a machine learning approach. Journal of Accounting Research, 58(1), 199-235.
- 2.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877-1901.
- 3.Cao, S., Jiang, W., Yang, B., & Zhang, A. L. (2021). How to talk when a machine is listening: Corporate disclosure in the age of AI. Review of Financial Studies, 34(8), 3852-3897.
- 4.Deloitte. (2025). Agentic AI in finance and accounting: 2025 industry outlook. Deloitte Insights.
- 5.EY. (2024). The future of the finance function: Autonomous agents and the modern controller. EY Global.