Reducing Errors in AI-Generated Financial Reports Through Explainable Artificial Intelligence: A Proposed Human-AI Verification Framework
Keywords:
explainable artificial intelligence; financial reporting; generative AI; accounting errors; human-AI verification; audit trail; internal controlAbstract
Artificial intelligence can draft financial reports quickly, but fluent language does not prove numerical accuracy. This study examines whether explainable artificial intelligence can reduce errors in generated financial reports. It proposes a Human-AI Verification Framework that treats each material statement as a testable financial claim. Each claim receives a source reference, recalculation, risk score, readable explanation, and recorded reviewer decision. The empirical study uses public Online Retail II transactions and official United Kingdom inflation data. Verified monthly and weekly measures formed the correct reporting base. Eight controlled error classes were inserted into those measures. The final benchmark contained 1,275 claims, including 850 errors and 425 valid claims. A time-based split produced 897 training claims and 378 test claims. Random forest, gradient boosting, logistic regression, and accounting rules were compared. Tree-based models achieved F1 scores above 97 percent on the later holdout period. The random forest detected 239 of 252 test errors without rejecting any valid claim. A 0.25 risk threshold captured 96.0 percent of errors under a production simulation. That setting routed 10.3 percent of claims for review. The residual error rate fell below one percent when reviewers corrected 95 percent of flagged errors. The findings support selective human review, evidence-linked explanations, and retained audit trails. The experiment uses controlled errors, so live organizational testing remains necessary.

