The Model Made the Estimate. Management Still Signs the 302.
AI did not arrive in the financial close through a project with a control-design phase. It arrived because accounting teams started using it, inside processes that were already in scope, and nobody wrote down what it had become.
There are only three possible answers to the question of what a tool is doing in a financial reporting process, and every SOX conversation about AI eventually reduces to picking one:
It is a control. It produces information relied on in a control. Or it is outside the boundary of financial reporting entirely.
The third answer is the only one that requires no further work, which is why it is the one most often assumed and least often defensible. A tool nobody scoped is not out of scope. It is unassessed.
Six Places It Is Already Operating
The table below is not a forecast. Each row is a use that has already spread through controllership teams on the strength of being obviously helpful — which it is. The column that matters is the middle one, because classification drives everything downstream: the evidence you retain, the testing your auditor performs, and whether a problem found later is a finding or a restatement conversation.
Usually a control, and often a key one. If unmatched items are cleared on the strength of a generated rationale, the tool is performing the detection, not assisting it.
Match rule logic, the population it ran against, the exception queue with dispositions, and the threshold below which nothing was investigated.
Information produced by the entity, feeding a management review control. The commentary is not the control; the reviewer's testing of it is.
The data the narrative was generated from, and evidence the reviewer independently corroborated explanations rather than accepting them.
An input to a significant estimate, which is inherently higher risk and attracts more testing than routine processes.
Method, data completeness, assumption support, back-testing against actuals, and documentation of management's own judgment on top of the suggestion.
A control over a well-known fraud risk. Scoping something down out of review is a control decision even when it is framed as prioritisation.
Coverage of the scoring population, the score threshold and who set it, and what happens to entries that fall below it.
Information produced by the entity feeding revenue recognition — the highest-scrutiny cycle in most SOX programmes.
Extraction accuracy testing on a sample, treatment of non-standard clauses, and the escalation path when a contract does not fit the template.
Disclosure controls and procedures, a separate assertion from ICFR and one the certifying officers sign to directly.
Source-to-disclosure tie-out, review of generated figures against the ledger, and control over which prior-period language was carried forward.
The Review Control Problem, Made Worse
Management review controls have been the most commonly criticised control type in inspection findings for a long time, and the criticism is always the same shape: the review happened, but nothing in the file shows what would have caused the reviewer to reject the item. No threshold, no list of what was investigated, no record of a question asked and answered. The control is asserted rather than evidenced.
Generated output makes this materially harder, for a reason that has nothing to do with accuracy rates. A well-formed narrative that ties to the numbers on its face invites confirmation. A junior analyst's rough draft has seams, and seams prompt questions. When the seams disappear, the reviewer's natural scepticism has less to catch on, and the only remaining defence is a review procedure specific enough to be performed the same way by someone who is not suspicious that day.
- A stated threshold, and a stated basis for it in relation to materiality.
- Identification of the items that met the threshold — including the ones that were investigated and cleared, not only the ones adjusted.
- Corroboration to a source independent of the tool that produced the analysis.
- Evidence of at least one rejection or challenge over the period. A control that has never once said no is difficult to distinguish from one that cannot.
Where a Shortfall Lands on the Severity Ladder
Severity is not determined by whether an error actually occurred. It is determined by what could have occurred and how likely that was — which is why a clean period does not settle the question. Two attributes of model-assisted controls tend to push an issue up this ladder rather than down: they operate at a scope no individual reviewer does, and their failures are plausible rather than obvious.
The control's design or operation does not allow management or employees to prevent or detect misstatements on a timely basis in the normal course.
In practice: The reviewer of a generated flux narrative leaves no evidence of the threshold applied, but variances are separately analysed in the FP&A cycle at a lower threshold.
Less severe than a material weakness, yet important enough to merit attention by those responsible for oversight of financial reporting.
In practice: Prompt and configuration changes to a reconciliation matcher are made without change-management approval, though every exception it produces is still worked manually.
A deficiency, or combination of them, such that there is a reasonable possibility that a material misstatement will not be prevented or detected on a timely basis.
In practice: Reserve estimates across all entities rest on a model whose inputs cannot be shown complete, with a review control that has never rejected a suggested figure.
The Vendor Upgrade Nobody Told You About
Operating effectiveness is concluded over a period, not at a point in time. That assertion assumes the control operated consistently throughout. A hosted model that was silently upgraded in the second month of the quarter breaks the assumption, and you will usually learn about it from a changelog rather than from a notice.
Treat prompts, configuration and model version as program-change objects: versioned, approved, and tied to a date so the population tested can be split if a change lands mid-period. Where the vendor offers version pinning, use it and document that you have. Where it does not, the contract is the only place the notice obligation can live — and that clause is far easier to obtain during procurement than after an auditor asks which version produced the March reconciliations.
Disclosure Controls Are a Separate Signature
Drafting assistance for MD&A and footnotes sits under disclosure controls and procedures, which the certifying officers attest to separately from internal control over financial reporting. Two failure modes recur. Figures that appear in narrative text but were never tied out to the ledger, because the tie-out control was designed around tables. And prior-period language carried forward and lightly edited, where the edit preserved a statement that has quietly stopped being true. Both are cheap to control for and expensive to find in a comment letter.
Related Reading
- AI washing and the claims in your marketing — the same statements, read by a different regulator.
- Model validation for AML transaction monitoring — a mature validation regime worth borrowing from.
- AI logs, e-discovery and litigation hold — retention decisions made for one reason, tested for another.
Your Investor Pages Make Claims Too
Automation and AI statements on product pages, trust centres and investor sections are read alongside your filings — and they age badly when the underlying process changes and the page does not.
See every claim your site is currently making. Run a free scan and check each one against what your close process actually does.
This article is general information and not legal, accounting or audit advice. Control classification, severity assessment and materiality are judgments that depend on your specific facts, your auditor's methodology and the applicable professional standards. Consult qualified advisors before relying on any conclusion here.