Get monthly updates on FP&A best-practices and more.
AI Agents in Finance: What They Actually Do All Day
There is a number from Deloitte's CFO Signals survey that has been showing up in vendor decks all year: 54 percent of CFOs say integrating AI agents into their finance departments is a transformation priority. It gets quoted that often because it reads like permission.
The number that doesn't make it onto the slide comes from a different survey. RGP asked 200 U.S. CFOs, at roughly the same time, how much value they were actually getting from the AI they had already bought. Two thirds expected significant returns within two years, and 14 percent reported meaningful value today.
The distance between those two figures is the real story of AI in finance at the moment. The question is no longer whether the technology works, but whether the thing you bought is doing anything useful on an ordinary Tuesday afternoon.
One caveat is worth registering before going any further, because both surveys polled finance leaders at companies considerably larger than most of the ones reading this, which makes them useful context rather than a direct comparison: if a two-billion-dollar enterprise with a dedicated data team is sitting at 14 percent, that tells you something about what a forty-person company should expect from a tool it connects on a Thursday.
What the word "agent" is supposed to mean
The term got stretched well past usefulness over the past year, so it helps to be specific about the three different things being sold under the same label.
A chatbot answers the question you ask it. You type "what was gross margin in Q2," it gives you the number, and the exchange ends there. That's genuinely useful, but you are still the person who had to know which question to ask.
Automation executes a rule you wrote in advance. If an invoice matches this vendor, code it to that account. It's reliable, and it's completely blind to anything you didn't anticipate when you set it up.
An agent gets handed a job rather than a question or a rule. It has a task, such as explaining this month's variances, along with access to your data and the ability to take several steps on its own to finish the work: pulling actuals, comparing them to plan, isolating what moved, checking whether the movement is a timing issue or a real change, and drafting the write-up. Your role shifts from assembling the analysis to reviewing it.
That distinction matters when you're evaluating software, because a fair amount of what gets marketed as an agent turns out to be a chatbot with better onboarding.
The four jobs agents handle well today
Deloitte's most recent read, from June of this year, has 44 percent of finance functions using AI for planning and budgeting and 41 percent using it to analyze financial data, so none of this is hypothetical anymore. The honest list of what works is still shorter than the marketing suggests.
Variance explanation is the clearest win of the group. Comparing actuals to plan, isolating what drove the difference, and drafting the commentary is structured, repetitive work, and it happens to be the task most finance teams spend the most hours on for the least strategic credit.
Anomaly detection in the ledger is a close second. Catching the vendor that got reclassified, the invoice coded twice, or the expense category that tripled while nobody was watching it is the sort of thing people normally find in November. An agent sitting on the ledger finds it during the week it happens.
Baseline forecast generation comes third on the list. Producing a first version of a rolling forecast from your historicals and drivers gives you a starting point to argue with, which is a different thing from a final answer and considerably more useful than a blank tab.
First-draft reporting rounds out the list. Monthly packages, board deck exhibits, and narrative sections rarely go out the way the software wrote them, but they're also rarely worse than what you would have written at eleven o'clock the night before they were due.
What those four have in common is worth noticing. They are all detection work and first drafts. The software handles the assembly and you supply the judgment.
The three jobs they shouldn't own yet
Anything that turns on intent belongs to a person. How to allocate shared costs, whether a contract gets recognized across twelve months or twenty-four, what qualifies as one-time: those aren't data problems, they're decisions with consequences, and whoever signs the statements should be the one making them.
Anything built on data you don't trust is the second category, and it deserves the section below.
Anything you can't show your work on is the third. If a number can't be traced back to the transactions that produced it, it has no business appearing in front of a board or an auditor, and any tool that can't answer the question "where did this come from" isn't ready for finance work regardless of how polished the output looks.
The real obstacle is usually your chart of accounts
The survey data here is uncomfortable and fairly consistent. In that same RGP study, 35 percent of CFOs named data trust as their single biggest barrier to getting a return on AI, only 10 percent said they fully trust their enterprise data, and 86 percent said legacy systems are limiting how ready they are to use AI at all.
What that suggests is that most disappointing AI projects in finance didn't fail at the AI step. They failed because the software got pointed at a general ledger where the same expense type lives in four different accounts, half the revenue sits in one catch-all, classes were applied inconsistently for two years running, and nobody ever cleaned it up because no single reason was urgent enough to justify the week it would take.
Connecting an agent supplies that reason, since it inherits whatever data hygiene you have and hands it back to you at speed, which is occasionally embarrassing and consistently useful.
The sequence that tends to work is unglamorous. Clean up the accounts you actually report on, get your systems pulling live data instead of exporting it, and then add the agent. Jumping straight to the third step is how organizations end up in that 86 percent.
A thirty-day test that costs you almost nothing
You don't need a transformation program to find out whether any of this holds up in your own shop. You need one recurring deliverable and an honest stopwatch.
Start by picking a task you already do every month and timing it truthfully, including the parts you finish on Sunday evening. Monthly variance commentary is usually the best candidate. Then hand that one task over for a full close cycle, using your own ledger and your own accounts rather than a demo dataset, because the demo dataset is always clean and yours is not.
Check the work line by line through the first month, tracing every number back to the transactions behind it. You aren't testing whether the output reads well, you're testing whether it's correct and whether you'd be able to tell if it weren't.
At the end, look at two things: hours returned to you, and how many corrections you had to make in the first month compared with the third. If the corrections aren't dropping, the tool isn't learning your business. If you come out of it having saved four hours a month and caught two errors you would otherwise have missed, that's a real result worth building on. And if you spent longer checking the work than you would have spent doing it, you found that out for the price of one close cycle instead of an annual contract.
Where we land on this
We built Mira, Clockwork's AI analyst, for the four jobs on that first list: variance analysis, anomaly flags, baseline forecasts, and first-draft reporting, all pulled straight from QuickBooks and Xero so that every figure traces back to the transaction behind it. We're deliberately not claiming it replaces the judgment part of the job. It replaces the assembly part, which is where the hours actually go.
Fourteen days is enough to run the test above on your own numbers and see what it turns up during a real close.
Conclusion
Where we land on this
We built Mira, Clockwork's AI analyst, for the four jobs on that first list: variance analysis, anomaly flags, baseline forecasts, and first-draft reporting, all pulled straight from QuickBooks and Xero so that every figure traces back to the transaction behind it. We're deliberately not claiming it replaces the judgment part of the job. It replaces the assembly part, which is where the hours actually go.
Fourteen days is enough to run the test above on your own numbers and see what it turns up during a real close.



.png)





.png)



.png)
.png)
.avif)



.avif)













.avif)


.webp)













.avif)


.png)
.avif)






.avif)









.avif)






