The short version: Abacum works with finance teams approving AI budget requests across the business this planning season. Split every request into automation, which replaces a process that already has a cost and gets a unit metric, or experimentation, which buys information and gets a question, a spend cap and a kill date.

The FY27 asks are in and the AI lines are bigger and vaguer than last year. Sales wants an agent nobody has scoped. Engineering has a token number that changes every time you ask for it. Somewhere in the stack there is a slide with a vendor’s ROI calculator output pasted onto it.

SpendHound surveyed 172 finance and procurement leads. 46% went over budget on AI last year. 57% are not confident they are paying a fair price. Almost twice as many are raising the line again this year.

Deloitte had 91% of organisations planning to spend more in the next twelve months. KPMG asked 2,145 leaders in May and 7% could point to an established return.

Chart: 91% of organisations plan to increase AI investment in the next 12 months (Deloitte, 1,854 senior executives, Europe and Middle East, October 2025) while only 7% report an established return on AI (KPMG Global AI Pulse Q2, 2,145 C-suite and senior leaders across 20 markets, June 2026).

Investment is up, evidence has not followed. Two separate surveys of different populations and dates, shown at the same scale for comparison only. Sources: Deloitte, State of Generative AI (Oct 2025); KPMG Global AI Pulse Q2 (Jun 2026).

I do not read that gap as the tools failing. We are running two different kinds of spending through one approval process.

Automation spend and experimentation spend are not the same request

Some of what you are being asked for is automation. It replaces a process that already exists, that somebody already does, that already costs something. Invoice matching. Contract compliance checks. First-line tickets. There is a before-state, which means there is a number.

The rest is experimentation. Somebody wants to find out whether a model can do something the company has never done. No before-state. Asking for ROI there is asking for a forecast of a discovery. You will get one and it will be fiction.

Both are worth funding. The test has to be named when the money is approved, not reconstructed in April when you are working out what happened.

What to hold each one to

Automation gets a unit. Cost per invoice, days to close, touchless rate, tickets resolved without a human. Written down with today’s number and a date, before the money moves. If the unit has not moved by the second quarter of running, the tool is wrong for the job. Stop paying for it.

Most of these do not fail on the model. Only 18% of companies have AI workflows running across every department and 84% have not redesigned a single job around what the tool can do. If nobody changed how the work happens, the unit was never going to move.

Experiments get a question, a cap and a kill date. What are we trying to find out, what is the most we will spend finding out, when do we stop. Then measure the department on how many experiments reached production and how many were killed on time. If nothing got killed in twelve months, nobody was experimenting.

Getting the baseline right is the same discipline as driver-based budgeting and forecasting. If you cannot state today’s number, you are not ready to approve the spend.

None of this survives a single AI cost centre

KPMG found leaders with strong cost visibility were five times more likely to have established a return, 15% against 3%. In the same survey, 42% could not say where their AI money was going.

This is what an FP&A platform is actually for at budget time. Abacum holds the baseline and the plan in the same place, so a tag applied in September still means something in June. Tag it by team and by use case at the point of approval. Doing it afterwards never works, because by then it is one line from one vendor and nobody remembers who asked for what.

How to model AI cost in the plan

On modelling the cost itself, the most useful framing I have read comes from CJ Gustafson at Mostly Metrics: budget tokens the way you budget benefits, as a per-head cost that varies by function rather than a project line. His poll of 250 public company CFOs landed near 20% of cash compensation. a16z’s Ivan Makarov puts per-employee token budgets at 20 to 50% of R&D salaries.

What matters for your plan is that it stacks on top of headcount instead of replacing it. If your FY27 model assumes AI spend offsets people, find out whether anyone has actually committed to that offset. I wrote about why I do not expect it in AI will not shrink your finance team. Uber’s CTO went through the whole 2026 token budget by April.

This is also where scenario planning earns its keep. Model the plan with the offset and without it, and see which one you would still sign.

What tells you it worked, by mid-2027

For automation, one question: did the unit move against the baseline you wrote down. For experiments, two: what share reached production, and what share were killed on schedule.

Then put the next dollar where the graduation rate is highest. It is the only signal I have found that is hard to game in either direction.

One thing to avoid on the way. You will be tempted to bring the MIT number into the room, the one where 95% of AI pilots fail. I would not. It measured bespoke builds against a six-month payback window, it came out of a preliminary working paper by a group with an interest in the answer, and the same report had general purpose tools converting from pilot to implementation above 80%. Somebody in that meeting has read it properly.

The case nobody has solved

A lot of what is working shows up as “we can do things we could not do before.” That is real and I do not have a clean way to put it on a P&L. Usage will not stand in for it either. All consumption tells you is that somebody opened the thing.

So budget it as capability, cap it, give it a graduation path, and say exactly that to the board. Naming the unproven spend yourself is what keeps you credible on the rest of it.

If you change one thing this cycle, make every AI ask say which of the two it is before you approve it.

Get ready for budgeting season with Abacum

In this article

Automation spend and experimentation spend are not the same request
What to hold each one to
None of this survives a single AI cost centre
How to model AI cost in the plan
What tells you it worked, by mid-2027
The case nobody has solved

Frequently Asked Questions

How should a CFO evaluate AI budget requests?

Split every request into automation or experimentation before approving it. Automation replaces an existing process, so require a unit metric and a written baseline with today's number and a date. Experimentation has no before-state, so require a question, a spend cap and a kill date instead of an ROI forecast. Naming which type it is at approval is what makes the spend reviewable later. In Abacum that baseline sits alongside the plan rather than in a side spreadsheet.

What is the difference between AI automation spend and AI experimentation spend?

Automation spend replaces a process that already exists and already has a cost, such as invoice matching or first-line support tickets. Because there is a before-state, it can be measured against a unit metric. Experimentation spend funds finding out whether a model can do something the organisation has never done, so there is no baseline and no honest ROI forecast to give. Abacum keeps both the baseline and the budget and forecast in one system, which is what lets a CFO tell the two apart at approval.

How can a CFO measure improved forecast accuracy?

Abacum measures forecast accuracy as the variance between each forecast version and the actuals it was predicting, tracked by version and by owner rather than in aggregate. That matters because a company-level accuracy number hides which department is consistently wrong and in which direction. Set the baseline before any new tool goes in, then compare like-for-like periods. See budgeting and forecasting in Abacum.

How much time can CFOs save with an automated FP&A tool?

Abacum customers typically recover days per reporting cycle, but hours saved is the weakest case a CFO can take to a board. The stronger measure is decision latency: how long it takes from a variance appearing to somebody acting on it, and how many planning rounds it takes to reach an answer. Time saved only counts once it changes a decision. See the platform.

Why can most companies not show a return on AI?

KPMG's Global AI Pulse found only 7% of leaders could point to an established return, while 42% could not say where their AI money was going. Leaders with strong cost visibility were five times more likely to have established a return, 15% against 3%. The constraint is usually instrumentation rather than technology, because spend pooled into a single cost centre cannot be attributed to any use case.

How much should we budget per employee for AI?

One useful approach is to budget tokens like benefits: a per-head cost that varies by function rather than a project line item. A poll of 250 public company CFOs by CJ Gustafson landed near 20% of cash compensation, and a16z's Ivan Makarov puts per-employee token budgets at 20 to 50% of R&D salaries. It stacks on top of headcount rather than replacing it, so it belongs in the same model as your headcount plan.

Is it true that 95% of AI pilots fail?

That figure comes from an MIT working paper and is widely misread. It measured bespoke internal builds against a six-month payback window, and the same report found general-purpose tools converting from pilot to implementation at above 80%. It is not a general finding that enterprise AI fails, and quoting it that way in a budget meeting is risky. More on how finance teams are actually adopting AI is in FP&A from the Trenches.

Explore FP&A with Abacum

See how connected planning, reporting, and data workflows help finance teams make faster, more confident decisions.

Financial planning

Build connected budgets, forecasts, and scenarios that stay current as the business changes.

Financial reporting

Turn trusted planning data into clear reporting, variance analysis, and stakeholder-ready narratives.

Finance integrations

Connect ERP, CRM, HRIS, and operational data so the plan reflects what is happening now.

Stop managing your platform and start managing the business.

Get the AI-native foundation you need to keep data trustworthy, models current, and decisions aligned—all on your own terms.

Stop managing your platform and start managing the business.

Get the AI-native foundation you need to keep data trustworthy, models current, and decisions aligned—all on your own terms.

Stop managing your platform and start managing the business.

Get the AI-native foundation you need to keep data trustworthy, models current, and decisions aligned—all on your own terms.

Webinar series with Christian Wattig: FP&A intelligence, one industry at a time
Webinar series with Christian Wattig: FP&A intelligence, one industry at a time
Webinar series with Christian Wattig: FP&A intelligence, one industry at a time