The cheapest part of an AI finance workflow may be the model call. The expensive parts can be establishing the right data, connecting the systems, designing the controls and dealing with the work that remains outside the automated path.
That creates a budgeting problem. A quoted usage price is easy to compare. The cost of delivering a correct, controlled outcome across several systems takes more work to establish. Finance should evaluate the whole process using a consistent boundary and a realistic operating period.
This article sets out a model for that assessment. The figures are illustrative, not Curia prices or a market-rate survey.
Falling model prices do not price the workflow
Stanford's 2025 AI Index reported that the price of querying a model at a GPT-3.5-equivalent MMLU benchmark level fell from $20 per million tokens in November 2022 to $0.07 in October 2024—more than a 280-fold reduction. This is a historical comparison at a specified benchmark level, not the current price of every model or a forecast for finance automation. Stanford HAI, AI Index research and development
It illustrates an important budgeting distinction: a change in one technical input price does not remove the cost of delivering the surrounding service.
Architecture also affects usage. Anthropic's engineering guidance notes that additional agentic complexity can trade higher cost and latency for improved performance. The appropriate design depends on whether that extra work improves the outcome enough to justify it. Anthropic, building effective agents
For example, an invoice workflow might use a rule to calculate a tolerance, a model to interpret a supplier explanation and a person to approve a commercial exception. Cost each part of that design instead of assuming every transaction needs the same amount of AI.
Define a complete cost boundary
Separate one-off investment from recurring operation. Then identify costs already included in a provider's fee to avoid double counting.
| Cost area | What to include | Question for the proposal |
|---|---|---|
| Assessment and design | Baseline measurement, process mapping, controls and scope | What decision will this stage support? |
| Implementation | Integration, data preparation, workflow construction and testing | Which systems and exception types are included? |
| Adoption | Training, procedure updates and transition effort | What time must the client team contribute? |
| Technical operation | Software, hosting, model calls, monitoring and support | What scales with volume or complexity? |
| Retained human work | Review, approvals, unresolved cases and corrections | How was this effort estimated? |
| Maintenance and change | Policy updates, integration changes and evaluation | What is included, and what triggers additional fees? |
Monitoring belongs in the recurring budget. NIST's March 2026 report discusses the challenges of monitoring AI after deployment, reinforcing the need to plan for ongoing observation rather than treating launch as the end of the work. NIST, monitoring deployed AI systems
The amount of monitoring required will vary with the process. The budget should identify an owner and an activity, rather than a generic contingency that nobody expects to use.
An illustrative first-year model
Assume a process handles 8,000 cases a month at six minutes of human work per case. That is 800 hours. At an assumed loaded cost of £35 an hour, the baseline capacity value is £28,000 a month.
Assume the proposed workflow reduces human effort to two minutes per case, including routine review and exception handling. Retained effort is approximately 267 hours, valued at £9,333. Gross capacity released is approximately 533 hours, or £18,667 a month.
Now assume £40,000 of one-off implementation expenditure and £8,000 a month of incremental operating costs covering the proposed service and technology. The economic benefit after recurring costs is approximately £10,667 a month. Across 12 months of full operation, that is £128,000 before the initial investment and £88,000 after it.
Dividing £40,000 by the £10,667 monthly net benefit gives a simple economic payback of about 3.75 months after full operation begins. This excludes deployment lead time, ramp-up, discounting and tax. A cash payback calculation would need cash savings rather than the value of released capacity.
Every input in this scenario is an assumption. It is not a quote, a typical project or a promised outcome. The loaded labour figure values available capacity; it becomes a cash benefit only where spending is actually reduced or avoided.
Stress-test the human effort assumption
Retained handling time has a large effect. If it is three minutes rather than two, monthly human effort is 400 hours and gross capacity released is worth £14,000. After the assumed £8,000 operating cost, the monthly net benefit is £6,000. Simple economic payback becomes approximately 6.7 months.
At four retained minutes, the monthly net benefit falls to about £1,333 and simple payback stretches to 30 months. The model has not changed its transaction volume or technology fee. The difference is entirely in the human work left behind.
This is why a proof should measure checking, correction and exception handling time. A demonstration that measures only the automated step cannot substantiate a claim about total labour released.
The same model produces a useful break-even volume. At two retained minutes, each case releases four minutes worth approximately £2.33 at the assumed hourly rate. Covering £8,000 of recurring cost therefore requires roughly 3,429 cases a month, before recovering the initial investment. At lower volumes, a simpler process change might be the better option.
Keep capacity, cash and working capital separate
Capacity can absorb growth, reduce overtime or give experienced staff more time for difficult cases. Those are legitimate benefits, but the route to value should be named. A salary already being paid does not disappear because a task takes less time.
Working capital needs a separate calculation. Applying an existing receipt faster can improve the accuracy of the customer ledger; it does not create a second cash receipt. Resolving a genuine payment dispute sooner may accelerate collection, but that effect needs to be observed and attributed.
Avoid adding overlapping benefits. If a reduction in overtime is already counted as a cash saving, do not also value those same hours at the full loaded rate as an additional benefit.
What Curia's assessment should establish
Curia starts with an assessment using the client's numbers, then proves a defined workflow before operating it as a managed service. The commercial case should specify the baseline, retained work, recurring costs and assumptions that could change the conclusion.
A strong proposal shows where the economics fail as well as where they work. That gives finance a practical decision: proceed, narrow the scope, fix the upstream process or leave the current approach in place. The pilot-to-production guide explains how to test those assumptions.
Bring Curia your volumes, handling times and current operating costs. We will assess whether a workflow has an economic case worth proving. Assess a workflow's economics