Building an AI agent is becoming easier. Deciding whether your business should own and operate one still requires a proper investment case.

For a finance leader, the decision extends beyond development time. Someone must define the process, connect the systems, test the controls, investigate failures and keep the workflow working as the business changes. Those responsibilities determine whether an impressive demonstration becomes a dependable part of the finance function.

Curia offers a managed approach: assess the opportunity, redesign and build the workflow, prove it on the client's data, then operate and improve it. An internal team can perform those same activities. The comparison is between two ways of providing the complete service, including everything that happens after launch.

The case for building internally has become stronger

AI has materially changed the economics of software development. In McKinsey's August 2026 global survey, 32% of respondents said their organisation had decided against buying at least one software product or feature because it could build the functionality internally using agentic coding tools. The survey covered 1,719 participants in 97 countries. That is evidence of changing buying decisions, rather than evidence that every internal build produces a better return. McKinsey, The state of AI in 2026

Model access has also become substantially cheaper at a given level of capability. Stanford's 2025 AI Index reported a fall of more than 280 times in the inference cost of achieving roughly GPT-3.5 performance on the MMLU benchmark between November 2022 and October 2024. This measures a particular benchmark and period; it does not measure the total cost of operating a finance workflow. Stanford HAI, AI Index 2025: State of AI in 10 Charts

For a company with capable engineers, strong finance process ownership and a well-maintained data environment, building internally deserves serious consideration. The organisation can shape the workflow around its requirements and develop expertise it can reuse.

The investment question is whether it wants to fund that capability for the life of the process.

The demonstration covers only part of the job

Consider an agent that investigates a mismatch between an invoice and a purchase order. A demonstration might show it reading the invoice, finding the order and explaining the difference.

Production introduces further questions. Which system holds the approved price? Is a later email an authorised contract variation? What happens when a receipt is missing, an approver is away or the same invoice arrives twice? If a system times out after accepting an update, how does the workflow avoid repeating the action?

These questions require decisions about process, authority and recovery. They need finance input as well as engineering work.

NIST's March 2026 report on deployed AI systems identifies ongoing monitoring as an essential discipline, including whether a system still performs its intended function and whether its infrastructure provides consistent service. Deployment creates an operating responsibility that continues after the initial project. NIST, Challenges to the Monitoring of Deployed AI Systems

That responsibility should appear explicitly in the comparison.

Responsibility Building internally Working with Curia
Process design Finance and technology teams map and redesign the work Curia leads the redesign with the finance team
Integration and development Internal engineers build and maintain the connections and workflow Curia builds the agreed workflow around the client's existing systems
Testing The business establishes and maintains its own evaluation process Curia tests on the client's data; performance and controls are agreed before launch
Operation Internal owners monitor failures, investigate exceptions and release changes Curia monitors, maintains and improves the managed workflow
Business decisions Finance retains approval authority and judgement Finance retains approval authority and judgement
Economics Internal delivery and operating costs, plus infrastructure and model usage Assessment, implementation and managed service costs, plus retained client effort and any separately charged technology costs

The Curia column describes the service model. The detailed scope and allocation of responsibilities belong in the engagement agreement.

Compare the cost of running the process over time

A credible internal estimate includes engineering, finance subject matter expertise, testing, infrastructure, model usage, monitoring, support, security work and subsequent changes. It also includes the opportunity cost of diverting people from other priorities.

Apply the same discipline to a managed service. Include implementation, recurring fees, internal supervision and work that remains with finance. Make the treatment of third-party technology costs explicit. Ask what counts as routine maintenance and what becomes a separately priced change.

Both options should be evaluated over the same period, at the same volume and against the same acceptance criteria. Comparing a supplier's annual fee with an internal team's initial development estimate hides much of the internal commitment.

One useful measure is total cost per correctly completed case, including review and rework. It connects expenditure to a business outcome and makes it harder for a low model cost to conceal an expensive operating process.

Time saved needs a route into business value

The distinction between individual productivity and company performance matters. In the same 2026 McKinsey survey, 80% of respondents reported improvements in their own productivity, while 37% attributed some positive enterprise EBIT impact to AI. These are different, self-reported measures; they show why a financial case needs more than a list of tasks made faster. McKinsey, The state of AI in 2026

Take an illustrative workflow with 6,000 cases each month. Reducing human effort by five minutes per case releases 500 hours. At an assumed loaded labour cost of £35 an hour, that represents £17,500 of monthly capacity, or £210,000 annually, before automation costs.

That capacity becomes a cash benefit only where expenditure actually changes: fewer contractor hours, lower overtime or a planned hire that is no longer required. If the team uses the time for analysis, customer queries or control improvements, the business may still gain substantially, but the benefit should be described and measured accordingly.

If only half the released capacity can be put to productive use, the modelled annual capacity value falls to £105,000 before costs. This is why Curia starts with the client's volumes, effort and economics. A benchmark can suggest where to investigate; a decision to proceed needs a local baseline.

This calculation is illustrative. It is not a Curia client result, price quotation or savings forecast.

The pilot should test the cases that make finance work difficult

A useful pilot contains representative transactions and deliberately difficult cases: missing documents, conflicting records, duplicates, unsupported requests and interrupted system connections. Historical examples used to develop the workflow should be separated from the cases used to evaluate it.

The evaluation should examine the completed business outcome. An agent can provide a convincing explanation while selecting the wrong record or leaving the intended update unfinished. Anthropic's January 2026 guidance on agent evaluation makes a related distinction between what an agent says it has done and the resulting state of the environment. Anthropic, Demystifying evals for AI agents

For finance, this means checking accuracy, policy compliance, escalation quality, evidence completeness, human handling time and successful recovery from failures. The thresholds should reflect the consequence of getting each action wrong.

Curia's assess–prove–run approach puts that evaluation before wider commitment. The first workflow earns the case for the next through observed performance.

When each approach makes sense

Building internally is attractive when the capability is strategically important, the organisation has a funded team to maintain it and there are enough related workflows to justify a continuing investment. The business should be able to name both the technical owner and the finance owner before development starts.

Curia is a stronger fit when finance has a clear operational problem but limited capacity to lead a software delivery and support function. That can be particularly relevant to a growing group managing several systems, or a team whose experienced people are already occupied with closing, reporting and acquisitions.

The case for Curia rests on the combined responsibility for process redesign, implementation and ongoing operation. The finance team defines the controls and retains the decisions; Curia takes on the work of building and running the agreed workflow.

Bring Curia one workflow, its monthly volume and the work your team still performs around it. We will assess the economics and identify what needs to be proven before you commit further. Start a conversation with Curia