A successful demonstration shows that a workflow can complete a task under the conditions presented. A production decision requires evidence that it can handle the real case mix, respect controls, recover from interruptions and create value after the remaining human work is counted.
That gap can be closed with a clear acceptance plan. Define the intended outcome, the cases to test, the failure conditions and the operating responsibilities before the pilot starts. The result should support a decision to launch, revise the scope or stop.
Curia's staged approach begins with an assessment and a contained proof on client data. The purpose of that proof is to establish whether the workflow has earned the next commitment.
Define completion in the business system
An invoice agent saying “processed” is insufficient evidence that the invoice was recorded correctly. Completion might require a specific ERP record, an approval, an attached source document and no duplicate entry.
Anthropic's agent-evaluation guidance distinguishes the agent's recorded interaction from the final state of the environment. It also recommends repeated trials because model behaviour can vary between runs. These principles translate directly into testing finance workflows. Anthropic, evaluating AI agents
Write an acceptance statement for each outcome. For example: the correct invoice is linked to the correct entity, the authorised treatment is recorded, the supporting evidence is available and any unresolved decision is routed to the named owner.
Some cases should remain unresolved automatically. Refusing to post when authority is missing can be a successful control outcome, even though a person must finish the work.
Build two test populations
The first population should represent the expected production mix. Include the actual range of suppliers, entities, formats, values and exception types. This sample helps estimate coverage, retained effort and operating cost.
The second should deliberately challenge the controls: duplicate submissions, misleading documents, unavailable systems, expired approvals and conflicting records. Its purpose is to reveal failure modes, not to estimate how frequently those cases occur in normal operation.
Keep some examples separate from development so they test generalisation rather than memory of the cases used to tune the workflow. n8n's evaluation documentation similarly describes checking AI workflows against datasets with known expected results. n8n, why evaluate AI workflows
Record how the sample was selected and which parts of the operation it excludes. A proof on one entity and one document type can justify a narrow launch; it cannot substantiate a claim about every finance workflow.
Agree the acceptance scorecard
| Dimension | Measure | Evidence that supports a launch decision |
|---|---|---|
| Outcome quality | Correct final results by case type and value | Inspection against agreed records and expected outcomes |
| Controls | Unauthorised actions, missing approvals and duplicate prevention | Deliberate challenge cases and verified system responses |
| Coverage | Share completed within the agreed scope | All test cases included, with exclusions and escalations reported |
| Human effort | Review, correction and exception time | Timed handling across the complete workflow |
| Economics | Cost per correctly completed case | Usage, service costs and retained work counted consistently |
| Recovery | Behaviour after outages and partial completion | Demonstrated safe resume, fallback and reconciliation |
| Operation | Monitoring, ownership and change process | Named responsibilities and a usable operating procedure |
Set thresholds before reviewing results. Otherwise a team can unconsciously redefine success around whatever the prototype happens to do well. The thresholds should reflect the consequence of each error type and the economics of the workflow.
Understand what zero observed failures means
Suppose a pilot completes 100 independent, representative trials and observes no failures. That is encouraging, but it does not establish a zero failure rate.
For a simple binomial model, the one-sided 95% upper bound after zero failures is approximately 3 divided by the number of trials. More precisely, it is 1 minus 0.05 raised to the power of 1 divided by that number. For 100 trials, the upper bound is about 2.95%; for 1,000, about 0.30%.
This is an illustrative statistical calculation. It assumes independent trials with a stable failure probability and a representative sample. Repeated near-identical invoices, changing configurations or systematically omitted edge cases weaken that interpretation.
The practical response is to combine representative testing with targeted control challenges and ongoing monitoring. A larger count of easy cases cannot replace a test of whether an unauthorised action is blocked.
Include the cost of successful escalation
Imagine 1,000 representative cases. The workflow completes 800 automatically and sends 200 to reviewers. If each reviewed case needs six minutes, retained handling time is 20 hours. If every automatic case also needs a one-minute spot check in the pilot, checking adds approximately 13 hours.
The total is approximately 33 hours, before correction or administration. Comparing that with the baseline gives a more useful result than reporting “80% automated”. If the future operating plan reduces checking, it should state the evidence and controls that justify the change.
These are illustrative quantities. They demonstrate the measurement boundary rather than a target automation rate. The workflow cost guide shows how retained human effort affects payback.
Stage the release around authority
Start by running the workflow without permission to make consequential changes and compare its proposals with the actual decisions. Then introduce supervised actions within a narrow, agreed scope. Expand only when the evidence supports the additional responsibility.
For each stage, identify a stop condition and a fallback. If a data feed is incomplete, a required approval service fails or a quality measure breaches its threshold, the process should have an explicit response. Finance also needs to know which cases are in flight and which can be handled manually.
NIST's work on monitoring deployed AI highlights challenges that continue after launch. A pilot therefore needs to establish how operating data, incidents and changes will be observed, not just how a one-time test will be passed. NIST, monitoring deployed AI systems
Make change part of the acceptance decision
Models, integrations, policies and transaction populations can change. Keep the accepted test cases and rerun relevant checks when a material component changes. Investigate new failure types and add them to the test set where appropriate.
The operating agreement should say who approves changes, maintains tests, investigates incidents and decides whether to pause the workflow. It should also define which maintenance is included in the service and what requires a revised scope.
Curia's proposition connects assessment, proof and managed operation. The value of that arrangement is continuity of responsibility through those stages, with the client's team retaining decisions and sign-off. A successful proof produces a measurable operating plan that finance can hold the service to.
Bring Curia one candidate workflow and the outcomes your team would need to trust. We will define the proof, controls and economics before a wider commitment. Plan a finance workflow proof