Interactive controls are loading. Phone and email links are available.

Skip to main content
A practical evaluation workbook for finance teams

Choose invoice AI using the documents your team actually receives

A convincing invoice demonstration shows that one document can be read. A purchasing decision needs different evidence: whether your invoices arrive intact, whether each amount belongs to the right job, and whether uncertain results stop before they create work in another system. This guide gives you a repeatable evaluation method built around those decisions.

Use it before buying an extraction tool, replacing an existing model or expanding an automation into a new supplier group. The examples are hypothetical test cases, not client results. Keep the first comparison away from live posting and payment permissions.

Invoice evaluation workflow showing approved samples, independent answers, candidate testing and a rollout decision
Invoice evaluation workflow showing approved samples, independent answers, candidate testing and a rollout decision. Select the diagram to view it full size.

Four decisions your trial should answer

Whole invoice
Score the complete business outcome
A correct total does not excuse an incorrect supplier, missing credit or wrong allocation.
Held-out set
Reserve unseen documents for the decision
Documents used to improve the setup cannot provide independent evidence that the improvement generalises.
Review effort
Measure the work left behind
Count checking, corrections and exception handling alongside extraction time and software charges.
Stop conditions
Agree them before testing
Define which errors prevent rollout even when the overall score looks favourable.

Start with the decision, then design the test

The useful comparison is between working processes on the same evidence. A model name or broad benchmark score does not describe your invoice path.

Define a usable invoice

Write down what the next system needs before comparing tools. Include supplier identity, invoice reference, currency, tax fields, line descriptions, quantities, job references and any approval requirements your team actually uses. Distinguish information printed on the document from information supplied by purchasing records. Otherwise a system that guesses a missing job number can look more complete than one that correctly asks for review.

Preserve difficult documents

Keep rotated scans, poor photographs, revised invoices, statements and attachments with several invoices in the test inventory. Record why each case is difficult. Do not let the person building the workflow quietly replace an awkward file with a clearer version. If staff really request a replacement in that situation, requesting a replacement is a valid expected outcome that should be scored explicitly.

Build a reference independently

Have a person familiar with the process create the expected record directly from the document and approved supporting records. Preserve uncertainty rather than forcing a single answer. A second reviewer should settle material disagreements before seeing candidate outputs. If the reference is copied from the incumbent system, its mistakes become the standard and a challenger can be penalised for finding them.

Separate recognition from authorisation

Reading an invoice accurately does not establish that the goods arrived, that the price was agreed or that the supplier should be paid. Score extraction and the downstream decision separately. This distinction also helps locate the fix: a parser problem, an unclear business rule and a missing purchase record require different work. The evaluation should leave you with those causes, not just a league table.

A six-part invoice evaluation protocol

Keep the source files, expected outcomes and run records together under controlled access. This makes failures reviewable and a later comparison reproducible.

A documented sampling frame

Inventory the intake

Take a representative slice of the channels you intend to automate. Note whether each file came from an email attachment, portal export, scan or photograph. Record supplier group, document type, page count and the reason it entered the sample. Keep repeated deliveries linked so that resends do not masquerade as independent evidence. Describe any groups missing from the sample before drawing a conclusion.

One shared scoring contract

Write the answer contract

Define field names, allowed empty values and how multi-page or multi-invoice files should be represented. Decide whether document dates stay as written or use one agreed format. Require supporting page references for fields a reviewer may challenge. Treat a missing mandatory field as an explicit review reason instead of silently substituting a convenient default. Version the contract before running candidates.

An honest boundary around tuning

Create development and decision sets

Use the development set to clarify prompts, mappings and exception reasons. Put the decision set aside and restrict access to its answers. After a candidate has been tuned against a failed decision case, that case becomes regression material rather than a clean unseen test. Document the split and the reason for it so the trial cannot become an exercise in repeatedly improving the same displayed score.

Comparable run records

Run the same complete workflow

Give each candidate the same source files and permitted context. Include document conversion, text extraction, validation and retry behaviour that would be used in service. Capture unsuccessful requests and timeouts rather than dropping them from the denominator. If one route uses a photograph and another uses text, label that as a workflow comparison. It is not a clean model-only experiment.

Field and invoice-level results

Score money-relevant mistakes first

Use separate outcomes for correct extraction, correct abstention, incorrect acceptance and unnecessary review. Examine supplier, reference and allocation errors individually, because they can be hidden by an average over many easy fields. Add an invoice-level result that fails when any agreed critical requirement fails. Keep reviewer notes explaining the failure so a future change can be checked against the actual cause.

A decision with explicit limits

Compare operational cost and choose

Combine measured service usage with the time people spend checking results and clearing exceptions. Show observed volumes rather than presenting a small sample as a monthly forecast. If you extrapolate, display the assumed volume and price basis separately. Recommend keeping the current method when the proposed improvement is too small, the sample is incomplete or a serious acceptance error remains unresolved.

Include these cases in the test inventory

TaskTraditionalExpected evaluation treatmentNotes
Readable digital invoiceA tidy demonstration documentTest all required fields and downstream validationKeep ordinary documents in the sample so the evaluation is not only a collection of disasters.
Photograph with a cropped edgeStaff infer the missing contentRequire review when a critical value is absentDo not reward an invented total simply because it matches a historical record by chance.
Invoice and statement togetherBoth documents look financially relevantRecognise the document roles separatelyThe statement should not become a second payable because it repeats invoice amounts.
Supplier sends a corrected versionThe latest email is assumed authoritativeIdentify the relationship and request the approved handlingYour business decides whether the existing record is replaced, credited or held.
Multiple jobs on one invoiceThe header job is copied to every lineCheck each allocation against the agreed referenceA correct grand total can conceal the most expensive operational mistake.
Credit note attached to invoiceSigns are normalised casuallyPreserve document type and the intended directionInclude the corresponding original invoice only when the production workflow would have it.
Unknown supplier layoutManual reviewer recognises the businessTest identity matching and uncertainty explicitlyDo not rely on the supplier name alone when the master record contains similar names.
Attachment cannot be readA failed request disappears from reportingCount a controlled failure and create a review taskThe trial must distinguish service failure from a valid decision to abstain.

Ways a promising score becomes misleading

Easy fields dominate the average

An invoice can contain many correctly read descriptions and one wrong allocation. Report critical-field errors and whole-invoice outcomes separately. Explain the scoring weights before presenting totals, and retain raw counts so the result can be reconstructed without trusting a chart.

The expected answer leaks into the test

Hints in filenames, reviewer comments or prompts can disclose the answer. Keep candidate inputs separate from the scoring records. Inspect a saved request from the run to verify what the system actually received, including metadata added by conversion tools.

Retries receive unlimited help

If a person edits a prompt after every failed document, you are measuring a supervised exercise. Define the retry policy in advance and include its cost and delay. Report assisted recoveries separately from first-pass results rather than presenting both as unattended success.

Historical truth is assumed

A previously posted invoice may have been corrected later or allocated incorrectly. Confirm reference values from the relevant evidence. Where the business rule changed, state which version applies to the test rather than marking either outcome correct retrospectively.

Privacy is ignored for convenience

Use only documents your organisation has approved for the selected evaluation environment. Remove unnecessary personal information while preserving the layout features being tested. Confirm access, retention and deletion arrangements as part of the evaluation scope, without assuming every vendor handles them alike.

A small win becomes a broad rollout

Success on one supplier group does not establish suitability for every document stream. State exclusions and use them to limit any pilot. A decision to keep testing is a valid result when the evidence cannot yet support a wider change.

Worked example: an allocation error hidden by a correct total

Consider a hypothetical maintenance business testing two invoice-reading routes. An invoice has separate lines for repairs at two properties. The supplier total is clear and both routes read it correctly. One route assigns every line to the first property reference in the header. The other preserves the separate references but flags a handwritten amendment for review. A simple total-only score marks both as successful. A field-average score may still favour the first route because most descriptions and amounts are correct.

The reference record should show each line, its printed reference, the approved supporting record and the expected handling of the amendment. Before running candidates, the business decides that a wrong accepted property allocation is a critical error. It also decides that review is appropriate when the amendment cannot be resolved from the permitted evidence. Those decisions turn the comparison into something meaningful: the first route has an incorrect acceptance; the second has a correct referral.

Now add the operational observations. Measure how long a reviewer needs to find and correct the first result when the source is displayed beside it. Separately measure the time needed to resolve the second result using the identified amendment. Do not count a fast extraction followed by hidden checking as a time saving. Record what the reviewer actually did and whether the proposed interface made the problem visible without reopening several systems.

Finally, ask what the proposed fix proves. Adding an instruction about property references may repair this particular invoice. Move the repaired example into the regression set and test the change on other unseen allocations. Retain the original failure and the changed version of the workflow. The purchasing decision should be based on the reserved evidence and the severity of remaining errors, not on a demonstration that the known example now looks right.

A reusable review sheet for the comparison meeting

Give each document a stable identifier and record its intake channel, category and inclusion reason. Add the expected document type, the critical fields, supporting evidence references and the authorised next outcome. Put candidate names in separate columns only after the reference is settled. This arrangement helps reviewers discuss whether an answer is defensible before they become invested in a preferred product.

For each candidate, capture first-pass outcome, final outcome after the permitted retry policy, critical error category, reviewer correction and elapsed handling time. Preserve failed requests as rows. Add a notes column for ambiguous source material and process questions that neither candidate can resolve. The meeting should distinguish a tool deficiency from an unresolved business decision rather than asking the software to settle both.

End the review sheet with a decision record. State the workflow version, sample boundaries, unresolved critical errors, observed operating effort and permitted next stage. Name the owner who can approve expansion and the evidence that would change the decision. Keep the document short enough to revisit when a supplier layout changes. Its value is that the next person can understand why a boundary exists without recreating the entire evaluation.

When the trial finds a process gap rather than a reading problem

Suppose both candidates read an invoice correctly but neither can assign the work to an approved job. The supporting record is missing from the evidence available to the workflow. That is not a tie between equally poor models. It is a finding about the process boundary. Record the missing evidence and identify who can supply it. Do not reward the candidate that invents a plausible job reference or searches outside the agreed data scope.

Use the finding to refine the operating design. The appropriate next step may be a required supplier reference, a purchasing-record lookup or a review task for the job owner. Each option changes the workflow being evaluated. State that change explicitly, update the answer contract and retain the original case. Then compare the revised process with the previous one under equivalent conditions. This allows the trial to improve the business process without falsely attributing the improvement to the model.

The same distinction applies when a PDF contains an invoice and unrelated supporting pages. If the document splitter sends only the first page to every candidate, all of them may miss a later credit. Inspect the actual candidate input before changing models. A failure caused by intake preparation should be repaired at intake and retested with mixed attachments. The evaluation record should preserve both the original failure and the evidence showing why it occurred.

At the decision meeting, group findings by the change they require: source quality, preparation, extraction, business rules, interface or permissions. Assign an owner to unresolved process questions. This prevents a software purchase from becoming a promise to solve missing approvals or undocumented allocation rules. It also makes the evaluation useful when the sensible outcome is to keep the current model and repair a narrower part of the workflow.

What Yes AI can help you evaluate

A test set built around your workflow

We can scope the document inventory, the business decisions and a practical reference format with your finance team. The deliverable should show the source of each expected answer and preserve disputed cases. This provides a basis for comparison that remains useful when the software shortlist changes.

A comparison you can inspect

A proposed evaluation can retain the candidate inputs, outputs, validation results and reviewer decisions. We can separate model quality from conversion failures and process gaps. Access to a particular accounting integration or vendor environment needs confirmation before it is included in the scope.

An operational recommendation

The recommendation should explain which documents could progress, which need a person and which cannot be processed. It should also identify the residual work and the owner of each exception. A lower service bill alone does not justify switching a process that creates financial records.

Advice when no build is justified

If your document volume is modest, the current process is dependable or the available test sample is too weak, continuing with the existing method may be the best decision. We can define a smaller evidence-gathering exercise instead of proposing a full replacement without a demonstrated need.

Take one document stream from question to evidence

The work can be scoped in stages. No stage below implies permission to post invoices or initiate payments.

Agree the intended decision

Name the exact workflow being considered, the allowed documents and the downstream action. Set critical errors and stopping rules with the person accountable for the finance process.

Prepare evidence

Collect approved documents and supporting records. Build the expected outcomes independently, review disagreements and set aside an unseen decision set before candidate tuning starts.

Run and retain

Execute the agreed candidates using the same contract. Retain errors and timeouts, measure the complete path and keep versions of instructions, mappings and validation rules.

Review the disagreements

Inspect every critical error and a sample of apparent successes. Separate conversion, extraction, allocation and business-rule causes. Check whether a proposed fix addresses the underlying issue.

Choose a limited next step

Produce a keep, revise or pilot decision with exclusions. If a pilot is justified, define monitoring, review ownership and rollback before giving it any live action permission.

Turn this guide into your next steps

Use these steps to prepare your own review. Tick a step once you have recorded its evidence. Ticks are temporary and are not saved or sent to us.

Bring one example of the process you want to improve. We can help define the scope, checks and next decision. Consultation options and any fee are shown before you book.

Scope an invoice AI evaluation

FAQ

How many invoices should we test?

There is no universal sample size that proves suitability for your business. Start by listing the document groups and decisions you must cover, then collect enough examples to expose variation within each group. Report the actual count, selection method and missing categories. A small test can reject a poor option quickly, but it should not be described as proof that a rare serious failure cannot happen.

Should we choose the cheapest extraction model?

Only after comparing the complete workflow on your documents. Conversion, retries and review can change the economics, while an inexpensive wrong acceptance may create more work than a higher service charge. Keep observed request cost separate from estimated monthly cost. If both candidates meet the agreed quality requirements, compare the operational cost and maintenance burden that remain.

Can text extraction replace image-based AI?

It is a candidate to test when documents contain usable text. Include layout-sensitive cases, scanned pages and mixed attachments in the comparison. A text route may preserve words while losing the relationships between them. Your evaluation should establish which document groups it handles and what happens when the text is incomplete, rather than assuming all PDFs behave alike.

What should count as a successful invoice?

An invoice succeeds when it satisfies the answer contract and the authorised next action is appropriate. That may mean a correct record ready for review, a properly identified duplicate or an explicit request for a clearer document. Keep these outcomes distinct. The label successful should never conceal a guessed field or a validation rule that was skipped.

Can we use our existing posted records as the answers?

Use them as leads, then check them against original documents and approved supporting records. Historical records may contain mistakes or later changes that are invisible in an export. Record the source of truth for each disputed field. If there is no defensible expected answer, mark the case unscorable for that decision and explain the gap.

Does a good result justify automatic payment?

No. This guide evaluates reading and handling invoices. Payment authorisation requires its own business controls and explicit approval. A limited pilot can produce drafts or review suggestions while existing approval arrangements continue. Any broader action should be assessed separately against the organisation's purchasing, supplier and payment procedures.

What should we bring to an evaluation discussion?

Bring a description of the intake channels, the destination system, examples of costly mistakes and the fields staff currently check. Do not send sensitive attachments before an appropriate transfer method is agreed. A redacted sample and a clear description of the decision can be enough to scope the next step without sharing an entire mailbox.

Get evidence before changing your invoice workflow

Bring the decision you need to make and the mistakes you cannot accept. Yes AI can scope a document-based comparison with an inspectable result, clear exclusions and a practical next step.

All discussions held in confidence. Australian-based consultants.