Interactive controls are loading. Phone and email links are available.

Skip to main content
Decision rules for document review

Set extraction thresholds around the errors your business cannot accept

A document-reading system can produce a tidy answer without enough evidence to act on it. The practical question is not whether the answer sounds confident. It is whether your workflow has a tested basis for accepting that field, referring the document or stopping the next action. Different fields and different consequences need different controls.

This guide describes a method for designing review thresholds from your own labelled examples. It does not offer a universal percentage or claim that one score establishes correctness. The worked examples are hypothetical and should be adapted to an approved test environment.

Extraction confidence workflow linking labelled examples, field checks, review decisions and outcome monitoring
Extraction confidence workflow linking labelled examples, field checks, review decisions and outcome monitoring. Select the diagram to view it full size.

Four checks before choosing a threshold

Evidence
What does the score actually describe?
A field-reading score, a rule result and a model-written confidence statement are different signals.
Consequence
What happens if acceptance is wrong?
Treat a wrong supplier or destination differently from a cosmetic description issue.
Calibration
How does the signal behave on your cases?
Compare score ranges with independently reviewed outcomes before relying on a cutoff.
Coverage
Which documents were actually tested?
Keep unrepresented layouts and document groups outside the acceptance claim.

Confidence needs a business decision behind it

Use the score as one input to a defined acceptance policy. The surrounding evidence and the action being authorised determine what that policy should require.

Name the source of each signal

An extraction provider may return a documented field score, while a language model may simply write high confidence in its response. Those are not interchangeable. During scoping, establish what each signal means, when it is absent and whether it has been tested against outcomes on your documents. Preserve the raw signal rather than renaming everything confidence and losing its origin.

Separate field quality from record suitability

A date can be read clearly while referring to the wrong event. An amount can match the printed total while the invoice belongs to another entity. Evaluate recognition, interpretation and business validation separately. A strong signal at the first stage must not silently stand in for evidence required at the later stages, particularly when the result creates or changes a record.

Measure incorrect acceptance and unnecessary review

A threshold that refers every document avoids automated acceptance errors by doing no useful acceptance work. A permissive threshold may look efficient while passing the failures you care about. Report both sides using actual counts and categories. The decision should state the acceptable operating boundary rather than hiding the trade-off inside one accuracy score.

Keep rejection and uncertainty distinct

A missing required document, a failed lookup and a low-quality scan all prevent progression for different reasons. Do not force them onto one numerical scale. Explicit rules can require review regardless of the model signal. This makes the policy easier to explain and ensures that a confident extraction cannot override an unavailable source or an unresolved approval requirement.

Build a review policy in six stages

Treat this as a proposed evaluation and control process. Any provider-specific score or integration behaviour must be confirmed for the system under consideration.

An explicit acceptance decision

Define the action boundary

Start with the action the extraction result would enable. Drafting a searchable record, creating an invoice and changing supplier details have different consequences. List the mandatory evidence for that action and the fields whose failure should stop it. Keep permissions separate from the quality score. A result that meets an extraction threshold is still subject to the business's approval and access rules.

A defensible reference set

Label a representative sample

Have qualified reviewers establish expected fields and expected handling from original sources. Include readable documents, ambiguity, missing information and cases that should be refused. Retain unscorable fields as an evidence limitation rather than inventing an answer. Divide development cases from the set used to assess the final policy. Record document groups so performance is not averaged across incompatible inputs.

Evidence about the signal

Inspect signal behaviour

Compare the available scores and rule results with actual correct and incorrect outcomes. Look for confidently wrong cases and low-scored but usable cases. Analyse critical fields separately instead of relying on a document-wide mean. If the signal does not distinguish outcomes usefully on the sample, do not force a numerical cutoff to do a job it cannot support.

An inspectable acceptance rule

Combine thresholds with checks

Use arithmetic, required-field, identity and cross-record checks where appropriate to the workflow. Keep each check's result visible and preserve failures to obtain evidence. A numerical threshold should not cancel a failed mandatory rule. Define the combined policy in plain language so reviewers understand why a document passed or was referred, and test that the implementation follows that policy.

Quality and workload together

Evaluate review workload

Run candidate policies on the reserved cases and report correct acceptance, incorrect acceptance, correct referral and unnecessary referral. Measure the review effort for referred cases rather than assuming all reviews take the same time. A useful policy should make the reason for review actionable. If every referral still requires full re-entry, the interface may need work before the threshold does.

A maintained decision policy

Monitor and revisit deliberately

Retain the policy version applied to each outcome. Watch for new document groups, changes in layouts and failures that pass existing checks. Review a sample of accepted work as well as referred work. Changes to the threshold should go through a repeatable comparison and approval process, with a way to return to the previous rule if the new boundary performs poorly.

Different fields need different evidence

TaskTraditionalProposed acceptance treatmentNotes
Supplier identityAccept the closest extracted nameRequire an approved identity match or reviewA clear reading of the printed name does not establish which master record should be used.
Invoice totalTrust a confident numberCheck the field and relevant arithmeticA total that reconciles can still belong to the wrong document or entity.
Job allocationCopy the most prominent referenceCheck each required allocation against permitted evidenceTreat missing or conflicting references as a specific business review reason.
Document dateUse the first visible dateDistinguish issue, service and due datesClarity of reading is different from choosing the correct date role.
Free-text descriptionRequire exact wordingDefine what meaning and detail must be preservedA cosmetic difference may not justify the same review rule as a money-relevant error.
Bank details on an invoiceHigh confidence permits an updateKeep supplier-change authorisation separateDocument extraction must not become permission to alter payment instructions.
Unreadable attachmentAssume low confidence means missing fieldsRecord a processing failure or request replacementNo usable source should not become a completed record with default values.
New supplier layoutApply the existing cutoff automaticallyLimit acceptance until its behaviour is evaluatedState which categories remain outside the evidence supporting the policy.

Threshold mistakes that can look mathematically tidy

A self-reported score is treated as a probability

A model writing a percentage does not by itself establish an empirically calibrated probability of correctness. Identify how the number was produced and test its relationship to outcomes. If that relationship is unknown, describe it as an unvalidated signal and keep the decision supported by other checks.

An average hides a critical field

A record containing many easy fields can receive a strong overall score despite an incorrect critical identifier. Gate critical requirements individually. Show the reason a record passed so the reviewer can see whether each required field and business check was actually considered.

The test set is repeatedly tuned

A cutoff chosen after inspecting every reserved failure is no longer independently evaluated on that set. Keep the repaired examples for regression and obtain fresh decision evidence. Record the tuning history so a rising score is not mistaken for broader reliability.

Rare failures are declared impossible

Not observing a particular error in a limited sample does not establish that it cannot occur. Describe the tested coverage and unresolved risks. Use conservative action boundaries where the potential consequence is significant and the available evidence is limited.

Missing checks default to pass

An unavailable supplier lookup or absent provider score should have an explicit handling rule. Do not let empty values become successful validation through a convenient default. Include missing-signal and failed-check fixtures in the evaluation and inspect their final business outcomes.

Review feedback is incomplete

If only rejected documents are reviewed, you cannot see confidently wrong accepted cases. Sample accepted outputs under the agreed oversight process. Preserve corrections and their causes so the policy can be assessed against both sides of its own decision boundary.

Worked example: a confident total with an uncertain destination

Consider a hypothetical invoice where the total is clearly printed and the extraction route returns a strong signal for that field. The document also names a trading business that could match two supplier records. A policy based on the average field score might accept the whole invoice because most values are easy to read. The correct design question is whether the supplier identity requirement has been satisfied, regardless of the total score.

The business defines separate acceptance conditions. The amount must be read and pass the applicable arithmetic check. The supplier must resolve to an approved record using the permitted evidence. A failure of either condition creates review. The system records that the amount passed and identity remained unresolved. The reviewer receives a supplier-match question rather than being asked to recheck every number on the page.

Now compare two candidate policies on the same reserved cases. One allows a high overall score to override identity ambiguity. The other requires the identity check independently. Record the incorrect acceptances and unnecessary referrals produced by each, including what a reviewer needs to resolve them. Do not claim a benefit from a hypothetical percentage. The decision should use the actual test results and the business consequence of each error.

If the second policy creates too many referrals, inspect the supplier evidence first. Perhaps the approved master data lacks an identifier that the invoices consistently contain. That is a data quality improvement to evaluate. Lowering the threshold would address the queue count without supplying the missing identity evidence. Keep the decision focused on the reason the document cannot safely progress.

A threshold policy worksheet you can reuse

For each required field, record the source signal, its documented meaning, the expected value source and the consequence of error. Add the explicit validation checks and the handling of absent signals. State whether the field is mandatory for the intended next action or only useful for later search. This stops a convenient score from replacing an unanswered business requirement.

For each candidate policy, record the sample boundaries and counts of correct acceptance, incorrect acceptance, correct referral and unnecessary referral. Include critical-error categories separately. Add measured review effort and examples of misleading signals. Keep the raw outcomes available so another reviewer can reconstruct the conclusion without relying on a chart or aggregate score.

The final policy record should name the approved document group, excluded cases, mandatory checks, owner and version. Include what happens when a check is unavailable and what evidence triggers reassessment. When the workflow changes, compare the new version against this record. This makes the threshold a maintained business control rather than a number somebody chose during a demonstration.

Distinguish three different reasons to stop

An unreadable source, an uncertain interpretation and an unavailable validation check can all prevent acceptance, but they should not be collapsed into one low-confidence outcome. The remedy for an unreadable source is usually better evidence. The remedy for an uncertain interpretation may be a reviewer comparing the relevant fields. The remedy for an unavailable check is to restore or investigate the checking path. A useful policy names these differences so the queue routes work correctly.

Consider a document with a clear printed reference while the approved supplier lookup is offline. The extractor may provide a strong signal and the printed text may be correct. The business identity requirement is nevertheless unverified. Record that the lookup could not be completed and prevent any action that requires it. Lowering or raising the extraction cutoff cannot supply the missing validation result. Test this exact absence case when evaluating the combined policy.

Now consider a poor scan where the supplier is known but one amount is obscured. Arithmetic may suggest a possible value, yet your contract may require the amount to be supported by the document. Decide in advance whether inference is permitted for that field and under what evidence. If it is not, the expected outcome is referral or a replacement request. A candidate should not receive a correct-acceptance label merely because its guess happens to match the reference.

Finally, consider a readable document whose business classification is ambiguous. This may need a process owner to decide what the document means, rather than a data-entry correction. Preserve the ambiguity and the relevant evidence. When reporting threshold performance, group these causes separately. Otherwise a model-quality chart can conceal missing infrastructure or unresolved business rules and encourage the team to adjust a numerical cutoff for a problem it cannot solve.

How Yes AI can help define a defensible boundary

Clarify the required evidence

We can map the fields and checks needed for one document action with the people accountable for the process. The result should distinguish extracted facts, inferred values and authorisation. This makes it possible to discuss thresholds without treating every document field as equally consequential.

Evaluate the available signals

A scoped trial can compare provider signals and validation outcomes with independently labelled examples. We can report where those signals help and where they do not. The trial should retain critical failures and limits rather than turning a small sample into a universal quality promise.

Design understandable review reasons

We can translate the acceptance policy into reasons a reviewer can act on. The proposed interface should show the disputed field and relevant source, with controlled correction actions. A threshold that sends work to an unusable queue merely moves the problem to staff.

Recommend simpler controls when appropriate

Some document tasks are better served by explicit validation and review than a numerical confidence system. If the signal is poorly defined or the volume does not justify calibration work, we can scope a simpler policy. There is no requirement to add AI scoring where it provides no demonstrated value.

From a confidence number to a tested review policy

Keep the evaluation separate from live permissions. A threshold proposal should become operational only after the business owner accepts the evidence and limits.

Name the decision

Document the next action, critical fields and mandatory checks. Identify the owner who can approve the acceptance boundary and the cases that must always be reviewed.

Prepare labelled evidence

Collect approved examples, establish expected outcomes and reserve independent decision cases. Record missing document categories and unresolved reference disagreements.

Compare candidate policies

Test the available scores alongside explicit validation rules. Report critical incorrect acceptance separately from ordinary extraction differences and unnecessary referrals.

Test the review path

Have staff handle representative referrals with the proposed evidence panel. Confirm the reasons are actionable and that completing review does not bypass unrelated approval controls.

Release within coverage

Apply only to the evaluated document group, retain policy versions and sample accepted work. Reassess on new evidence rather than changing the cutoff informally to reduce queue size.

Turn this guide into your next steps

Use these steps to prepare your own review. Tick a step once you have recorded its evidence. Ticks are temporary and are not saved or sent to us.

Bring one example of the process you want to improve. We can help define the scope, checks and next decision. Consultation options and any fee are shown before you book.

Design your document review thresholds

FAQ

What confidence threshold should we use for invoice extraction?

There is no percentage that can be recommended responsibly without knowing what the signal means, which documents it was tested on and what acceptance permits. Evaluate candidate policies on labelled examples and inspect critical failures. Use field-specific checks and explicit review rules where a single score does not capture the decision.

Is a model saying it is confident useful?

It may be a signal to investigate, but it should not be treated as a validated probability merely because it is expressed numerically. Keep its origin clear and compare it with independent outcomes. A well-written answer and a plausible explanation are not substitutes for source evidence or successful business validation.

What does calibration mean here?

In this context, calibration means examining how a score relates to observed correctness on relevant examples. The practical objective is to understand whether score ranges support a useful decision boundary. State the sample and method rather than claiming the relationship holds for every future document or for a different provider's score.

Should every field have the same threshold?

Not necessarily. Fields differ in available signals, ambiguity and consequence. A cosmetic description difference may be acceptable while an uncertain supplier identity requires review. Establish the field requirements and record-level business checks explicitly. Avoid combining unlike signals into a single number unless you can explain and evaluate the resulting policy.

Can arithmetic validation replace confidence scoring?

Arithmetic is a useful explicit check for tasks where amounts should reconcile, but it does not establish supplier identity, commercial approval or correct allocation. Use it for the claim it can support and keep other required checks separate. Whether a numerical confidence signal adds value should be tested rather than assumed.

How do we reduce too many referrals?

Inspect referral causes and the effort needed to resolve them before lowering a cutoff. Some referrals reveal missing source records, poor scans or unclear rules. Improving evidence and the review interface may help more than relaxing acceptance. Test any changed policy on reserved cases and report what new incorrect acceptances it introduces.

When should thresholds be reviewed again?

Review when the document population, extraction method, validation rules or intended action changes, and when oversight finds a material failure. Keep the previous policy and its evidence available for comparison. A quieter queue alone is not a reason to declare the new setting better; inspect the outcomes behind the change.

Replace a guessed cutoff with a tested acceptance rule

Yes AI can help scope the document sample, critical checks and review boundary for one workflow. Bring the action you want to enable and the errors that must stop it.

All discussions held in confidence. Australian-based consultants.