PRACTICAL
INTELLIGENCE.
← All guides

ORIGINAL GUIDE / Language & documents

Turn a document into a checked record

Define an extraction contract, retain source evidence and validate arithmetic before a proposed value enters another system.

Practical Intelligence · Published · Illustrative examples, not customer performance claims

Related research signals ↗

Document extraction becomes useful when a reviewer can trace a field back to the original and understand what happens if it is wrong. Begin with one document type and one destination. Mixing invoices, contracts and handwritten notes into a single pilot makes it harder to see why a field failed.

1. Define the record before the prompt

List the fields you actually need. For an illustrative invoice intake, that might be document reference, supplier, invoice number, invoice date, currency, line amounts, tax and total. Specify types and formats, whether a field is required, and what an absent value looks like. An unknown currency must stay unknown; it should not silently become your office’s default currency.

Keep extraction separate from interpretation. A document date is a value printed in the source. Whether it represents a payment deadline, an issue date or a service date is a separate meaning that should be supported by a label or reviewed.

2. Carry evidence with the field

Store the document identifier and page or source location alongside each proposed field. Retain the exact source text when appropriate. A reviewer should be able to compare the proposed value with its evidence without hunting through unrelated files. Restrict access to the document and derived data to the people who need it.

Worked example: a mismatch you can calculate

Suppose a synthetic invoice lists two items at ₹500 each, tax of ₹180, and a printed total of ₹1,280. The component arithmetic gives ₹1,180. Preserve both the extracted printed total and the computed total, flag the difference of ₹100, and require review. Do not “correct” the source by replacing the printed value without an explanation.

3. Put deterministic checks around the proposal

A model’s self-reported confidence is not a measured probability of correctness. Track your observed field errors on reviewed examples instead. The NIST AI Risk Management Framework provides broader context for evaluating and managing risks around an AI system; the field contract here is a practical exercise, not a claim of NIST certification.

4. Evaluate the destination as well

Test whether a reviewer can reject a proposal, edit a value, see the original, and retry a failed save without producing a duplicate. Measure both field correctness and the effort needed to accept a whole record. A system that reads most fields correctly but creates duplicate entries can still increase the workload.

A small next step

Make five synthetic documents: one ordinary, one missing a field, one with ambiguous dates, one with incorrect arithmetic and one duplicate. Write the expected outcome for each before running the extractor. That is a more useful starting point than collecting a large pile of sensitive documents.