Skip to content
Ventric
Automate

What is document automation, and where should you start?

Two different jobs get called document automation. Knowing which one you need changes the whole project.

Dan KennedyCo-founderPublished 7 min read

Document automation is the use of software to handle documents that currently pass through someone's hands. That covers creating them, reading them, filing them, and getting the information on them into the systems that need it.

The confusion worth clearing up first is that two quite different jobs share the name, and they need different approaches.

Two different jobs, one name

  • Document generation. Producing a document from data. Contracts, quotes, engagement letters, reports. The information exists in a system and you need it in a formatted document.
  • Document processing. Extracting data from a document. Supplier invoices, timesheets, application forms, certificates. The information arrives as a document and you need it in a system.

Generation is the easier problem, because you control the format. Processing is harder, because someone else does. Most businesses need processing, and most of the difficulty in this field lives there.

Document generation

The pattern: a template with placeholders, data from a system, and a step that merges the two and files the result.

This is a good early automation because it is predictable and the saving is easy to see. If someone produces fifteen quotes a week by copying last week's, changing the details and hoping they caught them all, generation removes both the time and a specific class of embarrassing error.

The things that make it worth doing properly:

  • One template that is genuinely the current one, rather than nine versions in different folders
  • Data pulled from the system of record rather than retyped, so a document cannot contradict the CRM
  • Consistent naming and filing, so the finished document is findable
  • A version history, which matters for anything contractual

The main pitfall is templates with too many conditional variations. If a document has thirty optional clauses depending on circumstances, you are building a rules engine rather than a merge, and the project is larger than it looks.

Document processing

Here you receive a document created by someone else, in a format you do not control, and need reliable data out of it.

Documents fall into three groups, and the group determines the difficulty:

  • Structured. Same layout every time. A form your own business issues. Straightforward: the data is always in the same place.
  • Semi-structured. Same information, different layouts. Supplier invoices are the classic case. Every invoice has a total and a date and a reference; every supplier puts them somewhere different. This is the most common and most valuable category.
  • Unstructured. Free text with information buried in it. Contracts, tender documents, correspondence. Hardest, and the category where AI genuinely earns its place.

For semi-structured documents, modern document intelligence services handle the general case well. They are trained to find a total, a date, a supplier name and line items across varied layouts, without needing a template per supplier. That is a meaningful change from a few years ago, when template-per-supplier was the only reliable route.

For unstructured documents, a language model can extract "what is the contract termination notice period" from a document that never uses those words in that order. Useful, and it needs verification, which is the next section.

Document workflow automation: the whole path

Extraction is one step. The reason a project delivers or disappoints is usually the rest of the path.

The full journey for an inbound document:

  1. Arrival. Email, upload, scan, portal. Multiple routes usually, which needs deciding rather than discovering.
  2. Identification. What kind of document is this? An invoice, a credit note and a statement need different handling.
  3. Filing. Consistent naming and location, immediately, so the document is findable even if later steps fail.
  4. Extraction. Pull the fields you need.
  5. Validation. The step people skip. Does the total match the purchase order? Is the supplier known? Does the date make sense? Is this a duplicate?
  6. Decision. Straight through if everything checks out, to a person if not.
  7. Delivery. Data written into the finance system, CRM or wherever it belongs.
  8. Exception handling. Anything that failed goes to a named person with the reason attached and enough context to act.
  9. Retention. How long it is kept and what happens at the end.

Validation and exception handling are what separate a document automation that works from one that quietly creates a mess. Extraction without validation means wrong data enters your systems faster than before, and takes longer to find.

Where to start

Pick the document type where all of these are true:

  • High volume. Dozens a month at least, or the arithmetic will not work.
  • Predictable purpose, even if the layout varies.
  • Currently rekeyed by a person.
  • Checkable. There is something to validate against, such as a purchase order or an existing record. This is what makes automation safe.
  • Digital on arrival. A PDF by email is a much easier start than paper through the door.

For most businesses that means supplier invoices, and supplier invoices are a good first project precisely because there is a purchase order or an expected amount to check against. Timesheets and delivery notes are close behind for similar reasons.

A deliberately unglamorous suggestion: start with the filing and naming even before extraction. Getting every document consistently named and in the right place is a small piece of work that makes everything afterwards easier, and on its own it removes a surprising amount of daily searching.

How accurate is it, really?

The honest answer is that accuracy depends heavily on document quality and type, and that you should not accept a percentage from a vendor as a promise about your documents. A clean digital PDF from a regular supplier is a different proposition from a phone photograph of a crumpled delivery note.

What matters more than the headline accuracy figure is that the system knows when it is unsure. A process that extracts a field with low confidence and flags it for checking is far more useful than one that is slightly more accurate overall but silently wrong. Confidence thresholds, and a validation step that catches what confidence misses, are what make this safe to rely on.

So design for it: automate the confident majority, route the rest to a person, and measure how big each group is. If ninety per cent goes straight through, you have removed ninety per cent of the work, and you know about the rest rather than discovering it in an audit.

A note on sensitive documents

Many documents worth automating contain personal data: application forms, HR paperwork, anything with customer details. Two things follow.

First, know where the processing happens. Using AI services inside your own Microsoft 365 tenant keeps the data within your environment. Sending documents to a third-party service means a supplier who processes personal data on your behalf, which needs the appropriate agreement and a look at where they store it.

Second, automation is a good moment to fix retention. If you are already redesigning how a document type is filed, adding a retention rule costs almost nothing extra and is easier than retrofitting later.

Questions about this topic

What is the difference between document automation and document management?

Document management is about storing, organising and controlling access to documents: where they live, who can see them, version history and retention. Document automation is about the work done to and with documents: generating them, extracting data from them and moving that data onward. They complement each other, and automation is considerably easier if the management side is already reasonably tidy.

Do we need AI for document automation?

Not always. If your documents have a fixed layout, straightforward rules will read them reliably and be easier to verify. AI earns its place with semi-structured documents such as invoices from many suppliers, where layouts vary, and with unstructured text where the information has to be interpreted rather than located. A good build uses rules where rules work and AI only where they do not.

Can it handle scanned paper documents?

Yes, with the caveat that quality matters much more than with digital files. A flatbed scan at a sensible resolution is usually fine. A photograph taken at an angle in poor light is not, and no amount of software fully compensates. If paper is a significant part of your volume, improving how it gets captured is often the highest-value first step.

Ready to take repetitive admin out of the week?

Automate covers process automation, SharePoint knowledge, Copilot adoption and practical AI on Microsoft 365.