How Businesses Can Automate Manual Workflows With AI
Contents
There is a version of AI automation that works, and it looks considerably less impressive than the demos. It does not replace a department. It removes a specific, repetitive, well-bounded task that a person currently does forty times a day — and it does so with a human still checking the output.
The projects that fail usually fail for the same reason: they started from the technology rather than from the process.
Start with the process, not the model
Before any tooling decision, the workflow has to be written down: what triggers it, what information it needs, what decision gets made, what happens to the output, and — crucially — what the exceptions are. Almost every manual business process is 70% routine and 30% exception, and the exceptions are where the value and the risk both sit.
This exercise regularly reveals that the bottleneck is not the task anyone assumed it was. It is worth doing even if the project stops there.
Three categories that genuinely work
Language models are good at a narrower set of things than the marketing suggests, but within that set they are very good and getting cheaper.
- Extraction from unstructured documents. Pulling line items from invoices, fields from application forms, terms from contracts. This used to require rigid templates per supplier; it no longer does. This is the highest-return category for most businesses by a distance.
- Classification and routing. Deciding which team an incoming email, ticket or enquiry belongs to, and how urgent it is. The task is bounded, the output is a label from a fixed list, and errors are cheap to correct.
- Drafting a first version. Reply drafts, summaries of long threads, first-pass report text. The output is explicitly a draft that a person edits — which is exactly the right framing, because it keeps a human accountable for what goes out.
Notice what these have in common. Each has a bounded input, a checkable output, and a person who remains responsible for the result.
Where it fails
Language models are not deterministic. The same input can produce a different output, and the model will produce a confident answer when it does not know. That rules out certain categories regardless of how good the model gets.
- Arithmetic and financial calculation. Do not ask a model to compute a total. Have it extract the figures and let ordinary code do the arithmetic, where the result is reproducible and testable.
- Anything requiring an identical answer every time. Compliance determinations, pricing rules, eligibility decisions. Encode these as rules; use the model only to read the inputs the rules consume.
- Irreversible actions without review. Sending payment, deleting records, dispatching customer communications. The step that commits should be a person or a deterministic rule, not a generated decision.
Design the human step deliberately
"Human in the loop" becomes theatre if the human is shown fifty confident-looking outputs an hour and asked to approve them. Nobody reads the fiftieth one carefully. Review has to be designed, not just present.
In practice that means surfacing the model's uncertainty rather than hiding it, routing low-confidence cases to a person and letting high-confidence ones through, showing the source passage next to every extracted field so verification takes a glance rather than a search, and making corrections cheap to record — because those corrections are the measurement data for whether the system is improving.
Measure it honestly
Before launch, record how long the task takes today and how often it is currently done wrong. Without that baseline there is no way to tell whether the automation helped, and the conversation defaults to impressions.
After launch, the number that matters is not model accuracy in isolation. It is end-to-end time and error rate for the whole process including the review step. A system that is 95% accurate but requires every output to be checked closely may save nothing at all. A system that is 85% accurate but flags its own uncertain cases reliably can save a great deal.
On cost
Per-call model costs have fallen sharply and are rarely the deciding factor at business volumes. The real costs are integration with the systems the process already touches, the review interface, and ongoing evaluation as documents, formats and suppliers change. Budget for the third one specifically — an automation that was accurate at launch and unmonitored for a year is a liability, not an asset.
Start with one process. Pick the one that is high-volume, low-variety and currently annoying — extraction from a recurring document type is the usual best first candidate. Measure it, ship it with a review step, and expand only once it has held up in production for a month.