Most of the AI automation work I do now isn’t about whether a model can do a task. It usually can, most of the time. The real question is what happens on the occasions it gets it wrong, and whether anyone will notice before it matters.
So I use a simple rule, and I use it on every workflow before I write any code.
The rule
A model can act on its own when the step is cheap to check and cheap to undo. When a step is expensive to undo, or someone outside the business will see the result, the model drafts and a person approves.
That gives you three buckets:
- Do it: reading, sorting, tagging, summarising, filling in internal fields. If it’s wrong, someone fixes it in ten seconds and nobody outside ever knew.
- Draft it: customer emails, quotes, journal entries, anything that changes a balance. The model does the slow part, a person clicks approve.
- Don’t touch it: refunds over a threshold, anything legal, firing off payments. The model can flag these. It doesn’t act.
What this looks like in Flo
Flo is my own product. It reads invoices out of a Gmail inbox, pulls out the supplier, amount, GST and due date, and files them as draft bills in Xero. Notice the word draft. Flo never approves a bill, and it never pays one.
Reading the invoice and filling in the fields is firmly in the first bucket. If it gets the due date wrong, the bookkeeper sees it when reviewing the draft and fixes it. Creating an approved bill would be in the second bucket, because once it’s approved it feeds into payment runs. So Flo stops there and lets a person look.
Flo also scores how confident it is. When a supplier it has seen fifty times sends the usual invoice, the draft goes straight in. When it’s a new supplier, or the total doesn’t match the line items, or the GST looks off, it gets flagged for a closer look. Roughly one invoice in twelve ends up flagged, and that’s about right. Too many flags and people stop reading them.
A workflow that got it wrong
A while back I inherited an automation for a trades business that drafted replies to quote requests and sent them straight away. It was quick, it was polite, and about once a fortnight it quoted a job the business didn’t do at a price it had made up. A customer turned up with a printout of one of those quotes and expected it honoured.
Nothing about the model changed in the fix. I moved the send step from bucket one to bucket two. Replies now land in a shared inbox as drafts, the office manager reads them over coffee, and most go out within the hour with a single click. Response time went from four minutes to about forty. Nobody has complained, and nobody has turned up with a printout since.
Questions to ask about your own workflow
- If this step is wrong, who finds out, and how long does it take?
- Can it be undone without an apology?
- Does the output leave the business, or touch money?
- If a person has to approve it, is the approval quick enough that they will actually do it properly?
That last one matters more than people think. A review step that takes five minutes per item gets rubber-stamped by week three. Good human in the loop design makes the check fast: show the source next to the draft, highlight what the model was unsure about, and make approve a single click.
If you’re thinking about automating something and aren’t sure where the line should sit, walk me through it. I’ll tell you which bucket each step belongs in.

Liam Hillier
Software, AI and integrations for businesses that have outgrown their spreadsheets. More about Liam.