Building an AI automation that works in a demo takes days. Keeping it working for a year is a different job, and it's the one that decides whether the automation actually saves time.
The pattern we see repeatedly: an automation launches, performs well for a few weeks, and then degrades. Nobody notices at first because it doesn't fail loudly. Invoices get miscategorized. Tickets go to the wrong queue. A report quietly stops including one data source. By the time someone spots it, the team has lost trust and gone back to doing the work by hand.
That last part is the real cost. A broken automation doesn't just stop saving time. It makes the next automation harder to get approved, because the team remembers the last one.
Here's why automations break, and what to do about each cause.
1. The model changed
Model providers update and retire models regularly. A new version may be better on average and still behave differently on your specific task: different formatting, different judgment on edge cases, different handling of your prompt.
A common example: an extraction step that returned dates as 2026-03-01 starts returning "March 1, 2026" after a model update. Nothing errors. The next system simply can't parse the field, and records pile up with empty dates.
What helps:
- Pin model versions where the provider allows it, and upgrade deliberately.
- Keep a small evaluation set of real examples with known correct outputs, and run it before switching versions.
- Validate the model's output against a strict schema, so format drift fails loudly instead of silently.
- Track provider deprecation notices so upgrades happen on your schedule, not theirs.
2. The inputs changed
Automations are built on examples from the time they were designed. Then a supplier changes its invoice layout, marketing launches a new form, or customers start writing in a new language. The automation sees inputs it was never tested on.
Language models are good at handling variation, which is part of the problem. They rarely refuse. Faced with an unfamiliar input, they produce a plausible answer, and plausible-but-wrong is harder to catch than an error.
What helps:
- Validate inputs and route anything unexpected to a human instead of guessing.
- Ask the model to report its confidence or flag missing information, and treat low confidence as a reason for review.
- Sample a few processed items every week and check them.
- Watch the share of items flagged for review. A rising rate is an early warning.
3. The systems around it changed
Most of an automation isn't AI. It's integrations: reading from an inbox, writing to a CRM, calling an accounting API. APIs change, credentials expire, rate limits tighten, fields get renamed.
These failures are often the most damaging, because they can be partial. The automation still runs, still reports success, and quietly skips every record where the renamed field matters.
What helps:
- Treat each integration as something that will fail, and handle errors explicitly.
- Alert on failures rather than retrying silently forever.
- Reconcile counts: if 120 emails came in and 95 records were created, someone should know why.
- Store credentials centrally with clear ownership and expiry dates.
4. The process changed
The business moves on. A new approval step is added, a product line is discontinued, a policy changes. The automation still follows last year's process.
This is the hardest cause to detect technically, because nothing is broken from the system's point of view. The automation does exactly what it was designed to do. It's just no longer what the business needs.
What helps:
- Document what the automation does in plain language, next to the process it supports.
- Make someone in the business responsible for telling the automation owner when the process changes.
- Review each automation against the current process at least quarterly.
5. Nobody owns it
The root cause behind most of the above. The automation was built as a project, the project ended, and the person who built it moved on to other work. There's no one whose job it is to notice when it degrades.
This is the single biggest predictor of whether an automation survives its first year.
Ownership doesn't have to mean a full-time role. It means one named person who receives the alerts, looks at the weekly sample and has the authority to pause the automation when something looks wrong. Pick a backup for holidays, too.
Early warning signs
Automations rarely fail all at once. These signals usually appear weeks before anyone complains, and each one takes little effort to watch:
- Volume drops without a business reason. Fewer items processed than usual often means an integration is skipping records, not that there's less work.
- The review rate creeps up. More items flagged as uncertain suggests the inputs have shifted away from what the automation was built for.
- People start working around it. Staff quietly redoing outputs, or keeping a side spreadsheet "just in case," is a strong signal that trust is slipping.
- Downstream teams find errors first. If finance or support spots mistakes before the automation owner does, monitoring has a gap.
- Nobody can say when it was last checked. If the answer to "is it working?" is "I think so," it's time for a sample check.
Treat any one of these as a reason to look closer. Two together usually mean something has already changed.
What a healthy automation looks like
| Area | Minimum standard |
|---|---|
| Ownership | A named owner, and a backup |
| Monitoring | Alerts on failures, volume drops and review-rate spikes |
| Quality | Weekly sample checks against an evaluation set |
| Change control | Model and prompt changes tested before release |
| Documentation | What it does, what it touches, how to pause it |
| Human review | A clear path for items the automation isn't sure about |
| Reporting | Hours saved and error rate, reported monthly |
The last row matters more than it seems. If you're not measuring hours saved against a baseline, you can't tell whether the automation is still worth running, and you can't make the case for the next one.
A weekly 30-minute routine
Most of this can be kept up with a short, regular check. A routine that works for many teams:
- Look at the numbers. Volume processed, failures and items sent for review this week compared with the last four weeks. Investigate anything that moved sharply.
- Check a sample. Pick five to ten processed items at random and confirm the output is correct. Add any mistakes to the evaluation set.
- Read the review queue. What kinds of items is the automation unsure about? A new pattern often means the inputs changed.
- Check upcoming changes. Model deprecations, API changes, credential expiry dates and planned process changes in the next month.
- Write two lines. What you checked and what you changed. Over a year, this log becomes the most useful documentation the automation has.
Make pausing easy
Every automation should have a documented, one-step way to pause it and fall back to the manual process. Teams that can pause safely act quickly when something looks wrong. Teams that can't tend to let problems run.
Build for the second year
When you plan an automation, plan its operation at the same time: who owns it, how it's monitored, how changes are tested and what happens when it's unsure. It's less exciting than the build, and it's what turns a demo into something your team can rely on.
A useful test before launch: imagine the person who built the automation leaves next month. Could someone else tell whether it's working, find out why it isn't, and pause it safely? If the answer is yes, the automation is ready for its second year.