When Not to Use AI: Seven Workflows to Skip
When not to use AI: seven workflows where rules, process fixes or people beat a model, with public evidence and what to build instead.
AI Engineer, Viithiisys

Why do we turn down some AI requests?
We decline an AI request when a simpler tool does the job better, or when the cost of a wrong answer outweighs the gain. Viithiisys has not published a decline rate, so this guide relies on named public sources instead.
Viithiisys has built software since 2007 and shipped 500+ projects, most of them before large language models existed. That history shapes the default: reach for the plain tool first. Anthropic's own guidance on building effective agents says the same thing, which is to find the simplest solution and add complexity only when it demonstrably helps.
What test do we apply before saying yes?
Seven questions decide whether when not to use AI is the right answer for a workflow. Is the volume high enough to justify the build? Can a mistake be reversed? Is the logic already a rule? Does the data exist? Is the process sound? Can the decision be explained? Does the saving exceed the upkeep?
A single "no" is enough to stop and reconsider. The sections below take each question in turn, with the failure mode and the alternative.
1. Low volume, high variance
AI automation is not worth it when a task happens a few times a week and every instance differs. There is no pattern to learn, and building an evaluation set costs more than doing the work by hand.
Consider a team that reviews bespoke supplier contracts, perhaps twenty a month, each with different clauses. A model can draft a summary, but someone must still read the contract to check it. The reviewer has now done the work twice.
What is the break-even arithmetic?
Add the hours to build, test, monitor and update the system. Divide by the hours saved each month. If payback takes longer than the time before the process or the model changes, stop.
A worked example makes this concrete. Suppose the build takes 120 hours, and monitoring and prompt updates add another 6 hours a month. If the tool saves 10 hours a month, the net saving is 4 hours, and payback takes 30 months. Few processes, and no model version, stay unchanged for two and a half years. At 60 saved hours a month the same build pays back in about two months.
Variance matters as much as volume. Ten thousand near-identical invoices a month suit automation. Forty one-off requests do not, and a checklist with a shared template usually captures most of the benefit.
2. Where is a wrong answer unrecoverable?
Avoid AI where one wrong output causes harm that cannot be reversed: irreversible payments, deleting records, clinical dosing, legal filings. Models produce fluent errors, and fluency hides them from a tired reviewer.
The standard cautionary case is legal. Reuters reported in 2023 that a federal judge sanctioned New York lawyers who filed a brief containing cases invented by ChatGPT. The citations looked real, which is exactly the failure mode.
Does human review fix the problem?
Only partly. Review works when checking is much faster than producing, such as comparing a draft against a source document.
It fails when the reviewer must redo the task to verify it, or when volume is high enough that approval becomes a reflex. A reviewer who approves two hundred outputs a day, nearly all of them correct, stops reading closely, and the rare bad one passes with the rest. If you cannot undo the error and cannot cheaply check it, keep a person or a deterministic control in charge.
3. Where the rule is actually deterministic
If the decision follows written logic, such as a tax table, an eligibility threshold or a routing rule, use code. The comparison of rules vs AI is lopsided here: rules are testable, free per call and give the same answer every time.
A model can return different outputs for the same input, and every call costs tokens and latency. Teams sometimes route a refund-policy check through an LLM because the policy lives in a PDF. Transcribing the policy into twenty lines of conditions is faster to build and far easier to audit.
The cheapest AI project is the one a rule table or a process change made unnecessary.
Where does a hybrid make sense?
Use a model for the untidy input and rules for the decision. An LLM can pull the order number, product and complaint type from a rambling email, then a rule decides the outcome. Our document AI work follows this split: extraction is probabilistic, the decision is deterministic and logged.
4. Where does the data not exist yet?
AI fails when the history it needs was never recorded. A prediction model needs labelled outcomes, and a new product with three months of messy records cannot supply them.
Generative models can work zero-shot, which tempts teams to skip the data question. But you still need a set of real cases with known right answers to test against. Without one, you cannot tell whether the system works, only that it sounds confident.
What should happen first?
Fix the capture before the model. Define the fields, record the outcomes consistently and store the reasons behind each decision. A few months of clean records beats a year of free-text notes.
That is data engineering work, not AI work, and it is the honest first deliverable. Once the data exists, the model question is far cheaper to answer.
5. Where the real problem is a process, not a tool
If the workflow is slow because of six sign-offs, duplicate data entry or unclear ownership, AI will automate the mess. You get the same bottleneck at higher speed and with a new failure surface.
A common pattern is a team that re-keys data between two systems that should be integrated, then asks for an AI agent to do the re-keying. An API integration removes the task entirely. Another is an approval chain where four of the six approvers never reject anything, which no model should be trained to imitate.
How do you tell a process problem from a tool problem?
Map the workflow end to end and mark every wait, hand-off and re-entry. If most of the elapsed time is waiting, not working, the fix is structural. A request that takes nine days to clear but holds only forty minutes of actual work is a queueing problem, and no model shortens a queue.
Applying AI workflow automation to a sound process is worthwhile. Applying it to a broken one is how projects stall after the pilot.
6. Where compliance requires explainability you can't give
Where a regulator or a customer is entitled to a reason for a decision, a system that cannot produce one is a liability. Credit, hiring, insurance and benefits decisions sit in this group.
The US Consumer Financial Protection Bureau's Circular 2022-03 states that creditors must give specific reasons for an adverse decision even when they use complex algorithms. In Europe, the EU AI Act lists credit scoring and employment screening among its high-risk uses, and GDPR Article 22 restricts solely automated decisions with significant effects.
Is an LLM still usable nearby?
Yes, in a supporting role. A model can summarise a case file or draft a letter while a named person, following documented criteria, makes the decision. Check with counsel in your jurisdiction, because the boundary between assisting and deciding is where most compliance arguments happen.
7. Where the ROI is real but under the cost of maintenance
A workflow can save genuine hours and still lose money. The saving is visible on day one, while the upkeep arrives quietly over the following year.
That upkeep is concrete. Providers retire models on published schedules, as the Anthropic model deprecation page shows, and each migration means re-running your tests. Prompts drift as inputs change, and someone has to notice.
What does maintenance actually include?
- Re-running the evaluation set after every model or prompt change
- Monitoring output quality, cost and latency in production
- Updating prompts and retrieval sources as the business changes
- Handling the exceptions the system escalates
If the saving is ten hours a month and one engineer must stay familiar with the system, the sums rarely work. This is the unglamorous side of MLOps, and it belongs in the business case before the build starts.
What do we recommend instead in each case?
For each of the seven cases there is a cheaper, safer default. The table sets out the alternative we suggest and the signal that would justify revisiting AI later.
| Case | Recommend instead | Revisit AI when |
|---|---|---|
| Low volume, high variance | Checklist, shared template, a short script | Volume grows and cases converge |
| Unrecoverable errors | Person decides, deterministic controls | Errors become reversible and cheap to check |
| Deterministic rule | Rules engine or plain code | Input becomes unstructured |
| No data yet | Instrument the process, capture outcomes | A labelled history exists |
| Broken process | Redesign, integrate systems, clarify ownership | The process is stable |
| Explainability required | Documented criteria, model as assistant only | Regulator accepts the explanation method |
| Maintenance exceeds saving | Leave manual or use a script | Scale raises the saving |
How do we put this to a client?
We say it plainly on the first call. Viithiisys is an independent software vendor in Mohali with a Canadian office in Markham, and teams that have delivered for Paytm, Snapdeal, IKEA, Nestle, Shiprocket and Vikram Solar do not need to invent AI work to stay busy. You can read more about who we are and how we work.
If you suspect one of these seven applies to your plans, the broken workflow assessment gives you an honest read on whether to build, fix or leave it alone. For a narrower question, contact us directly.
FAQ
- When should you not use AI in a business process?
- Skip AI when the task is rare, when a wrong answer cannot be undone, when a fixed rule already decides the outcome, when no labelled data exists, when the process itself is broken, or when a regulator needs a specific explanation of each decision.
- Is a rules engine better than AI for some workflows?
- Yes. If the logic can be written down as thresholds, lookups and conditions, a rules engine gives the same answer every time, is cheap to run and can be tested line by line. Use a model only for the messy input that feeds those rules.
- How do I know if AI automation is not worth it for my team?
- Compare hours saved per month against the hours needed to build, evaluate, monitor and update the system. If the saving is small or the task changes often, the maintenance burden usually wins and a checklist or a script is the better answer.
- Can an AI project be saved if the data is not ready?
- Usually the better move is to pause the model work and fix the data first. Instrument the process, capture outcomes consistently and build an evaluation set. Models trained or tested on thin or inconsistent records produce confident answers nobody can trust.