FXNL
Back to Insights

Insight

What Business Processes Are Actually Worth Automating With AI?

Automate work when the improvement is worth the full cost, the result can be evaluated, and mistakes can be contained.

By FXNL Team10 min read

Just because AI can automate something doesn’t mean it should.

That’s true, but it isn’t very useful on its own. A business still has to decide which work deserves attention, and how to improve it.

Start by looking for work that is costly, slow, error-prone or frustrating—and for chances to deliver a better outcome. Then test whether AI is the right way to improve it.

Use AI for the parts that need interpretation. Use the simplest reliable method for everything else.

Possible Is Not the Same as Worthwhile

“AI can do this” and “this is worth changing” are different statements.

AI performance is uneven. Harvard Business School research found AI assistance improved some knowledge tasks and made others worse, even within seemingly similar work. Being able to do a task doesn’t settle whether changing the process pays off.

A worthwhile opportunity depends on:

  • How often the work happens, and in what volume
  • How long it takes, and how long it waits
  • How much rework it creates
  • Its effect on quality, service and growth
  • What an error costs
  • The effort to implement it
  • The ongoing support it needs

Repetition helps the economics, but it isn’t enough by itself, and it’s not a requirement. Valuable, irregular work like research synthesis or expert preparation can benefit too.

One more distinction matters from the start. An hour saved is capacity, not necessarily cash. We come back to that below.

Five Questions Before You Automate

These five questions aren’t a score. Adding points together lets large savings cancel out unacceptable risk, and hides uncertainty behind a number. Answer each question with evidence, and treat a clear “no” on the fourth or fifth as a stop.

  1. Is it worth improving?

    Ask about:

    • Cases per month or year
    • Handling time
    • Waiting time
    • Rework
    • Service impact
    • Revenue or growth impact
    • What people could do with the capacity returned

    Low-value work may not justify bespoke automation. If the value is weak, stop—or use an inexpensive tool you already have.

  2. Can we make it work reliably?

    Look at:

    • Representative inputs, including the awkward ones
    • Access to the information the work needs
    • Which sources are authoritative
    • What a good output looks like
    • How often exceptions occur
    • How it actually performs when tested

    Then ask the AI-specific questions. Can the system obtain the facts? Can correctness be checked at reasonable cost? What happens when a confident-looking answer is wrong? Can the workflow abstain or escalate? If the answers are unclear, the next step is a bounded feasibility test, not a production promise.

  3. What will it take to run?

    Include:

    • Implementation
    • Integrations
    • Software
    • AI usage
    • Review
    • Exception handling
    • Maintenance
    • Change management
    • Support

    A useful AI response can still be a poor operational investment.

  4. Can we contain mistakes?

    Consider:

    • How severe an error would be
    • Whether it can be reversed
    • How many cases would be exposed
    • Whether sensitive information is involved
    • What the system is allowed to do
    • How errors would be detected
    • Where problems escalate

    High projected savings don’t compensate for unacceptable residual risk. If the risk can’t be brought to an acceptable level, that’s a veto, not a weighting.

  5. Who owns the change?

    This is the readiness gate. Look for:

    • A process owner
    • An agreed process
    • Training
    • Support
    • A manual fallback

    Without ownership, the opportunity may belong in discovery rather than production.

Question three is usually where costs surprise people. What Actually Makes an AI Project Expensive? covers what drives them.

These questions draw on GSA’s process-improvement guidance, NIST’s risk framing and the lessons of tested performance. They’re a practical discussion tool, not a validated scientific instrument.

When AI Isn’t the Answer

AI is one option among several. The simplest reliable one usually wins.

Ways to improve a process, from simplest to most autonomous
Remove or simplify the workThe step exists because of duplication, habit or poor process design.Example: Stop rekeying data by collecting a required field once.
Rules, formulas, APIs or traditional automationDecisions can be explicitly specified and inputs are structured.Example: Match invoice IDs, sum amounts, apply approval thresholds, synchronize records.
AI assistanceA person benefits from interpretation, drafting or synthesis.Example: Prepare a cited proposal draft for someone to review.
AI inside a fixed workflowLanguage or documents vary, but the business process itself is predictable.Example: Extract invoice fields, validate totals and supplier with rules, queue exceptions, post once approved.
Model-directed agentThe next useful action genuinely depends on what the system discovers along the way.Example: Investigate an internal support issue across approved systems, with bounded actions and escalation.

A multi-step automation is not automatically an AI agent.

Decide separately which steps need probabilistic interpretation and which decisions need dynamic control. An agent is justified when its flexibility creates measurable value beyond a simpler design.

That’s also the advice in Anthropic’s engineering guidance on building agents, which recommends simpler solutions when they suffice. It’s informed vendor practice, not independent proof that agents are generally worse. And a federal inventory of robotic process automation shows data comparison, document collection and reporting being automated well before generative AI.

For a closer look at where each level of autonomy fits, see AI Agents vs. AI Automation: What Does Your Business Actually Need?

What Bad Candidates Look Like

These warning signs are defaults, not a blacklist. Each has an important exception.

  • Trivial economic value

    Defer bespoke automation.

    But: Rare work isn’t automatically low-value. A major proposal or regulatory submission can matter a great deal, and an existing assistant may cost little to try.

  • A broken or disputed process

    Clarify the purpose, decisions and ownership first.

    But: Understand why the process is failing before encoding it into automation. AI can even help analyze the process—just don’t automate a disagreement.

  • Inaccessible, conflicting or stale information

    Fix access and agree which source is authoritative.

    But: AI can help classify and clean information, but its output shouldn’t become the official record without validation.

  • Nobody can define a good result

    Do discovery first. Create examples or a rubric.

    But: Not every output needs an exact right answer. Creative work can still have useful, consistent criteria.

  • Serious, hard-to-reverse errors

    Keep consequential authority with qualified people, or defer.

    But: High-risk work isn’t automatically off-limits. AI may help prepare evidence or flag missing information while a qualified person decides.

  • No owner, support or fallback

    Not ready for production.

    But: A supervised discovery exercise can still investigate the opportunity.

  • Constantly changing requirements

    Avoid a large fixed build until the process settles.

    But: A configurable aid may still work if updates are cheap and someone owns them.

Notice what isn’t on the list: unstructured inputs. Emails, documents and other variable language are often exactly why AI is being considered. Stable, structured data more often favors rules.

The strongest research warning is that results vary. A study of AI-assisted customer support found substantial gains. A randomized trial by METR found that 16 experienced developers took 19% longer on 246 tasks with early-2025 AI tools, despite believing they were faster. Neither study lets anyone classify every support or coding process in advance. Both are reasons to measure.

Human in the Loop Is Not a Magic Safety Feature

Adding a reviewer is a design choice to evaluate, not a safety sticker. Review has its own time cost, its own error rate and its own effect on the queue.

A 2024 meta-analysis in Nature Human Behaviour covering 106 experiments found that, on average, human–AI combinations performed worse than the better of the human or the AI alone. Results varied by task, and were more promising for content creation than for decisions. The studies predate much of today’s generative AI.

That doesn’t mean removing reviewers from consequential decisions. It means the combined process has to be tested, not assumed. A reviewer needs:

  • The source evidence and uncertainties, not just a polished recommendation
  • The authority to disagree or stop the action
  • Sufficient expertise in the work
  • Enough time to review properly

Measure missed errors and incorrect overrides as well as review time. Route missing information, unusual cases and failed checks to people, and don’t rely on the model’s own confidence alone.

Review at every step isn’t always necessary either. For low-impact, reversible actions, control can come from a limited scope, testing before launch, monitoring and escalation.

A Simple Economic Test

The core calculation is straightforward:

Annual net hours released = eligible cases × reduction in total human handling time.

The catch is in “total.” Future handling time has to include:

  • Review
  • Correction
  • Exception handling

Then separate two things that are easy to blur: capacity value and cash savings.

Salaried hours don’t automatically become savings. Capacity creates financial value only when the organization can credibly use it, for example through:

  • Additional productive work
  • Growth
  • Avoided hiring
  • Reduced overtime
  • Improved service

Where that mechanism isn’t proven, report hours and potential capacity separately from cash. We go into this in more depth in How Do You Know If an AI Project Is Worth the Investment?

A worked example

Every number below is an illustrative assumption—not a market benchmark, a vendor quote or an FXNL client result.

Capacity (illustrative)

Eligible invoices
500 / month
Current human time per invoice
8 minutes
Future human time, including review and exceptions
3 minutes
Net hours released (500 × 12 × 5 ÷ 60)
500 hours / year
Assumed value per hour
$40
Share the business can realistically use
60%
Potential annual capacity value
$12,000

Cost (illustrative)

Software and support
$6,000 / year
Implementation (one-time)
$8,000
First-year net benefit ($12,000 − $6,000 − $8,000)
−$2,000

Five hundred hours a year sounds like a clear win. Under these assumptions, the first year loses money before counting any separately justified quality benefit. From year two, the recurring benefit would be about $6,000 a year.

An opportunity can look attractive on hours alone and still have weak first-year economics.

That doesn’t make it a bad idea. It means it needs stronger economics, a cheaper route, or a different objective.

Ordinary Examples

These show how the questions apply to common work. They’re illustrations of the method, not customer success stories.

  • Document processing

    When it fits

    Many readable documents, such as invoices, share the same required business fields.

    What to watch

    Let AI extract; let rules check amounts, duplicates and suppliers; have a person approve anything consequential before it’s posted.

  • Customer inquiries

    When it fits

    Questions recur, the knowledge behind them is current, and answer quality can be measured.

    What to watch

    Start by assisting staff or answering narrow questions. Route disputes and unusual cases to people. A published study of 5,172 support agents at one company found AI assistance increased issues resolved per hour by 15% on average, unevenly across workers—evidence for assistance, not for replacing a team.

  • Intake

    When it fits

    Emails or forms can be turned into a defined case record.

    What to watch

    Verify required fields and identity. Missing facts should become follow-up questions, not guesses.

  • Reporting

    When it fits

    The source data and definitions recur and are stable.

    What to watch

    Calculate the numbers with rules. Use AI for a draft explanation, then check it against the numbers.

  • Reconciliation

    When it fits

    Identifiers and matching rules are explicit.

    What to watch

    Use rules for matching and arithmetic. AI can help explain or categorize what doesn’t match. A plausible narrative is not a balanced ledger.

  • Knowledge retrieval

    When it fits

    The source content is authoritative, current and access-controlled.

    What to watch

    Test whether answers are complete, faithful to the sources, and whether the system says when it can’t answer. The UK Government Digital Service reported its GOV.UK Chat pilot improving from an early 76% answer-accuracy benchmark to 90% in later evaluation. The lesson is the need for scoped sources and ongoing evaluation—not that a 10% error rate is acceptable elsewhere.

  • Proposal preparation

    When it fits

    Approved content exists and proposals share recurring sections.

    What to watch

    Draft from approved material. People approve prices, promises, compliance statements and what makes the offer different.

Whichever you test, a useful pilot measures quality, total handling time, exceptions, adoption and cost—against how the work is done today.

If you’re still deciding whether process improvement should be your starting point at all, see Where Should a Business Start With AI?

If you have a candidate in mind, the AI Investment Estimator gives a planning range for what it might involve. It’s a starting estimate, not proof of return. Our Projects show what this kind of workflow work looks like in practice.

The question isn’t whether AI can do the work. It’s whether changing the work is worth it.