Soft failure: when AI makes progress feel real

Generative AI has an unusual failure mode: it can fail while appearing to succeed.

Traditional software usually has clearer limits. Ask Microsoft Word to make a financial transaction and it cannot do it. The failure is obvious.

Professional reviewing an AI dashboard that appears complete while hidden warnings and evidence gaps sit beneath the surface.

Generative AI behaves differently. Ask it for something outside its real capability and it will often still produce a plausible response. It may suggest options, explain a process or generate a polished answer.

The interaction continues. Nothing visibly breaks.

This is what I think of as soft failure: an interaction that creates the appearance of progress without necessarily achieving the underlying objective.

For organisations using AI in safety, security, risk and compliance, this matters. A plausible answer can be mistaken for a reliable one, and in high-consequence environments that distinction is important.

What is soft failure in AI?

A hard failure is easy to recognise. The software stops, rejects the request or tells us the task cannot be completed.

Soft failure is different.

An AI system may produce something useful-looking even when:

  • the evidence is incomplete;
  • the answer only partly addresses the task;
  • an assumption has been presented as fact;
  • important uncertainty has been missed;
  • the underlying objective has not actually been achieved.

The problem is not simply that generative AI can be wrong. All systems can fail.

The problem is that AI failure can look remarkably similar to AI success.

That makes it harder for users to know when additional checking is required.

When activity looks like progress

I see a similar effect in commercial meetings.

A prospective client shows interest. After an initial conversation, more people are invited to the next meeting. Eight or ten people join. The meeting lasts an hour. People ask questions and appear engaged.

It can feel like a successful meeting.

But afterwards there may be no agreed problem to solve, no defined use case, no discussion of budget, no owner and no next commercial step.

The meeting itself was real. The engagement was real.

But curiosity is not the same as buying intent.

The danger comes when we mistake the signals associated with progress for progress itself.

Generative AI can create the same effect.

Why AI can feel like progress

A session with generative AI can produce pages of analysis, ideas, recommendations and actions in a few minutes.

The experience feels productive because there is so much visible output.

This is where I find the analogy with dreams useful.

I do not mean that the AI is dreaming. I mean the experience for the user can become dream-like.

While we are inside a dream, events can feel coherent and convincing. It is only when we wake up that we test them against reality.

AI can create a similar experience.

We can finish an interaction feeling that a problem has been analysed, a decision improved or a piece of work completed.

The more useful question is what has actually changed.

Has the evidence been checked?

Has somebody made a decision?

Has an action been taken?

Has the underlying problem moved forward?

Or have we mainly produced a convincing description of progress?

Why fluent AI output makes failure harder to spot

Generative AI is very good at producing readable text quickly and at scale.

That is one of its major benefits. It is also part of the risk.

Easy reading is hard writing. People normally have to work to turn complex subjects into clear, accessible prose. Generative AI can produce that fluency almost instantly.

Fluent writing can therefore create an impression of competence that is stronger than the evidence behind the answer.

There is a wider question here too.

Spoken language has been central to human communication for thousands of years. Mass literacy across the general population is comparatively recent.

We are now encountering something historically unusual: machines capable of generating almost unlimited quantities of fluent written language.

We are still learning how to scrutinise it.

If nine consecutive AI answers appear useful and convincing, it is natural to approach the tenth with less suspicion.

The danger is that the tenth may concern something much more consequential.

Why soft failure matters in safety, security and risk

A soft failure in a low-consequence task may result in a poor recommendation or some wasted time.

A soft failure in a high-consequence environment can become part of a decision that later needs to be explained and defended.

Consider an AI-supported review of:

The AI output may look reasonable. A reviewer may accept it and move on.

If something later goes wrong, an internal investigation, regulator, insurer or court may ask different questions.

What evidence supported the decision?

What information was available at the time?

Which sources were used?

Where was uncertainty identified?

What assumptions were made?

Who reviewed the result?

At that point, a plausible answer is not enough.

The organisation needs to be able to show how the decision was reached.

AI governance should make failure visible

This is one of the important differences between using generative AI as a general productivity tool and using AI in high-consequence operational work.

The objective should not simply be to make AI produce better answers.

It should also be to make weak answers, missing evidence and uncertainty easier to see.

That means putting controls around the task.

Depending on the workflow, these may include:

  • defining precisely what the AI is being asked to do;
  • limiting it to approved information sources;
  • requiring important conclusions to be supported by evidence;
  • identifying uncertainty and missing information;
  • escalating where the evidence is insufficient;
  • keeping a human responsible for consequential decisions;
  • recording the information, assessment and final decision;
  • monitoring performance and learning from later outcomes.

These controls do not remove every possibility of AI failure.

They make failure easier to detect, challenge and understand.

Three questions to ask when AI appears to have helped

Whenever an AI interaction appears successful, three questions are useful:

  1. What has actually changed as a result of this interaction?
  2. What evidence tells us the output is reliable?
  3. Could we explain and defend the resulting decision later?

Generative AI is unusually good at making progress feel real.

In serious environments, the challenge is making sure it is.

Author bio: Andrew Tollinton

Andrew Tollinton Founder SIRV and author

Andrew Tollinton is CEO and Co-Founder of SIRV, which builds operational AI for safety, security and resilience teams. He focuses on practical, controlled AI use in serious environments, with particular interest in evidence, accountability and human judgement. Andrew chairs the Institute of Strategic Risk Management’s AI in Risk Management Special Interest Group and speaks regularly on AI governance and operational resilience.

"SIRV helped us move beyond basic reporting into a system that actively supports decision-making". Les O'Gorman, Director of Facilities, UCB - Pharma and Life Sciences

css.php