11 Jun 2026

LLMs fail fluently: verification and traceability

LLMs fail fluently: verification and traceability

ChatGPT lost it when I wanted to explain CPSC to an Italian friend.

It started with correct translation and a sensible structure. Then, mid-sentence, it collapsed into multilingual noise: fragments of Russian, Hindi, Chinese, emoji, and strings that looked uncomfortably like internal system artefacts.

This happened in ChatGPT, one of the most widely deployed AI products in the world. Not in a beta, not in some obscure wrapper, but 5.5 instant. And the detail that matters is not that it broke. Everything breaks sometime, it is how it broke.

Traditional software usually fails loudly. An error code, a stack trace, a blue screen of death. Something throws an exception and your monitoring lights up.

Large language models fail fluently.

The answer started out correct and confident. From the product's point of view, the request likely completed successfully. No alert fired, because nothing obvious had gone wrong at the system level. The model answered. It just answered badly.

When you integrate an LLM into your product, you inherit this failure mode along with the capability. The vendor cannot predict every case, and often cannot explain it afterwards. If your integration budget covers the perfect path only, the broken output goes straight to your customer wearing your logo.

Two disciplines follow from this. The first is verification. Your checks have to inspect the output itself, not only the plumbing around it.

Is the response in the expected language? Does it match the expected structure? Does it make claims that need a human before they ship?

A status code tells you the model answered. It tells you nothing about what it said.

The second is traceability. Log every output, with the input, timestamp and model version that produced it. Silent failures are only findable after the fact, and when a customer shows you a screenshot, you want to know exactly what your system produced. If you cannot reproduce it, you cannot fix it. If you cannot trace it, you won’t even know.

If your AI product returned fluent nonsense tomorrow, would your system notice before your customer did?

Originally published on LinkedIn

Want to apply this to your business?

If this sparked a useful question, let’s talk about where AI, automation, or product strategy can create practical leverage in your organisation.

Start the conversation