What survives the third week of an industrial AI pilot

A pattern is clear across our deployments: PoCs that look great on day one and disintegrate on day twenty share the same failure modes. We've mapped seven of them, and what to do instead.

Process plant at dusk, where industrial AI pilots meet real operating conditions
7 failure modes mapped

Industrial AI pilots do not usually fail at the demo. They fail somewhere around day twenty, once the novelty has worn off and the system has to survive a shift change, a maintenance window, and a week where nobody has time to look at it. We have run enough of these to see the same seven failure modes recur, almost independently of the use case.

1. The pilot was scoped around the model, not the workflow

A model that is 94% accurate and fires into nobody's inbox is worth nothing. The pilots that survive define, before kickoff, who receives the output, what they are expected to do with it, and what happens when they disagree. That is a workflow question, and it is usually answered too late.

2. Nobody owns the exceptions

Every system produces cases it cannot decide. If there is no named person for those, they pile up silently and the backlog becomes the reason the pilot is judged a failure. Assign the reviewer on day one, and measure their queue depth as a first-class metric.

3. The data was cleaner in the sample than in production

Sample extracts are almost always curated, often unintentionally. The tag list is the one that was maintained; the drawings are the ones that were digitised. Ask explicitly for the messiest available slice and build against that.

4. Second-shift conditions were never tested

Lighting changes, different operators, different line speeds, different supervisors. A vision pilot validated between 10am and 4pm on a Tuesday has tested roughly a third of the conditions it will actually meet.

5. Success was never defined numerically

"Better visibility" cannot be evaluated. A pilot needs a baseline measured before anything is installed and a target expressed in the same unit. Without the baseline, the pilot ends in a debate about whether it worked.

6. It was integrated as a demo, not as a system

Screens that read from a one-off extract are a prototype. If the pilot does not read live from the systems of record, everything learned during it has to be rebuilt for production, which is where the timeline quietly doubles.

7. The plant was never given a reason to keep it

The people who make a pilot succeed are usually not the people who bought it. If the system does not save the operator, the engineer, or the supervisor time in their own week, it will be tolerated for the duration of the pilot and then quietly abandoned.

What we do differently now

  1. Baseline first. Nothing is installed until the current number is measured and agreed.
  2. Name the reviewer and the escalation path before the first model runs.
  3. Read from the systems of record from week one, even if the model is trivial at that stage.
  4. Run through at least one full shift cycle, including night shift, before any go/no-go review.
  5. Report the exception queue alongside the accuracy figure, every week.

None of this makes the model better. It makes the difference between a model that gets used and one that gets a polite write-up.

Written from deployment experience across multiple sites. Patterns are described without identifying any customer — we are under confidentiality with the operators involved.

Talk to an engineer about this
Now onboarding new deployments

Bring one problem. We'll show you the deployment.

A 30-minute working session with our engineering team. No slideware. Bring a real operational problem and leave with a concrete view of what a PredCo deployment looks like for your plant.

4 hrs
Average response time
4–6 wks
Kickoff to live pilot
On-prem
Sovereign deployment supported