Industry Solutions

Why Industrial Analytics Projects Fail

Why Industrial Analytics Projects Fail

Data threads rising from a processing plant into a dashboard, several of them broken

The Model Was Right. The Production Report Won.

Most industrial analytics projects turn on a single moment, and it usually arrives about six weeks after the first dashboard goes live. Someone from the plant opens the new availability figure, compares it to the monthly production report, and finds they disagree by four per cent. The analytics team explains the difference. The explanation is correct. It changes nothing, because the production report is the number finance signs and the board reads, and a new tool that contradicts the signed number is a tool with a credibility problem rather than an insight.

From that week on, the dashboard gets opened less. Within a year it is a line item in a renewal conversation that nobody can defend.

This happens over and over in Australian mining, minerals processing, sugar, manufacturing and energy operations, and it almost never happens because the model was bad. The model is usually the most competent part of the whole exercise. The project fails somewhere underneath it, in the data layer that was assumed to be sorted and was not. Across 18 years of industrial data and operational technology work, including plant, production and reporting systems projects delivered for BHP, Rio Tinto and Senex Energy while employed by other consulting and engineering firms, the same five failure modes turn up with enough regularity to be worth naming.

The pattern is well documented outside industry too. Gartner's October 2024 survey of more than 3,100 technology executives found only 48 per cent of digital initiatives meet or exceed their business outcome targets, and BCG's 2020 study of 825 senior executives found 70 per cent of digital transformations fall short of their objectives. McKinsey's Global Industry 4.0 survey reported that as of late 2020 roughly 74 per cent of surveyed companies described themselves as stuck in pilot purgatory.

The cost is not abstract. An ABB survey of 3,215 plant maintenance decision makers, published in 2023, put unplanned downtime at AUD $349,000 an hour for a typical Australian industrial business, against a global average of AUD $194,000, with 69 per cent of respondents seeing unplanned outages at least monthly. Analytics that reached the supervisor making that call would be worth a great deal.

What follows is the view from underneath: why these projects stall, written by people who build the data layer rather than the slide deck.

Failure One: Building On Tags Nobody Reconciled

The most common failure is also the most avoidable. A team builds an analytics layer directly on historian tags, without first establishing that those tags produce a total anyone agrees with.

A historian tag is a raw measurement with a compression setting. Ask it for last month's tonnes and what you get back is an integral over a period, and that integral moves depending on the compression deadband, the interpolation method, how you handled the eight hours the instrument sat in calibration, and whether the query was time weighted. Two competent engineers querying the same tag with different summary settings will produce different tonnages, and both can defend their working.

A model built on top inherits every one of those choices silently. Its output looks authoritative because it arrives in a clean chart, and the first time anyone compares it to the reported figure, the gap is unexplainable to everyone who was not in the room when the query was written.

The production report wins that argument every time, and it should. It is reconciled against despatch, inventory movements and a physical stock count. The analytics layer is reconciled against nothing. Until the measurements underneath have been through a documented balance, an analytics project is building a second set of books, and the organisation already knows which set it trusts.

The fix is boring and it comes first. Establish one agreed figure per measurement, produced the same way every period, before anything is modelled on top of it. We have written the detail of this elsewhere, in a piece on why SCADA, the historian and the ERP disagree on tonnes and how to reconcile them. It is the least interesting work on the programme and it determines whether the rest survives.

Two Ways To Sequence The Same Programme

Metric
Analytics first
Reconciliation first
Improvement
Starting pointConnect to historian, start modellingAgree one number per measurement, then modelDefensible
First disagreementAnalytics layer loses to the signed reportBoth come from the same reconciled sourceNo conflict
Audit positionQuery logic lives in one person's notebookDocumented method, repeatable every periodAuditable
Typical outcomeDashboard retired in year twoFigure becomes the operational referenceDurable

Failure Two: No Asset Context, So Every Question Needs A Translator

The second failure is structural, and it kills the promise that gets these projects funded. Raw historian tags are a flat namespace of cryptic names. A tag called CV104_WT_PV means something to the control systems engineer who commissioned it and nothing to anyone else. There is no hierarchy in that name, no relationship to the circuit the conveyor sits in, and no way to ask "show me every weightometer on the secondary crushing circuit" without a human who already knows the answer.

So every analytics question routes through that person. A manager asks for a comparison across three lines, and the request queues behind whatever else the one person with the tag knowledge is doing. This is where self-service analytics dies, and it dies silently, because nobody files a defect saying "the promise did not happen". People just stop asking.

The remedy is an asset model: a maintained hierarchy that sits over the raw tags and gives them structure, units, equipment types and relationships. In a PI environment that means Asset Framework element templates built so they survive plant changes, which we have covered in detail in our note on designing PI Asset Framework templates that last. In an AVEVA or Citect estate the equivalent discipline lives in the equipment hierarchy and the naming standard.

ISA-95, published internationally as IEC 62264, describes enterprise and control system integration across levels running from sensors at the bottom to business planning at the top. Asset context belongs in the middle, at the manufacturing operations layer, underneath analytics and above the raw signals. Projects that skip it end up asking level four questions of level one data, and paying someone to bridge the gap by hand forever.

Failure Three: Nobody Owns The Model After Go-Live

An industrial analytics model describes a plant at a moment in time. The plant then changes. A screen is replaced, a circuit is rerouted, an instrument is moved, a new reason code is added because a new failure mode appeared, a line is debottlenecked and the nameplate rate moves.

Every one of those changes invalidates a small piece of the model, and none of them triggers a review, because the project closed and the team moved on. The drift is slow enough that nobody notices a single step, and after two years the analytics layer describes a plant that no longer exists. It still runs. It still produces charts. The charts are precise and wrong, which is worse than being visibly broken, because a visibly broken report gets fixed.

This is the same decay that hollows out delay accounting models, and for the same reason. Our guide to how Ampla delay accounting drifts away from the plant it was built for describes the mechanism in a system most Australian sites already run.

The answer is unglamorous and it separates a programme from a project. Somebody owns the model: a named person, an agreed response time, a change process that reaches the analytics model whenever the plant changes, and a periodic review against reality. Most organisations buy this as reactive support, which is why it yields no improvement. Our view of what an ongoing support and improvement arrangement should actually cover applies to analytics models as much as to production systems.

Failure Four: Scoped As Technology, When The Hard Part Is Definitions

Analytics projects get funded as technology projects, with a platform, an integration effort and a delivery date. The hard part is almost never technical.

Consider a business running three sites with three different histories of acquisition and commissioning. Ask what availability means. One site excludes planned maintenance from the denominator, another includes it, and the third excludes it only when the shutdown was in the annual plan. Ask what a delay is, and one site records anything over two minutes while another has a five minute threshold. Ask what a tonne is, and you get the wet figure at one site and the dry figure at the next.

None of that is a data problem. Every one of those definitions is defensible locally, and every one of them was chosen by someone competent for a reason that made sense at the time. Aggregate them without agreement and the group number is meaningless, which the site managers will work out faster than the project team does. The moment they do, each site goes back to its own figure and the consolidated view becomes decoration.

Getting agreement is slow, political work involving people who will lose an argument about their own numbers, and it decides whether a multi-site analytics investment returns anything. Scope it explicitly, give it time, and give it a decision maker. The alternative is finding out during user acceptance testing with no authority to settle it.

What Is Actually Wrong

Your analytics output is not being used. Which problem do you have?
Numbers disagree with the signed report
→ A reconciliation problem. Fix the measurement balance first.
Every question needs the one tag expert
→ An asset context problem. Build the hierarchy over the tags.
It was right at go-live, drifting now
→ An ownership problem. Nobody maintains the model.
Three sites will not accept the group figure
→ A definitions problem. Agree the terms before the platform.
Correct, agreed, and still nobody opens it
→ An audience problem. It was built for the wrong reader.

Failure Five: The Dashboard Built For The Steering Committee

The last failure is the one people find hardest to hear, because the artefact looks good. Most industrial dashboards are designed for the group that funded them. They are monthly, aggregated, colour coded against target, and built to answer the question a steering committee asks, which is whether things are broadly on track. That is a legitimate question to ask. Nobody acts on the answer.

The decisions that move an operation are made at six in the morning by a supervisor deciding which of three jobs the maintenance crew takes first, and at shift change by a metallurgist deciding whether to adjust a setpoint. Those people need a different artefact entirely: fewer numbers, current data rather than monthly, the specific equipment they are responsible for, and enough context to act without opening three other systems.

A dashboard that serves the committee and not the supervisor gets opened twelve times a year. A view that serves the supervisor gets opened twice a shift, and it generates the operational change that eventually shows up in the committee's monthly number anyway. Build the second one first.

The reporting stack matters here more than people expect. Pointing a reporting tool straight at a historian and querying millions of raw events is how teams accidentally build a slow second historian, and the result is too slow to open at 6am. Our note on getting historian data into Power BI without building a second historian covers the aggregation decisions that make an operational view fast enough to use.

Where AI And Machine Learning Genuinely Help

There is a real place for machine learning on operational data, and it is narrower and more useful than the marketing suggests.

Anomaly detection works when it is pointed at a named failure mode on a specific asset class. A model trained on vibration, temperature and load signatures across a population of similar pumps can flag a developing bearing fault earlier than a fixed threshold, because the signature is multivariate and a threshold is not. Success needs three things: a named failure mode, instrumentation that sees it, and enough historical examples of it happening. Where all three hold, the results are good. This is the ground covered in our piece on asset health programmes for Australian mining operations.

Classification is the other genuine win. Delay and downtime records arrive with a machine-measured duration and a human-supplied reason code, and the human half is where the quality problem lives. A model that suggests the likely reason code from the process conditions around the event, and lets the operator accept or override, improves coding consistency without taking the judgement away. We have written about where a trained model beats hand-written delay classification rules and where the rules still win.

Alarm rationalisation is the third. Alarm floods are a clustering and analysis problem over historical alarm records, and the work of finding chattering alarms, duplicate alarms and standing alarms responds well to analysis at scale. The standards work here already exists. ANSI/ISA-18.2, first published in 2009 and revised in 2016 and adopted internationally as IEC 62682, covers the management of alarm systems for the process industries across a full lifecycle from philosophy and rationalisation through to monitoring and audit. The EEMUA 191 guidance most Australian process operators work to sets the numbers: in steady state an average below one alarm per operator per ten minutes is likely acceptable, more than one a minute is likely unacceptable, and an alarm flood is more than ten new alarms in any ten minute window. The analysis tells you which alarms to rationalise. The standard tells you what good looks like when you are done.

Where They Do Not

Anything that needs a deterministic, auditable answer should not have a model in the middle of it. Production figures that feed a royalty calculation, emissions data that feeds a Safeguard Mechanism obligation, metal accounting that feeds a financial close: these need a documented method that produces the same answer from the same inputs every time, and that a regulator or an auditor can follow. A model that is right ninety five per cent of the time is the wrong tool for a number someone signs.

The second exclusion is broader. Machine learning on data that does not reconcile produces reconciled-looking nonsense, confidently and at speed. If the tonnes do not agree across systems, a model trained on those tonnes has learned the disagreement. Fix the measurement layer first, every time.

The third gets overlooked. Where the physics is known and a first-principles calculation exists, use the calculation. A mass balance beats a regression that approximates one: it is explainable, it extrapolates correctly outside the training range, and nobody has to defend it in an audit.

What Good Looks Like: One Decision, Worked Backwards

The projects that work invert the usual order. Instead of connecting to a data source and looking for value, they start with one decision somebody makes repeatedly and work backwards to the data that decision needs.

Working Back From The Decision

Name the decision
One recurring call, one named person, one time of day
List what they need
The few numbers that would change the call
Trace each to source
Which system, which tag, which period, summarised how
Reconcile first
Agree one figure per measurement before modelling
Build the narrow view
For that person, at that moment, nothing extra
Assign the owner
Named maintainer and a change process from day one

Naming the decision is harder than it sounds and it is where the value gets decided. "Improve reliability" is not a decision. "At the Monday planning meeting, choose which four of the eleven flagged work orders go into this week's window" is one, and you can work out exactly which numbers would change it.

Tracing to source is where the unpleasant surprises arrive, and week two is a better place to meet them than month six. It is common to find that of the four numbers a decision needs, one does not exist in any system, one is manually entered and unreliable, one exists in two systems with different values, and one is fine. That finding usually reshapes the scope, and it is the most valuable thing the project produces early.

Building narrow is a discipline in its own right. A view that answers one decision for one role can be delivered in weeks and judged honestly, because the test is whether the decision changed. A platform that could theoretically answer anything takes a year and can never be judged at all.

A Sequence That Holds Up

1
Weeks 1-2
Decision and source review
Name the recurring decision, trace every number it needs back to a system and a tag
2
Weeks 3-6
Reconcile the measurements
Agree one figure per measurement, documented and repeatable each period
3
Weeks 5-8
Asset context
Build the hierarchy and templates so questions can be asked without a translator
4
Weeks 9-12
The narrow view
Deliver for the one role, at the one moment, and test whether the decision changed
5
Ongoing
Ownership
Named maintainer, change process tied to plant changes, periodic review against reality

The overlap between reconciliation and asset context in weeks five to eight is deliberate. Working out which tag is the trusted source for a measurement is most of the work of deciding what the asset model should expose.

What The Sequence Changes

Time to first defensible numberWeeks, not quarters
Arguments about whose figure is rightSettled once, documented
Analytics questions needing the tag expertMost of them removed
Model still describing the plant in year threeOnly if someone owns it

Where To Start

If you have an analytics layer nobody opens, do not start by replacing it. Work out which of the five failures you have first, because they need different fixes and the wrong fix is expensive. A week of honest review settles it: compare the analytics output to the signed report, count how many questions last quarter needed the one tag expert, and check when the model was last reviewed against the plant.

If you are about to start one, spend the first fortnight on the decision and the sources rather than the platform. The platform matters far less than the industry assumes, and it is much easier to change later than a data layer built on unreconciled measurements.

This sits underneath the operational technology, analytics and production systems capability we build for asset-heavy Australian operators, alongside the system integration work needed to get plant data into enterprise reporting at all. For resources and processing businesses running these programmes from head office, our Brisbane practice covers the same ground. Where analytics has to support several sites at once, the data conditions get harder again, and our companion piece on what a remote operations centre needs from the data layer sets out what has to be true first.

If your dashboards are technically correct and commercially ignored, start a scoping conversation and we will review the data layer underneath them before anyone proposes a rebuild.

Common Questions

Why do industrial analytics projects fail so often?

Rarely because of the model. The common causes sit underneath it: analytics built on historian tags that were never reconciled, so the output disagrees with the signed production report; no asset context over the raw tags, so every question needs a human translator; nobody owning the model after go-live, so it drifts from the plant; and definitions of availability or tonnes that differ across sites.

What does it mean to reconcile data before building analytics?

It means establishing one agreed figure per measurement, produced the same way every reporting period, before anything is modelled on top of it. A historian total depends on the compression deadband, the interpolation method, how calibration gaps were handled and whether the query was time weighted. Two engineers can query the same tag and get different answers, and both can defend the working.

Why does asset context matter more than the analytics platform?

Because raw historian tags are a flat namespace of cryptic names with no hierarchy and no relationships. Without a maintained asset model over the top, nobody can ask for every weightometer on a circuit without a person who already knows the tag names. That is the point where self-service analytics stops, and it stops silently because people just stop asking rather than raising a defect.

Where does AI genuinely help with process data?

Three places hold up well. Anomaly detection on a named failure mode where you have instrumentation that sees it and enough historical examples. Classification, such as suggesting a delay reason code from process conditions and letting the operator override. And alarm rationalisation, where clustering over historical alarm records finds chattering, duplicate and standing alarms for review against ISA-18.2.

When should you not use machine learning on operational data?

Whenever the answer has to be deterministic and auditable. Production figures feeding a royalty calculation, emissions data feeding a Safeguard Mechanism obligation and metal accounting feeding a financial close all need a documented method that gives the same answer from the same inputs. Also avoid it where the underlying data does not reconcile, because the model simply learns the disagreement.

How should an industrial analytics project be scoped?

Start with one decision somebody makes repeatedly, name the person and the moment they make it, then work backwards to the few numbers that would change that call. Trace each number to a system, a tag and a summary method. Reconcile those measurements, build the narrow view for that one role, and assign a named owner with a change process before go-live rather than after.


Related Reading