Industry Solutions

Alarm Rationalisation from Historian Data

Alarm Rationalisation from Historian Data

Rationalising a control room alarm system from historian and alarm log data

The Operator Who Stopped Reading the Alarms

Consider a typical Australian processing plant on night shift. A single board operator watches three screens. Over an eight hour shift the alarm summary logs several thousand entries. Most are the same twenty tags cycling in and out of alarm, a level that sits one percent above a poorly set threshold, a pump that trips its high vibration alarm every time it starts. The operator learned months ago that acknowledging them all is the job, and reading them all is impossible. So the alarms get acknowledged in bulk, and the one alarm that mattered, the one that was different, scrolled past at 2am inside a flood of noise.

This is the failure mode behind a surprising number of incident reports. The alarm system was supposed to be the last engineered layer of protection before something goes wrong, and it has been demoted to background noise over time because nobody had the time to fix it. If you run a plant, a water utility, a smelter, a mine site or any continuous process in Australia, the alarm log on your historian is probably telling you this right now, and nobody is reading it either.

Alarm rationalisation is the discipline that fixes it. Most sites never finish one, and the reason is rarely that the method is unclear. The traditional method, a room full of engineers reviewing every alarm by hand, takes so long that it stalls. The evidence needed to do it faster is already sitting in the historian.

What Rationalisation Actually Means

Rationalisation is the review of every configured alarm against a documented philosophy, to confirm that each one is valid, has a defined operator response, is set at the right threshold and priority, and earns its place on the screen. An alarm with no action an operator can take is not an alarm. It is a distraction wearing an alarm's clothing.

The reference documents are real and worth naming precisely. EEMUA Publication 191, Alarm Systems: A Guide to Design, Management and Procurement, is the guideline most Australian sites benchmark against. The formal standard is ANSI/ISA-18.2, Management of Alarm Systems for the Process Industries, and its international equivalent is IEC 62682, aligned with both. These documents agree on the shape of a healthy alarm system and on the benchmarks that tell you whether yours is one.

The benchmarks are specific. EEMUA 191 treats an average of around one alarm every ten minutes, roughly 144 a day per operator position, as the upper bound of what a person can reliably manage in steady state, and high performing sites run well below it. It defines an alarm flood as more than ten alarms in ten minutes for a single operator, a rate at which no one can assess and act on each one. Against those numbers, the night shift above is not a busy plant. It is a plant whose alarm system has failed its basic design intent.

EEMUA 191 Alarm Performance Benchmarks

Manageable average alarm rate (steady state)~150/day
Target for a well-performing systemFewer per shift
Alarm flood threshold (one operator)>10 in 10 min
Priority distribution most sites aim forLow / med / high, weighted to low

These are published benchmarks, not a savings claim. What they give you is a target, and a way to measure the gap between where your alarm system is and where the guidance says it should be. Most sites that have never rationalised are many multiples over the manageable rate, and they do not know by how much because nobody has counted.

Why the Workshop Method Stalls

The textbook approach to rationalisation is a series of workshops. The alarm list is exported, and a team of process engineers, operators and a facilitator works down it row by row, deciding for each alarm whether it stays, what its priority should be and what response it demands. On a small skid this works. On a plant with fifteen thousand configured alarms it does not, because the arithmetic defeats it. At a realistic review rate of a few dozen alarms an hour, the full list is months of workshop time that the same engineers are supposed to find alongside their day jobs.

So the effort starts, covers the first few hundred alarms, and stops. The document that results describes an alarm philosophy that was never applied to most of the system. Meanwhile the plant changed, new alarms were added at commissioning of the last project, and the gap between the philosophy and the live configuration widened again. The discipline is sound. The method does not scale to the size of the problem.

The insight that unlocks it is that you do not have to review every alarm with equal effort. A small fraction of the tags generate the overwhelming majority of the alarm load, and the historian already knows which ones they are. Fix those first, with evidence, and the operator's screen gets quiet enough that the genuine review of the rest becomes possible.

The Evidence Is Already in Your Historian

Every alarm and event that has fired is recorded. Whether the system is OSIsoft PI with its event frames, an AVEVA Historian, a Citect alarm log or a DCS journal, the alarm history is a dataset with a timestamp, a tag, a state transition and usually a priority. Analysed properly it answers the questions a workshop guesses at.

It tells you which tags chatter, cycling in and out of alarm many times an hour because their threshold sits inside the normal process noise band. It tells you which alarms stand, sitting in the active-unacknowledged state for hours or days because the operator has learned to leave them there. It tells you which alarms are stale, fired once and never cleared. It tells you which fire together, the correlated groups that mean one process upset is generating twenty alarms. And it tells you, minute by minute, when the floods happen and what triggers them.

What the Alarm History Reveals

Metric
Alarm pathology
What the data shows
Improvement
ChatteringThreshold sits in the process noise bandSame tag, many transitions per hourDeadband / delay
Standing / staleOperator leaves it active permanentlyLong active-unacknowledged durationRe-set or remove
DuplicateSeveral alarms for one conditionTags that always fire togetherConsolidate
FloodOne upset floods the screen>10 in 10 min, correlatedSuppress on cause
Nuisance priorityEverything set to high at commissioningPriority distribution skewed highRe-prioritise

The first pass of an evidence-led rationalisation is not a workshop. It is a query. Rank every tag by the alarm load it has generated over a representative period, say ninety days, and the top of that list, usually a few dozen tags, is where most of the operator's noise comes from. That is where the review starts, and it is the fastest path to a screen an operator can actually read.

Where AI Helps, and Where It Does Not

There is a real place for machine learning here, and there is a lot of vendor noise pretending it is bigger than it is. The honest split matters, because an alarm system is a safety function and pretending a model can own decisions it cannot is how you make things worse.

Pattern detection over the alarm log is where analytics earns its place. Clustering correlated alarms to surface the groups that fire together is a genuine data problem that a person scrolling a log will miss. Detecting that a tag has started chattering more than it used to, flagging drift in alarm behaviour after a process change, and reading operators' free-text shift comments to connect a recurring alarm to what was happening on the floor, these are all things a model does well and a manual review does slowly or not at all. The comment field, as anyone who has worked with plant data knows, is often the most honest record of what the alarm system is actually doing.

What stays with human judgement is every decision that changes the alarm itself. Whether an alarm is safety related, what its priority should be, what threshold protects the equipment, whether it can be suppressed and under what condition, these are engineering and risk decisions with consequences. A model can rank the candidates and show the evidence. An engineer, against the site's alarm philosophy, makes the call. The same principle applies to any AI agent working over operational data: the value is in surfacing and prioritising, the boundary is a read-only one, and a human stays in the loop on anything that touches plant state.

Analytics or Engineer?

What is the task in front of you?
Find the tags generating the most load
→ Analytics: rank from the log
Group alarms that fire together
→ Analytics: correlation clustering
Set a threshold or priority
→ Engineer: against the philosophy
Decide if an alarm is safety related
→ Engineer: risk assessment

Getting this boundary right is the difference between an analytics layer that helps and one that introduces risk nobody sees. It is the same conversation Australian boards are starting to have about where AI advisory and oversight belong in an operational business, and the answer on a plant floor is more conservative than in a back office, because the failure modes are physical.

The Workflow That Scales

Put together, an evidence-led rationalisation follows a repeatable path that starts from data and ends with a maintained system rather than a shelved document.

Evidence-Led Rationalisation

Extract
Pull the full alarm and event history from the historian
Rank
Order tags by alarm load, find floods and correlations
Triage
Fix chattering and nuisance alarms first, with evidence
Review
Rationalise the rest against the alarm philosophy
Monitor
Track KPIs against EEMUA 191 every period, not once

The last step is the one that makes it stick. A rationalisation that happens once is a snapshot that decays. The plant changes, projects add alarms, and within a year the system drifts back toward noise. The maintainable version publishes the alarm KPIs on a schedule, the same way the plant already reports production, so that a rising average rate or a new chattering tag shows up as a number someone owns rather than as an incident eighteen months later. That reporting layer is a natural extension of getting historian data into Power BI for the rest of the operation.

A Realistic Sequence

A rationalisation done this way is measured in weeks for a first result and in an ongoing cadence after that, not in a single heroic project.

A Staged Rationalisation

1
Weeks 1-2
Baseline
Extract history, benchmark against EEMUA 191, quantify the gap
2
Weeks 3-5
Quick wins
Fix the top load generators: deadbands, delays, duplicates
3
Weeks 6-12
Full review
Rationalise remaining alarms against the philosophy
4
Ongoing
Sustain
Publish KPIs each period, manage change on alarms

The baseline alone is worth doing on its own. It converts a vague sense that there are too many alarms into a defensible number, benchmarked against a recognised guideline, that tells you how far off you are and where the load concentrates. That number is what turns alarm management from a complaint into a funded piece of work.

Why This Matters Beyond the Control Room

For operators of critical infrastructure the alarm system carries a regulatory weight alongside the operational one. Under the Security of Critical Infrastructure Act and the risk management program obligations that flow from it, operators are expected to demonstrate that the controls protecting essential services are effective and maintained. An alarm system that floods and gets acknowledged in bulk is a control that is not doing its job, and the evidence of that is sitting in the same logs a regulator or an auditor could ask to see. Reading those logs before someone else does is part of taking the CIRMP obligations under the SOCI Act seriously.

For mine sites the same evidence supports principal hazard management. Where an alarm is a control in a principal hazard management plan, being able to show that the alarm fires when it should, is not buried in noise and gets acted on is exactly the kind of control effectiveness evidence those plans are supposed to carry. The alarm history is the record that either supports that claim or undermines it.

How to Start Without Buying Anything

The first move is an extract and a baseline, ahead of any software purchase. Pull ninety days of alarm and event history off the historian, rank the load, count the floods and compare the result against the EEMUA 191 benchmarks. That analysis tells you whether you have a tuning problem in a handful of tags, a philosophy problem across the whole system, or a change-management problem where every project adds a little more noise. Those need different fixes, and the baseline is what tells them apart before anyone commits to a build.

Solve8 works with Australian operators to turn alarm and historian data into a rationalisation that actually finishes, drawing on 18 years of hands-on work with SCADA, historians and plant data on major Australian industrial sites, gained while working with previous consulting and engineering employers. If your operators have stopped reading the alarms, start a scoping conversation and we will baseline your alarm system against EEMUA 191 before anyone proposes a tool. You can see how this connects to the rest of an operation's data on the Solve8 home page and across our system integration work.

Alarm rationalisation sits inside the wider picture of getting plant data to a state where it can be trusted, which is what our operational technology and mining consulting work covers.

Common Questions

What is alarm rationalisation?

Alarm rationalisation is the structured review of every configured alarm against a documented alarm philosophy, to confirm each one is valid, set at the right threshold and priority, and has a defined operator response. Any alarm with no action an operator can take is removed or re-engineered. The recognised references are EEMUA Publication 191 and the ANSI/ISA-18.2 standard, with IEC 62682 as the international equivalent.

How many alarms per hour is too many?

EEMUA 191 treats an average of around one alarm every ten minutes, roughly 150 a day per operator position, as the upper limit of what a person can reliably manage in steady state, and well-performing systems run below it. It defines an alarm flood as more than ten alarms in ten minutes for one operator, a rate at which the operator cannot assess and act on each alarm. Most sites that have never rationalised run many multiples above the manageable rate.

Can AI do alarm rationalisation automatically?

No, and any vendor claiming it can is describing a risk. Machine learning is genuinely useful for ranking alarm load, clustering correlated alarms and detecting drift, which are data problems a manual review does slowly. The decisions that change an alarm, its priority, threshold, whether it is safety related and whether it can be suppressed, are engineering and risk judgements that stay with a person, made against the site's alarm philosophy. Analytics surfaces and prioritises. The engineer decides.

Where does the data for rationalisation come from?

From the alarm and event history already stored in your historian or DCS, whether that is OSIsoft PI event frames, an AVEVA Historian, a Citect alarm log or a DCS journal. Each record has a timestamp, a tag, a state transition and usually a priority. Ninety days of that history is enough to rank the load, find the floods and identify the chattering, standing and duplicate alarms that generate most of the noise.

How long does an alarm rationalisation take?

A first useful result, the baseline against EEMUA 191, takes a week or two. Fixing the top load generators takes a few more weeks and quiets the operator's screen the most. The full review of the remaining alarms runs over the following weeks depending on system size, and then it becomes an ongoing cadence: publishing alarm KPIs each period and managing change so the system does not drift back into noise.


Related Reading