Business Strategy

Agentic AI Needs a Real Data Foundation

Agentic AI Needs a Real Data Foundation

Building the data foundation an agentic AI system needs to work on operational data

The Pilot That Worked, Until It Had To Be Trusted

Consider a common situation in an Australian industrial business. A team builds an agentic AI proof of concept, an assistant that can answer questions about production, pull a figure from a report, draft a shift summary. In the demo it is impressive. It reads the data, reasons about it, produces a sensible answer, and the room agrees this changes how the plant will run. Then the question comes that always comes: can we trust it in production, on real decisions, with real money and real safety on the line. And the honest answer, months later, is no. The model reasoned well enough. The data underneath it could not carry the weight.

This is the most common way agentic AI projects end, and the numbers say it is about to get more common. Gartner has forecast that over 40 percent of agentic AI projects will be cancelled by the end of 2027, driven by rising costs, unclear value and inadequate controls. A separate Gartner projection holds that through 2026 around 60 percent of AI projects will be abandoned for a lack of AI-ready data. Read together, they describe a single problem. The intelligence is not the constraint. The foundation it stands on is.

For an Australian business weighing what an AI strategy engagement should actually buy, this is the most useful thing to understand before spending anything. The work that decides whether agents succeed is the work most vendors skip past to get to the demo.

Why the Foundation Fails, Not the Model

The models available today are more than capable of the tasks most businesses want to point them at. When an agentic system fails in production it is almost never because the model could not reason. It is because one of a small set of foundation problems made the reasoning unsafe to rely on.

The first is context. A model handed a number with no idea what it measures, what units it is in, when it was recorded or whether the sensor was healthy at the time will confidently use it wrong. On a plant, a tag called FIC_101.PV means nothing without the asset context that says it is the feed flow to mill two, in tonnes per hour, valid only when the mill is running. A person knows this. An agent knows only what the data model tells it, and on most sites the data model does not tell it.

The second is quality. Agents act on what they read. If the underlying data is stale, contradictory across systems or full of gaps that were quietly backfilled, the agent inherits every one of those flaws and acts on them at speed. The same three-systems-disagree problem that produces SCADA, historian and ERP numbers that do not reconcile becomes far more dangerous when an autonomous agent is making decisions off whichever number it happened to read first.

The third is access and boundary. An agent that can only answer questions is useful and low risk. An agent that can act needs a hard boundary defining what it may touch, and on operational systems that boundary is a safety requirement. The gap between a read-only advisory agent and one that can change plant state is the gap between a helpful tool and an incident waiting to happen.

Why the Demo Works and Production Does Not

Metric
In the demo
In production
Improvement
Data scopeA clean, curated sampleEverything, including the messy tagsGovernance
ContextExplained by the person demoingMust live in the data modelSemantics
QualityKnown-good period chosenGaps, drift, contradictionsQuality gates
FreshnessStatic snapshotLive data, measured in minutesPipelines
ActionsRead-only, watchedMay act, unwatchedHard boundary

None of these is a model problem, and none of them is fixed by choosing a different model. They are data and governance problems, and they are the reason the demo and the production system behave so differently.

What AI-Ready Data Means for Operations

The phrase AI-ready data gets used loosely. For an operational business it has a concrete meaning, and it is more demanding than the reporting-grade data most sites already have.

Reporting data is allowed to be a day or a week old, reconciled once a month, and understood by the handful of people who work with it. Data that an agent acts on has to be governed at the level of the individual asset, carry its own meaning so that context does not live only in someone's head, and be quality-assured continuously rather than at a monthly close. Where traditional data management runs at a reporting cadence, data feeding a production agent needs its quality signals measured in hours or minutes, because the agent is acting on that cadence too.

On a plant, four things carry most of that weight. Tag naming and an asset model, so that every signal knows what it is and what it belongs to. This is the same discipline that makes PI Asset Framework templates valuable: context modelled once, applied everywhere, rather than re-explained for every question. Data quality signals, so the agent knows when a sensor is unhealthy and can decline to act on it. Lineage, so any answer can be traced back to the measurement it came from, which is what makes an agent's output auditable. And a defined access boundary, so what the agent may read and what it may never touch is written down and enforced.

Is Your Data Ready for an Agent?

Which of these can you answer today?
Every tag knows its asset, units and validity
→ Context: ready
You can tell when a signal is bad in real time
→ Quality: ready
Any figure traces to its source measurement
→ Lineage: ready
None of these reliably
→ Fix the foundation first

If the honest answer to most of these is no, the foundation is the next step and the agent waits behind it. Deploying an agent onto data that cannot answer these questions is how a project joins the 40 percent that get cancelled, and it usually joins them after the money is spent.

The Read-Only Boundary Is Not Optional

There is a strong case for keeping the first generation of operational agents strictly read-only, and it is worth stating plainly because the pressure to let agents act arrives quickly once they prove useful.

An agent that reads plant data, reasons about it and surfaces what a human should look at delivers most of the value with almost none of the risk. It can spot the correlation a person missed, draft the shift report from the data, flag that a figure looks wrong before it reaches the board. Every one of those is advisory, and a person acts on it. The moment an agent can change a setpoint, acknowledge an alarm or write back to a control system, the risk profile changes completely, and the data foundation and the governance around it have to be far stronger before that is defensible.

This is the same principle that sits under AI agent governance and human override, applied to a setting where the failure mode is physical rather than reputational. On a plant, the conservative default is the correct engineering judgement, because the cost of a confident wrong action is measured in equipment, production and, at the extreme, safety. Start read-only, prove the foundation, and widen the boundary deliberately rather than by drift.

What to Fix First

The sequence that gets a business from a promising demo to an agent it can trust is a data and governance sequence, not a model-selection one.

From Demo to Trusted Agent

Scope
Pick one job the agent should do, and its data
Model
Give the data context: assets, units, validity
Assure
Add quality signals and lineage the agent can read
Bound
Define read-only access and the human override
Deploy
Run advisory first, measure, then widen carefully

The discipline that makes this work is starting narrow. A single, well-defined job on well-understood data, with a clear read-only boundary, is the version that survives contact with production. The all-purpose plant assistant that can answer anything is the version that impresses in the demo and gets cancelled a year later. Scope decides the outcome more than model choice does, and scoping it well is where a good system integration approach earns its place, because the hard part is connecting the agent to governed data rather than the reasoning itself.

A Realistic Sequence to Data-Ready

Getting operational data to the point where an agent can safely act on it is measured in months, and most of the value arrives before the agent does, because governed, contextualised, quality-assured data improves every report and dashboard the business already runs.

Getting Data-Ready for Agents

1
Weeks 1-3
Assess
Pick the target job, audit the data it needs against readiness
2
Weeks 4-8
Model
Build the asset context, units and validity into the data
3
Weeks 9-14
Assure
Add quality gates, lineage and the access boundary
4
Ongoing
Deploy
Advisory agent first, measured, boundary widened deliberately

The assessment on its own is worth doing before committing to an agent at all. It tells you whether your data can carry an autonomous system or whether you are one governance project away from being able to try, and it does so before the budget is spent on a pilot that was never going to reach production.

The Numbers Behind the Strategy

The published forecasts are useful as a planning input, not as a scare tactic. They quantify a bet that most businesses are making blind.

What the Research Says About Readiness

Agentic AI projects Gartner expects cancelled by 2027>40%
AI projects abandoned through 2026 for lack of AI-ready data~60%
The variable that separates themThe data foundation

The strategic conclusion is to spend the first portion of any agent budget on the foundation rather than avoiding agentic AI altogether, because that spend is what moves a project from the group that gets cancelled to the group that scales. It is also spend that pays for itself independently of the agent, since the same governed data makes existing reporting and decision-making better. This is the argument Australian boards are increasingly hearing when they ask what AI advisory should tell them before they fund a programme, and it is why the readiness question sits ahead of the tooling question in any serious plan.

How to Start

The first step is not to select a model or a platform. It is to pick one job worth automating and audit the data that job depends on against a readiness standard: does every signal carry its context, can you tell in real time when it is wrong, and can any answer be traced to its source. That audit tells you whether you have an agent opportunity now, a governance project to do first, or a reporting problem masquerading as an AI problem. They need different work, and knowing which one you have is what stops you spending a pilot budget on a project that could not have reached production.

Solve8 helps Australian businesses build the data foundation an agentic system actually needs, drawing on 18 years of hands-on work with industrial data, historians and enterprise systems on major Australian sites, gained while working with previous consulting employers. If you have a promising agent demo and a nagging doubt about production, start a scoping conversation and we will assess your data against what an agent requires before anyone proposes a build. You can see how this fits the wider picture on the Solve8 home page.

Common Questions

Why do most agentic AI projects fail?

Rarely the model itself. They fail on the data foundation underneath it: signals with no context, quality problems the agent inherits and acts on, and the absence of a defined access boundary. Gartner forecasts that over 40 percent of agentic AI projects will be cancelled by the end of 2027, and separately that around 60 percent of AI projects will be abandoned through 2026 for a lack of AI-ready data. The intelligence is rarely the constraint. The foundation is.

What does AI-ready data mean for an industrial business?

It means data governed at the level of the individual asset, carrying its own context so meaning does not live only in someone's head, quality-assured continuously rather than at a monthly close, and traceable back to its source. Reporting-grade data can be a week old and understood by a few people. Data an agent acts on needs its context and quality signals available in minutes, because the agent is acting on that cadence.

Should an AI agent be allowed to control plant equipment?

Not as a starting point. A read-only advisory agent that reads operational data, reasons about it and surfaces what a human should act on delivers most of the value with little of the risk. An agent that can change a setpoint, acknowledge an alarm or write to a control system carries a physical failure mode, and the data foundation and governance have to be far stronger before that is defensible. Start read-only, prove the foundation, widen the boundary deliberately.

How is agent-ready data different from data we already report on?

Reporting data is allowed to be periodic, reconciled monthly and understood by specialists. Agent-ready data has to be contextualised so every signal knows what it is, quality-assured in near real time so the agent knows when not to trust a signal, and lineage-tracked so its answers are auditable. Most businesses have reporting-grade data and assume it is agent-ready. The gap between the two is where the pilots fail.

What should we do before starting an agentic AI project?

Pick one specific job worth automating and audit the data that job depends on against a readiness standard: context, real-time quality, and lineage. That audit tells you whether you have an agent opportunity now, a governance project to do first, or a reporting problem in disguise. Doing it before committing to a pilot is what stops a project spending its budget and then discovering it could never have reached production.


Related Reading