Business Strategy

Running an AI Pilot That Survives Go/No-Go

Running an AI Pilot That Survives Go/No-Go

AI pilot 90-day framework for Australian businesses

The Pilot That Killed the Programme

Most Australian AI programmes do not die in production. They die in pilot. The board approved something measurable. The pilot delivered something inconclusive. Three months later, the steering committee quietly stops meeting and the budget gets reallocated.

Gartner forecasts that over 40% of agentic AI projects will be cancelled by end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. MIT Sloan Management Review research on enterprise AI consistently shows that the majority of pilots never reach production, with structural design issues (not model performance) cited as the leading cause.

If you have already worked through the board-level business case and the payback maths, and you have read about the operational reality after go-live, this is the missing middle article. This is the operational playbook for the 90 days between go-decision and go/no-go review.

This is for the COO, Head of Transformation, Programme Manager, or CIO at an Australian business about to launch its first or next AI pilot. The thesis is uncomfortable but verifiable: the difference between pilots that survive go/no-go and pilots that quietly die comes down to design decisions made in the first seven days, not technology choices made over three months.


Part 1: The Four Modes of Pilot Failure

Before we cover what good looks like, here are the four failure modes that account for the majority of cancelled pilots. If your pilot is heading toward one of these, the model and vendor will not save it.

Four Modes of Pilot Failure

Metric
Failure Pattern
Root Cause and What Good Looks Like
Improvement
Inconclusive pilotCannot prove improvement or harm. Steering committee asks 'did it work?' and the answer is 'sort of'.Root cause: no clean before/after baseline. Good: 4 weeks of pre-pilot measurement on the same metric the AI will move.Defensible
Successful demo, failed scaleSandbox results brilliant, production results disappointing. Vendor blames data, business blames vendor.Root cause: pilot ran on cherry-picked data with no integration load. Good: pilot uses production data, real volumes, real exception rates from week one of go-live.Scalable
Adoption collapsePilot team loves it, broader team ignores it. Daily usage drops below 20% within 60 days of handover.Root cause: change management treated as training, not as job redesign. Good: cohort selected for adoption representativeness, not enthusiasm.Sticky
Compliance retrofitPilot worked, but privacy, Fair Work, or industry-regulator review surfaces issues that cannot be retrofitted.Root cause: Privacy Impact Assessment and security review skipped at pilot stage. Good: OAIC-style PIA completed before any production data touches the system.Defensible

The four failures share a single pattern: the pilot was designed to demonstrate enthusiasm, not to produce a defensible go/no-go decision. The next section is about designing the inverse.


Part 2: Five Design Decisions That Decide the Outcome

The five decisions below are made (or avoided) in the first seven days. Every successful AU pilot we have studied or supported nails these. Every failure misses at least two.

The 5 Design Decisions Made at Kickoff

Use case scoping
One process, one data source, one cohort. Narrow enough that any movement in the metric is attributable to the pilot.
Baseline definition
Measure the same metric for at least 4 weeks before the AI does anything. No baseline equals no defensible result.
Cohort selection
Representative of normal operating conditions, not enthusiasts. Statistically large enough that results generalise to the rest of the business.
Success criteria
Pre-committed by exec sponsor in writing. Specific (cycle time, error rate, adoption %), measurable, with a target and a window.
Kill criteria
Pre-committed conditions under which the pilot stops. Includes safety, compliance, cost overrun, and adoption floor. Respected when triggered.

The most common omission, by a long margin, is the kill criteria. Most pilots have an implicit "we will continue if it looks promising" stance, which means the project never gets killed even when the data says it should. NIST's AI Risk Management Framework is explicit on this: pre-committed go/no-go thresholds are a baseline governance requirement.

For the staffing decisions behind these roles, see our piece on the AI agent staffing gap in Australian businesses. For vendor scoping at this stage, the vendor selection question set is worth running before kickoff.


Part 3: A 90-Day Pilot Framework

The 90-day window earns its length. It is short enough to hold executive attention and long enough to capture a real operational rhythm including month-end. Compress it to 30 days and you are running a demo. Stretch it past 120 and the sponsor disengages, the goalposts move, and you end up in what Gartner calls "AI pilot purgatory".

The 90-Day Pilot Framework

1
Days 0-7
Kickoff and charter
Charter signed by exec sponsor. Governance committee membership confirmed. Baseline measurement of the target metric begins (will need to backdate 4 weeks of historical data). PIA scoping decision made. Kill criteria pre-committed in writing.
2
Days 8-30
Build and integrate
Vendor or platform configured. Production integrations stood up (not synthetic). Sandbox testing on real but de-identified data. PIA completed per OAIC guidance. Security review against the [50-point AI security checklist](/blog/ai-security-checklist-50-points-australian-businesses/). Fair Work consultation initiated where workforce impact exists.
3
Days 31-60
Parallel run
AI operates alongside the existing manual process. No decisions are taken solely on AI output. Daily metric capture, weekly stand-up with the cohort, fortnightly steering review. Adoption is observed, not assumed. Exceptions are logged with reason codes.
4
Days 61-75
Assessment
Compare 30 days of parallel-run data to the pre-pilot baseline. Measure adoption rate (daily active use, not logins ever). Calculate cost per transaction including model, integration, and oversight time. Quality audit on a stratified sample.
5
Days 76-90
Go/no-go decision
Executive readout with three options: scale, iterate, kill. Each option has a pre-committed budget envelope and decision owner. Kill is genuinely on the table; if it is not, the framework has already failed.

The single most important calendar item is the Day 0 baseline measurement. If you cannot answer "what was the cycle time, error rate, and cost per transaction in the four weeks before the pilot started" with a number drawn from your operational systems, the pilot will not produce a defensible go/no-go regardless of how well the AI performs.


Part 4: Metrics That Survive a Go/No-Go Review

Boards do not accept "users liked it". They accept measurable, attributable change against a baseline. Below are the metrics that survive a steering committee challenge, and the anti-metrics that do not.

Metrics That Survive vs Anti-Metrics That Do Not

Metric
Defensible Metric
How to Measure, Target, Evidence
Improvement
Cycle timeMedian minutes from input received to output produced.Pull from workflow system, compare pilot cohort to 4-week pre-pilot baseline on same cohort. Target: pre-committed % reduction.Survives
Error rate% of outputs requiring correction within 30 days.Stratified audit sample, same audit method as pre-pilot. Track both AI errors and human override errors.Survives
Adoption rateDaily active use across cohort, not 'logins ever'.System telemetry, 30-day rolling window. Pre-committed floor (e.g. >70% daily active by day 60).Survives
Cost per transactionFully loaded: model + integration + oversight time + exception handling.Finance reconciles against actual AP and time-tracking data. Compare to pre-pilot fully loaded cost.Survives
Quality / accuracyDomain-specific (e.g. classification accuracy, extraction F1).Held-out test set, audited by SME independent of vendor. Target pre-committed.Survives
Time-to-outputP50 and P95 latency at production load.Measured at peak hour, not average. P95 matters for SLAs.Survives

Anti-Metrics: Do NOT Count These as Success

Metric
What People Try to Count
Why It Does Not Survive Scrutiny
Improvement
Users said it was greatQualitative survey, often from the enthusiast cohort.Not attributable, not generalisable, vulnerable to social desirability bias.Reject
Pilot completed on timeProject management metric.Measures the team, not the outcome. Many on-time pilots produce inconclusive results.Reject
Vendor was responsiveRelationship metric.Confuses delivery quality with vendor sales effort. Useful for vendor selection, not go/no-go.Reject
Demo went wellStakeholder presentation outcome.Demos are curated. Production is not.Reject
Number of users onboardedCounts access, not use.Logins ever is not adoption. Daily active use is.Reject

Part 5: Five Anti-Patterns Killing AU Pilots

These are the five anti-patterns we see most often in Australian businesses. None of them are about the model.

Five Anti-Patterns Killing AU Pilots

Metric
Anti-Pattern
What Goes Wrong and How to Avoid It
Improvement
1. No baselinePilot starts measuring on day one of go-live. There is nothing to compare against.Run 4 weeks of pre-pilot measurement on the same metric, same cohort, same definitions. Without it, no result is defensible.Critical
2. Cohort too small or too biasedPilot runs with 3 enthusiasts. Results do not generalise to the wider team that will inherit the system.Select a cohort that mirrors operating reality: a mix of high, mid, and low performers across at least two locations or teams.Critical
3. Free vendor pilot with no exit clauseVendor offers a 'free' 90-day pilot. By day 90 the integration is so deep that switching costs lock you in regardless of results.Pre-commit exit clauses, data export rights, and a documented switching playbook. See our piece on [why DIY without understanding the lock-in costs more](/blog/ai-agents-australian-businesses-why-not-diy-without-understanding/).Critical
4. Pilot ran on synthetic dataSandbox results brilliant, real-world results poor because production data has edge cases, formatting variance, and PII the synthetic data did not.Use production data (de-identified where required for the PIA) from week one of build. Synthetic data is for vendor demos, not pilot evidence.Critical
5. No change managementAdoption stalls because the workflow change was not co-designed with the people who do the work.Follow a real change management process. See our piece on [AI change management and employee adoption](/blog/ai-change-management-employee-adoption/). Without it, even a technically successful pilot dies on handover.Critical

Part 6: Governance and Compliance at Pilot Stage

A pilot is not a regulatory exemption. The Australian Government's Department of Industry, Science and Resources (DISR) Voluntary AI Safety Standard, published in 2024, applies from pilot through to production. Treating pilot as a "we will deal with compliance at scale" phase is the most common reason a pilot cannot be scaled even when it works.

Pilot-Stage Governance Sequence

Privacy Impact Assessment
OAIC recommends a PIA for any project handling personal information at pilot stage, not just production. Completed by Day 21 of the framework.
Security review
Run the [50-point AI security checklist](/blog/ai-security-checklist-50-points-australian-businesses/) before any pilot touches production data. Document gaps and accepted risks.
Fair Work consultation
Any pilot that materially changes workflow for employees triggers consultation obligations. See [Fair Work compliance for AI automation](/blog/fair-work-compliance-ai-automation-guide/).
Audit logging from day one
DISR Voluntary AI Safety Standard requires traceability. Capture inputs, outputs, model version, and human override decisions from the first transaction.
Industry-specific overlay
Finserv: align with APRA CPS 230 and ASIC INFO 225 (see [financial services AI compliance](/blog/financial-services-ai-compliance-apra-asic/)). Healthcare: [TGA and patient-data overlay](/blog/healthcare-practice-ai-patient-automation/). Government suppliers: [DTA AI assurance](/blog/ai-government-contracts-compliance-requirements/). Aged care: [Aged Care Quality and Safety Commission alignment](/blog/aged-care-ai-automation-compliance-care/). ACCC consumer-facing: [Consumer Guarantees obligations](/blog/accc-consumer-guarantees-ai-implementation/).

For data flows that involve cross-border processing, the gap between Australian Privacy Principles and other regimes is meaningful and worth understanding at pilot stage rather than retrofitting later. Our explainer on Privacy Act versus GDPR covers the pilot-stage implications.


Part 7: Is Your Pilot Designed to Produce a Defensible Go/No-Go?

Five-minute self-assessment for the COO or programme manager. If you cannot answer "yes" to all five, the pilot will not produce a decision the board accepts.

Defensible Go/No-Go Readiness Check

Can your pilot survive a steering committee challenge?
Baseline measured for at least 4 weeks pre-pilot on same metric, same cohort
→ Yes: defensible direction of travel. No: result will be inconclusive regardless of AI performance.
Success criteria pre-committed in writing by executive sponsor
→ Yes: protects against goalpost movement. No: results will be re-interpreted to favour continuation.
Kill criteria pre-committed and the sponsor has confirmed they will respect them
→ Yes: real go/no-go is possible. No: pilot enters Gartner's 'AI pilot purgatory' indefinitely.
Cohort large enough and representative enough that results generalise
→ Yes: scale decision is defensible. No: scale will surface failures the pilot missed.
PIA and security review completed before any production data was touched
→ Yes: results can be scaled. No: compliance retrofit will block scale even if performance is good.

If three or fewer are "yes", do not run the pilot yet. Spend two weeks fixing the gaps. The cost of fixing them before kickoff is a fraction of the cost of an inconclusive 90-day pilot.


Part 8: The 12-Question Pre-Pilot Readiness Checklist

Use this as the agenda for the kickoff meeting. If any question cannot be answered, stop and resolve before proceeding.

  1. What single business metric will this pilot move? Not "improve customer service". Specific: cycle time, error rate, cost per transaction, conversion.
  2. What is the pre-pilot baseline value of that metric? Drawn from operational systems, not estimates.
  3. What is the pre-committed target value and the date by which it must be hit? Signed by the executive sponsor.
  4. What are the pre-committed kill criteria? Cost, safety, compliance, adoption floor. Each with a threshold and an owner.
  5. Who is the executive sponsor and what is their fortnightly time commitment? "As needed" is not an answer.
  6. Who is the cohort, how were they selected, and is the selection representative of operating reality?
  7. What production data will the pilot use, and what is the PIA status?
  8. What is the integration plan and what is the rollback plan if integration destabilises a production system?
  9. Who owns change management, what training has been designed, and what does the workflow look like post-pilot?
  10. What is the audit logging design and who reviews the logs weekly?
  11. What is the vendor exit clause and what is the data export design?
  12. What is the executive readout format for Day 90 and who attends?

Pilot Economics Summary

A typical AU AI pilot, costed honestly across 90 days:

Indicative 90-Day Pilot Budget (Typical Business)

Vendor / platform fees (90 days)$15,000 to $40,000
Internal time (sponsor, owner, SME, change champion)$25,000 to $60,000
Integration and data prep$10,000 to $30,000
PIA, security review, compliance work$8,000 to $20,000
Change management and training$5,000 to $15,000
Total 90-day pilot cost (indicative)$63,000 to $165,000

These ranges align with Hays and Robert Half 2025 Australian salary data for the contributing roles and with published vendor pricing for enterprise AI platforms. A pilot that costs less than this range is almost certainly underspending on the governance, change, and baseline-measurement work that determines whether the result will survive go/no-go. For the full payback maths post-decision, see our payback period calculator article.


What This Article Replaces

If you are running a pilot using a vendor-supplied template, that template is almost certainly optimised to demonstrate the vendor's product, not to produce a defensible business decision. The 90-day framework above replaces the vendor template with an Australian governance-aligned alternative.

For pilots in regulated sectors, layer the industry-specific overlay on top: APRA-aligned finserv, healthcare, government supplier requirements, accounting firm overlay, construction, aged care.

For the prior decisions that determine whether the pilot is the right shape in the first place: the board business case and the AI model selection guide and the AWS/Azure/GCP services overview.


Next Step

If your business is preparing to launch an AI pilot in the next quarter and you want a second pair of eyes on the design before kickoff, that is the most useful conversation we can have with you. We help Australian businesses design pilots that produce defensible go/no-go decisions rather than inconclusive demos. Our services in AI strategy and managed AI services cover the design-through-handover lifecycle.

Book a 30-minute pilot design review. No pitch, no slides. Walk us through your draft charter and we will tell you where the framework above flags risk.


Related Reading:


Sources: Gartner forecasts on agentic AI project cancellation (2024-2025), MIT Sloan Management Review enterprise AI research, McKinsey State of AI annual reports, DISR Voluntary AI Safety Standard 2024, OAIC Privacy Impact Assessment guidance, NIST AI Risk Management Framework, CSIRO Data61 and Australia's National AI Centre publications, Productivity Commission automation research, Hays and Robert Half 2025 Australia salary guides.