Running an AI Pilot That Survives Go/No-Go

The Pilot That Killed the Programme
Most Australian AI programmes do not die in production. They die in pilot. The board approved something measurable. The pilot delivered something inconclusive. Three months later, the steering committee quietly stops meeting and the budget gets reallocated.
Gartner forecasts that over 40% of agentic AI projects will be cancelled by end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. MIT Sloan Management Review research on enterprise AI consistently shows that the majority of pilots never reach production, with structural design issues (not model performance) cited as the leading cause.
If you have already worked through the board-level business case and the payback maths, and you have read about the operational reality after go-live, this is the missing middle article. This is the operational playbook for the 90 days between go-decision and go/no-go review.
This is for the COO, Head of Transformation, Programme Manager, or CIO at an Australian business about to launch its first or next AI pilot. The thesis is uncomfortable but verifiable: the difference between pilots that survive go/no-go and pilots that quietly die comes down to design decisions made in the first seven days, not technology choices made over three months.
Part 1: The Four Modes of Pilot Failure
Before we cover what good looks like, here are the four failure modes that account for the majority of cancelled pilots. If your pilot is heading toward one of these, the model and vendor will not save it.
Four Modes of Pilot Failure
| Metric | Failure Pattern | Root Cause and What Good Looks Like | Improvement |
|---|---|---|---|
| Inconclusive pilot | Cannot prove improvement or harm. Steering committee asks 'did it work?' and the answer is 'sort of'. | Root cause: no clean before/after baseline. Good: 4 weeks of pre-pilot measurement on the same metric the AI will move. | Defensible |
| Successful demo, failed scale | Sandbox results brilliant, production results disappointing. Vendor blames data, business blames vendor. | Root cause: pilot ran on cherry-picked data with no integration load. Good: pilot uses production data, real volumes, real exception rates from week one of go-live. | Scalable |
| Adoption collapse | Pilot team loves it, broader team ignores it. Daily usage drops below 20% within 60 days of handover. | Root cause: change management treated as training, not as job redesign. Good: cohort selected for adoption representativeness, not enthusiasm. | Sticky |
| Compliance retrofit | Pilot worked, but privacy, Fair Work, or industry-regulator review surfaces issues that cannot be retrofitted. | Root cause: Privacy Impact Assessment and security review skipped at pilot stage. Good: OAIC-style PIA completed before any production data touches the system. | Defensible |
The four failures share a single pattern: the pilot was designed to demonstrate enthusiasm, not to produce a defensible go/no-go decision. The next section is about designing the inverse.
Part 2: Five Design Decisions That Decide the Outcome
The five decisions below are made (or avoided) in the first seven days. Every successful AU pilot we have studied or supported nails these. Every failure misses at least two.
The 5 Design Decisions Made at Kickoff
The most common omission, by a long margin, is the kill criteria. Most pilots have an implicit "we will continue if it looks promising" stance, which means the project never gets killed even when the data says it should. NIST's AI Risk Management Framework is explicit on this: pre-committed go/no-go thresholds are a baseline governance requirement.
For the staffing decisions behind these roles, see our piece on the AI agent staffing gap in Australian businesses. For vendor scoping at this stage, the vendor selection question set is worth running before kickoff.
Part 3: A 90-Day Pilot Framework
The 90-day window earns its length. It is short enough to hold executive attention and long enough to capture a real operational rhythm including month-end. Compress it to 30 days and you are running a demo. Stretch it past 120 and the sponsor disengages, the goalposts move, and you end up in what Gartner calls "AI pilot purgatory".
The 90-Day Pilot Framework
The single most important calendar item is the Day 0 baseline measurement. If you cannot answer "what was the cycle time, error rate, and cost per transaction in the four weeks before the pilot started" with a number drawn from your operational systems, the pilot will not produce a defensible go/no-go regardless of how well the AI performs.
Part 4: Metrics That Survive a Go/No-Go Review
Boards do not accept "users liked it". They accept measurable, attributable change against a baseline. Below are the metrics that survive a steering committee challenge, and the anti-metrics that do not.
Metrics That Survive vs Anti-Metrics That Do Not
| Metric | Defensible Metric | How to Measure, Target, Evidence | Improvement |
|---|---|---|---|
| Cycle time | Median minutes from input received to output produced. | Pull from workflow system, compare pilot cohort to 4-week pre-pilot baseline on same cohort. Target: pre-committed % reduction. | Survives |
| Error rate | % of outputs requiring correction within 30 days. | Stratified audit sample, same audit method as pre-pilot. Track both AI errors and human override errors. | Survives |
| Adoption rate | Daily active use across cohort, not 'logins ever'. | System telemetry, 30-day rolling window. Pre-committed floor (e.g. >70% daily active by day 60). | Survives |
| Cost per transaction | Fully loaded: model + integration + oversight time + exception handling. | Finance reconciles against actual AP and time-tracking data. Compare to pre-pilot fully loaded cost. | Survives |
| Quality / accuracy | Domain-specific (e.g. classification accuracy, extraction F1). | Held-out test set, audited by SME independent of vendor. Target pre-committed. | Survives |
| Time-to-output | P50 and P95 latency at production load. | Measured at peak hour, not average. P95 matters for SLAs. | Survives |
Anti-Metrics: Do NOT Count These as Success
| Metric | What People Try to Count | Why It Does Not Survive Scrutiny | Improvement |
|---|---|---|---|
| Users said it was great | Qualitative survey, often from the enthusiast cohort. | Not attributable, not generalisable, vulnerable to social desirability bias. | Reject |
| Pilot completed on time | Project management metric. | Measures the team, not the outcome. Many on-time pilots produce inconclusive results. | Reject |
| Vendor was responsive | Relationship metric. | Confuses delivery quality with vendor sales effort. Useful for vendor selection, not go/no-go. | Reject |
| Demo went well | Stakeholder presentation outcome. | Demos are curated. Production is not. | Reject |
| Number of users onboarded | Counts access, not use. | Logins ever is not adoption. Daily active use is. | Reject |
Part 5: Five Anti-Patterns Killing AU Pilots
These are the five anti-patterns we see most often in Australian businesses. None of them are about the model.
Five Anti-Patterns Killing AU Pilots
| Metric | Anti-Pattern | What Goes Wrong and How to Avoid It | Improvement |
|---|---|---|---|
| 1. No baseline | Pilot starts measuring on day one of go-live. There is nothing to compare against. | Run 4 weeks of pre-pilot measurement on the same metric, same cohort, same definitions. Without it, no result is defensible. | Critical |
| 2. Cohort too small or too biased | Pilot runs with 3 enthusiasts. Results do not generalise to the wider team that will inherit the system. | Select a cohort that mirrors operating reality: a mix of high, mid, and low performers across at least two locations or teams. | Critical |
| 3. Free vendor pilot with no exit clause | Vendor offers a 'free' 90-day pilot. By day 90 the integration is so deep that switching costs lock you in regardless of results. | Pre-commit exit clauses, data export rights, and a documented switching playbook. See our piece on [why DIY without understanding the lock-in costs more](/blog/ai-agents-australian-businesses-why-not-diy-without-understanding/). | Critical |
| 4. Pilot ran on synthetic data | Sandbox results brilliant, real-world results poor because production data has edge cases, formatting variance, and PII the synthetic data did not. | Use production data (de-identified where required for the PIA) from week one of build. Synthetic data is for vendor demos, not pilot evidence. | Critical |
| 5. No change management | Adoption stalls because the workflow change was not co-designed with the people who do the work. | Follow a real change management process. See our piece on [AI change management and employee adoption](/blog/ai-change-management-employee-adoption/). Without it, even a technically successful pilot dies on handover. | Critical |
Part 6: Governance and Compliance at Pilot Stage
A pilot is not a regulatory exemption. The Australian Government's Department of Industry, Science and Resources (DISR) Voluntary AI Safety Standard, published in 2024, applies from pilot through to production. Treating pilot as a "we will deal with compliance at scale" phase is the most common reason a pilot cannot be scaled even when it works.
Pilot-Stage Governance Sequence
For data flows that involve cross-border processing, the gap between Australian Privacy Principles and other regimes is meaningful and worth understanding at pilot stage rather than retrofitting later. Our explainer on Privacy Act versus GDPR covers the pilot-stage implications.
Part 7: Is Your Pilot Designed to Produce a Defensible Go/No-Go?
Five-minute self-assessment for the COO or programme manager. If you cannot answer "yes" to all five, the pilot will not produce a decision the board accepts.
Defensible Go/No-Go Readiness Check
If three or fewer are "yes", do not run the pilot yet. Spend two weeks fixing the gaps. The cost of fixing them before kickoff is a fraction of the cost of an inconclusive 90-day pilot.
Part 8: The 12-Question Pre-Pilot Readiness Checklist
Use this as the agenda for the kickoff meeting. If any question cannot be answered, stop and resolve before proceeding.
- What single business metric will this pilot move? Not "improve customer service". Specific: cycle time, error rate, cost per transaction, conversion.
- What is the pre-pilot baseline value of that metric? Drawn from operational systems, not estimates.
- What is the pre-committed target value and the date by which it must be hit? Signed by the executive sponsor.
- What are the pre-committed kill criteria? Cost, safety, compliance, adoption floor. Each with a threshold and an owner.
- Who is the executive sponsor and what is their fortnightly time commitment? "As needed" is not an answer.
- Who is the cohort, how were they selected, and is the selection representative of operating reality?
- What production data will the pilot use, and what is the PIA status?
- What is the integration plan and what is the rollback plan if integration destabilises a production system?
- Who owns change management, what training has been designed, and what does the workflow look like post-pilot?
- What is the audit logging design and who reviews the logs weekly?
- What is the vendor exit clause and what is the data export design?
- What is the executive readout format for Day 90 and who attends?
Pilot Economics Summary
A typical AU AI pilot, costed honestly across 90 days:
Indicative 90-Day Pilot Budget (Typical Business)
These ranges align with Hays and Robert Half 2025 Australian salary data for the contributing roles and with published vendor pricing for enterprise AI platforms. A pilot that costs less than this range is almost certainly underspending on the governance, change, and baseline-measurement work that determines whether the result will survive go/no-go. For the full payback maths post-decision, see our payback period calculator article.
What This Article Replaces
If you are running a pilot using a vendor-supplied template, that template is almost certainly optimised to demonstrate the vendor's product, not to produce a defensible business decision. The 90-day framework above replaces the vendor template with an Australian governance-aligned alternative.
For pilots in regulated sectors, layer the industry-specific overlay on top: APRA-aligned finserv, healthcare, government supplier requirements, accounting firm overlay, construction, aged care.
For the prior decisions that determine whether the pilot is the right shape in the first place: the board business case and the AI model selection guide and the AWS/Azure/GCP services overview.
Next Step
If your business is preparing to launch an AI pilot in the next quarter and you want a second pair of eyes on the design before kickoff, that is the most useful conversation we can have with you. We help Australian businesses design pilots that produce defensible go/no-go decisions rather than inconclusive demos. Our services in AI strategy and managed AI services cover the design-through-handover lifecycle.
Book a 30-minute pilot design review. No pitch, no slides. Walk us through your draft charter and we will tell you where the framework above flags risk.
Related Reading:
- The Automation Business Case Template for the Australian Board - The approval framework that precedes pilot kickoff.
- Operating AI Agents in Production: The Reality for Australian Business - What happens in months 3 to 24 after a successful go decision.
- AI Change Management and Employee Adoption - The discipline that makes the difference between adoption and collapse at handover.
- AI Vendor Selection Questions for Australian Business - The question set to run before signing a pilot contract.
- The AI Agent Staffing Gap in Australian Business - Who needs to be on the pilot team, and what to do if they do not exist internally.
Sources: Gartner forecasts on agentic AI project cancellation (2024-2025), MIT Sloan Management Review enterprise AI research, McKinsey State of AI annual reports, DISR Voluntary AI Safety Standard 2024, OAIC Privacy Impact Assessment guidance, NIST AI Risk Management Framework, CSIRO Data61 and Australia's National AI Centre publications, Productivity Commission automation research, Hays and Robert Half 2025 Australia salary guides.