AI Prototype to Production: What AI Coding Tools Miss


The Short Answer
AI coding tools such as Claude Code, Cursor and GitHub Copilot make the first version of a piece of software dramatically cheaper. They do far less for the work that takes an AI prototype to production: access control, data protection, behaviour under load, and the support and onboarding that let people use it without someone standing beside them. The prototype is now the cheap part. The production work costs about what it always did, and the research says the tools can make that work harder if nobody owns it.
This post is for two kinds of reader. The first is an operations or business leader whose staff member or contractor built an internal tool, client portal or customer-facing feature with an AI coding tool or a prompt-to-app builder, and who now has people depending on it. The second is a product team holding an AI-built prototype that customers have seen and want. Both face the same question: what has to change before this is something the business can stand behind?
When AI Coding Tools Speed Things Up, and When They Don't
The most rigorous study to date is a randomised controlled trial by METR, published on 10 July 2025. Sixteen experienced open-source developers completed 246 real tasks in repositories they knew well, with and without AI tools. With AI they took 19% longer. Before starting they expected a 24% speed-up, and afterwards they still believed AI had sped them up by 20% (METR, July 2025). The authors caution against reading it as proof that AI fails most developers; what it does show is that developers can feel faster while measurably working slower.
Google's DORA research points the same way at team level. The 2024 report found that a 25% increase in AI adoption was associated with a 1.5% decrease in delivery throughput and a 7.2% reduction in delivery stability (Google Cloud, 23 October 2024). The 2025 report, drawing on nearly 5,000 respondents, found throughput had turned positive but the negative relationship with delivery stability remained. Its summary: AI amplifies what is already there, so strong teams get better and struggling teams see their problems intensify (Google Cloud, 24 September 2025).
Developers themselves are wary. In the 2025 Stack Overflow Developer Survey, 84% of respondents use or plan to use AI tools, yet 46% distrust the accuracy of the output against 33% who trust it, and only 3.1% highly trust it. The top frustration, cited by 66%, is "AI solutions that are almost right, but not quite", and 45.2% say debugging AI-generated code takes longer (Stack Overflow, 2025).
What the Research Says About AI Coding Tools
Where AI coding tools genuinely help
- Scaffolding a new project, its folder structure, build configuration and a first working screen
- Boilerplate such as form validation, API client wrappers and data transfer objects
- Writing unit tests against code that already exists and is understood
- Mechanical migrations: a library upgrade, a rename across a codebase, a format conversion
- Internal tools with a small blast radius, where a bad day means an inconvenienced colleague rather than an exposed customer record
Where they slow you down or mislead
- Large, unfamiliar codebases, where the tool lacks the context a long-serving engineer carries (the METR setting)
- Security-sensitive code: authentication, session handling, payment flows, anything touching personal information
- Ambiguous requirements, where the tool fills gaps with confident guesses instead of asking
- Long-lived systems, where duplicated code compounds into maintenance cost
That last point has data behind it. GitClear analysed 211 million changed lines of code from 2020 to 2024 and found the share of lines classed as refactoring fell from 25% in 2021 to under 10% in 2024, while copy-pasted code rose from 8.3% to 12.3% and overtook moved code for the first time (GitClear, 2025). Duplicated code costs little on the day it is written and keeps costing every time someone has to change it in three places.
For a deeper look at review practices that catch these problems, see our guide to AI code review for software teams.
Prototype vs Production: What Actually Changes
A prototype has to prove the idea works. A production feature has to keep working when the person who built it is on leave, when a customer does something unexpected, and when an auditor asks who could see what. Most AI-built prototypes handle the first and leave the second unaddressed.
Prototype That Impresses vs Feature Customers Rely On
| Metric | Prototype | Production |
|---|---|---|
| Users | The builder and a few colleagues | Every customer or staff member, including ones who misuse it |
| Authentication | Shared login or none | Individual accounts, MFA, roles, session expiry |
| Secrets | API keys in code or a .env file in the repo | Secrets manager, rotation, least-privilege keys |
| Data | Sample data, one database | Real personal information, backups, retention and deletion rules |
| Failure handling | Error shows on screen, builder fixes it | Graceful errors, retries, alerts to an on-call owner |
| Monitoring | None | Logs, metrics, traces and cost tracking per feature |
| Change control | Edits straight to the live version | Version control, review, test suite, staged releases |
| Support | Message the person who built it | Documentation, support process, defined response times |
None of the right-hand column is exotic. It is the ordinary craft of software engineering, and it is the part prompt-to-app builders tend to skip because nobody asked for it in the prompt.
The Three Production Roadblocks
Every team that moves software from demo to dependable runs into the same three walls. AI-built prototypes hit them sooner, because the build was fast enough that nobody stopped to plan for them.
Roadblock 1: Security
Veracode tested code from more than 100 large language models across Java, Python, C# and JavaScript. Forty-five per cent of samples failed security tests and introduced OWASP Top 10 vulnerabilities. Java failed 72% of the time, and the models failed to defend against cross-site scripting in 86% of relevant samples. Newer and larger models wrote more functional code, but their security performance stayed flat (Veracode, 30 July 2025).
In practice the security gaps in an AI-built tool cluster in five places:
- Authentication and authorisation. The login screen works, but the API behind it often does not check whether this user may see this record.
- Secrets. Keys for the database, email service or AI provider end up committed to the repository.
- Input handling. Form fields and file uploads pass straight into queries or pages.
- Dependencies. Tools pull in packages freely, sometimes outdated, occasionally ones that do not exist.
- Data residency. Customer data flows to a third-party AI API or a hosting region nobody chose deliberately.
If the feature itself calls a language model, the OWASP Top 10 for LLM Applications 2025 adds its own list, led by prompt injection and sensitive information disclosure, with excessive agency and unbounded consumption further down.
For Australian organisations this is a legal matter as well as a technical one. Australian Privacy Principle 11 requires reasonable steps to protect personal information from misuse, interference, loss and unauthorised access, and since the Privacy and Other Legislation Amendment Act 2024, APP 11.3 states that those steps include technical and organisational measures (OAIC, APP 11 guidelines). The OAIC received 532 data breach notifications from January to June 2025. Malicious or criminal attacks caused 59% of them and human error caused 37% (OAIC, 4 November 2025). The ASD Essential Eight controls, including multi-factor authentication, restricting administrative privileges, patching applications and regular backups, are a sensible baseline for anything a business tool touches. Our AI security checklist goes through those controls point by point.
Roadblock 2: Scale
A prototype that works for five users can fail at fifty for reasons that have nothing to do with features. Queries that scanned a 200-row table now scan two million rows. Two people edit the same record and one change disappears. A background job that ran in seconds now overlaps with its next run.
AI features add a cost dimension that traditional software lacks. Every model call is metered, so a chatty prompt or an unbounded loop can turn a modest monthly bill into an alarming one. OWASP lists unbounded consumption as a top-ten risk for this reason. Scaling well means caching, batching, rate limits per user, and a dashboard that shows cost per feature before finance asks.
Observability is the other half. Without logs, metrics and traces, a production incident becomes guesswork. We cover the running costs in more detail in operating AI agents in production.
Roadblock 3: Self-Service
This is the roadblock teams underestimate most. A tool is self-service when a new user can sign up, get the right access, learn it, get help and pay for it without the builder's involvement. That requires:
- Onboarding that works without a walkthrough
- Permissions and roles an administrator can manage without a developer
- Multi-tenancy, so one customer can never see another customer's data
- A support path with documentation and someone accountable for answers
- Billing, invoicing and plan limits, if customers pay
Multi-tenancy deserves particular care. Retrofitting tenant isolation into a single-tenant prototype touches nearly every query in the application, and it is the change most likely to introduce a data leak between customers if done in a hurry.
A Framework: What AI Builds, What Engineering Owns
Most businesses will keep using AI coding tools, so the practical question is where a mistake would land. Low-consequence, easily verified work can go to the tool with a human reviewing the result. High-consequence or hard-to-verify work needs an engineer who owns the design and uses AI as an assistant at most.
Who Should Own This Piece of Work?
Applied to a typical AI-built tool, the split usually looks like this. AI can carry user interface screens, test scaffolding, documentation drafts, data import scripts and internal admin views. Engineering should own the data model, authentication and authorisation, tenant isolation, integrations with core systems such as the ERP, CRM or accounting platform, anything handling payments or personal information, and the deployment pipeline.
The same reasoning sits behind our build vs buy TCO guide: the visible build cost is rarely the cost that decides whether a system is worth having.
A Typical Hardening Path
Taking an AI-built prototype to production follows a recognisable sequence. The durations below are a typical shape for a focused tool with one core workflow and are not a quote. A tool with payments, many integrations or several customer tenants will take longer.
Production Readiness Gates
Typical Shape of a Hardening Engagement
Gotchas that come up repeatedly
- The rebuild that should have been a refactor, and the reverse. The assessment should decide part by part. Often the user interface is worth keeping and the data layer is not.
- Nobody owns it after go-live. The person who prompted it into existence has a day job. Software without a named owner degrades.
- The test suite tests the wrong thing. AI-written tests often confirm the code does whatever it currently does. Whether that matches what the business needs is a separate question. Acceptance criteria have to come from the people who use it.
- Integration surprises. The prototype talked to a spreadsheet. Production has to talk to the finance system, with its own authentication, rate limits and data quality problems. That is system integration work and should be scoped as such.
How We Approach This at Solve8
Solve8 builds software and runs its own products in production: CallMate, ReportingMate, SupportAgent and Despatchy. Before Solve8, the founder built a workforce management SaaS that grew past 1,500 users and is still running (background here). Our principal consultant also spent 18 years on enterprise data platform, reporting and operational technology programs in Australian mining and energy, delivered while working with previous consulting employers. Across that work the pattern holds: the proof of concept is the easy part, and access control, data lineage and support handover decide whether anyone still trusts the system a year later.
We use AI coding tools in our own development, and we treat their output the way you would treat a capable junior developer's pull request, which gets reviewed before it ships.
A custom software development engagement for an AI-built prototype normally starts with the assessment week above. If what you need first is an independent view of the risk, an IT consulting review covers that, and managed AI services cover ongoing support once the tool is live.
Your next steps this week:
- List every tool in your business built with an AI coding tool or app builder, and who relies on each.
- For each one, answer the decision tree question above: if it is wrong, what happens?
- For anything touching customer data or money, check where the secrets live and who can log in.
- Book a scoping conversation and we will review the code and give you a keep, refactor or rebuild view before anyone proposes a build.
Common Questions
Are AI coding tools making developers faster?
Sometimes. A randomised study by METR in 2025 found experienced developers working in their own large codebases took 19% longer with AI tools, even though they believed they were faster. AI tools help most on scaffolding, boilerplate, tests and mechanical changes, and least on unfamiliar or security-sensitive code.
Is AI-generated code secure enough for production?
Not without review. Veracode's 2025 testing of more than 100 language models found 45% of code samples failed security tests and introduced OWASP Top 10 vulnerabilities, and newer models were no more secure. Treat AI output as a draft that needs security review, automated scanning and tests before release.
Can we keep our AI-built prototype or do we need to rebuild it?
It depends on the part. An assessment usually finds some parts worth keeping, often the user interface, and some that need rework, often the data model, authentication and tenant isolation. Deciding keep, refactor or rebuild component by component is cheaper than either keeping everything or starting over.
What does the Privacy Act require of an internal tool built with AI?
If the tool holds personal information, APP 11 requires reasonable steps to protect it from misuse, loss and unauthorised access, and since 2024 that explicitly includes technical and organisational measures. An AI-built tool carries the same duty as any other system.
How long does it take to move an AI prototype to production?
For a focused tool with one core workflow, a typical shape is around six weeks from assessment to handover, covering security fixes, stability and scale work, and monitoring. Payments, several integrations or multiple customer tenants extend that, and the assessment week is where a realistic estimate comes from.
What should AI build and what should engineers own?
Ask what happens if the code is wrong. Low-consequence, easily checked work such as screens, tests and internal admin views can be AI-built with review. Authentication, data models, tenant isolation, payments, integrations with core systems and anything handling personal information should be owned by an engineer.
Related Reading:
- From Claude Code to Production in 30 Days: hardening AI features themselves with evals, structured outputs and observability.
- Build vs Buy AI: Complete TCO Guide for Australia: why the visible build cost rarely decides whether a system is worth owning.
- AI Code Review: A Guide for Australian Dev Teams: review practices that catch the problems AI-written code introduces.
- Operating AI Agents: The Real Cost After You Build: the running costs and support load once AI is live.
Sources: METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (10 Jul 2025); Google Cloud DORA reports (23 Oct 2024 and 24 Sep 2025); Veracode 2025 GenAI Code Security Report (30 Jul 2025); Stack Overflow 2025 Developer Survey; GitClear AI Copilot Code Quality 2025 research; OWASP Top 10 for LLM Applications 2025; OAIC APP 11 guidelines and Notifiable Data Breach statistics January to June 2025 (4 Nov 2025); ASD Essential Eight.