Getting an AI Pilot Into Production
A lot of the Australian businesses we walk into have one: a pilot that worked, that everyone was pleased with, and that is still sitting in the same shared folder months later. The technology was rarely the problem. What stopped it was that nobody was named as the owner, nobody agreed what success looked like before it started, IT met the project for the first time when it asked for production access, and the demo ran on an exported copy of the data rather than the live system. This page is the checklist that gets one of those over the line, or stops it honestly.
Written for Australian businesses with a proof of concept sitting idle: the gates you have to clear on baseline, accuracy, security, privacy, integration, monitoring, rollback, run cost, support and training, plus the go or no-go decision and who signs it. Australian consultancy, with the Privacy Act and the Australian Privacy Principles built into the sign-off rather than bolted on at the end.
Realistic ROI
Why Yes AI for the Run From Pilot to Production
The build was the easy half. The half that kills pilots is the unglamorous run through security review, privacy sign-off, integration into the real systems, monitoring, run cost and the conversation with the people whose job it touches. That work needs someone who has done it before and who is willing to say no-go out loud. Four reasons Australian businesses hand us the stalled pilot.
We agree the number before anything runs
A pilot with no agreed success measure cannot pass or fail, it can only be argued about. Before we touch it we go and measure the job as it is done today: how many items a week, how many minutes each takes end to end, how many need a second person, how many get reworked, and what that costs per month at your real wage rates. That baseline is written down and agreed by whoever will make the go or no-go call. From then on the argument is about a number rather than about whether the demo felt impressive in the room.
We run the gates as a checklist, not as a feeling
Every gate on this page gets a written pass condition, a person who owns it and a date. The security review is a real review with your IT people in the room. Privacy sign-off names which personal information moves, where it rests and how long it is kept. Monitoring means a view somebody actually looks at and an alert that reaches a rostered person. Nothing is waved through because the pilot went well, because a pilot going well is not evidence about production.
Australian, and honest about which rules actually bind you
We are an Australian consultancy, so the sign-off is written against the Privacy Act and the Australian Privacy Principles, your own privacy policy and any obligations specific to your sector, not a generic overseas template. We will also tell you plainly which parts are legal requirements, which are your own internal policy and which are simply good practice, because treating everything as mandatory is how a sign-off stalls for months over something nobody was ever obliged to do. We are not a law firm, so what we write is general information and the documentation to support your own decision rather than legal advice, and we will say plainly when a question belongs with your solicitor.
We build the production half that is missing
Most stalled pilots need real work to ship, not another slide. The integration into your live systems instead of a spreadsheet export, the error handling for the cases the demo never saw, the logging, the alerting, the kill switch, and the monthly reporting the owner needs. We scope, build, integrate and support that half and stay on afterwards, so a go decision is not just handing your team a prototype and hoping for the best.
The Six Gates Between a Working Pilot and Production
Each gate has a written pass condition and a named owner. A pilot that clears all six goes live with a rehearsed rollback and a known monthly cost. A pilot that fails one gets fixed or gets stopped, on the record, this month, instead of drifting into next year as an unanswered email thread.
Baseline measured, then compared with the pilot result
You cannot claim an improvement against a number you never took. Before go-live we measure the current process on the same sample the pilot will run: volume per week, minutes per item end to end, how many need a second pair of eyes, how many get reworked, and the wage cost of all of it. Then we run the pilot over that same sample and put the two side by side. The comparison has to use the same period and the same mix of work, because a pilot tested on last March's tidy records and compared with this June's messy ones will flatter itself, and nobody notices until the board asks.
Accuracy and failure rate on real edge cases
Demos are built on the happy path. Production is mostly not the happy path. Before this gate clears we assemble a test set out of the ugly real records: the scanned document that is on its side, the customer who trades under three names, the credit note, the part-paid invoice, the enquiry written in two sentences with no punctuation, the job that was cancelled then rebooked under a new number. We measure how often it is right, how often it is wrong, and above all how often it is confidently wrong, because a system that says it does not know is manageable and one that guesses with conviction is not.
Security review, access rights and privacy sign-off
This is where a pilot built on somebody's personal login meets reality. We get your IT people in the room, replace shared or personal credentials with a dedicated service account holding only the access rights the job needs, and document what data leaves your systems, where it travels, where it rests and how long it is kept. Where personal information is involved we write the privacy position against the Privacy Act and the Australian Privacy Principles, update your collection notice if it needs it, and have a named person sign it. Leaving this to the end is what turns a two week go-live into a two month one.
Integrated into the real systems, not a spreadsheet export
A pilot that ends with somebody downloading a CSV and pasting it into Xero, MYOB, HubSpot or your practice management system has not shipped, it has only moved the manual work to a different desk. Production means it writes into the system of record through a supported connection, with the awkward parts handled: which system owns each record, what happens when a write fails halfway through a batch, what happens when the same record is edited on both sides, and how a person corrects a mistake the automation made without anyone hand editing a database.
Monitoring, alerting, rollback and a kill switch
Before it goes live, somebody has to be able to see it working and somebody has to be able to stop it. We put in a plain view of what it processed, what it skipped and what it queued, an alert that reaches a rostered person rather than an unwatched inbox when the error rate or the queue depth crosses a line, a documented path back to the manual process, and a kill switch a non-technical manager can throw at four on a Friday without ringing us first. Then we rehearse that rollback once before go-live, because an untested rollback is a wish.
Run cost, support ownership, training and the go or no-go
The last gate is the one skipped most often. What does it cost to run per month in AUD, counting usage, licences, hosting, monitoring and the human review time it still needs. Who supports it on the Tuesday it breaks, and what response can they expect. Who trains the staff whose work it touches, and what were those people actually told about what happens to their role. Then a go or no-go decision, in writing, with a named decision maker, a date, and the condition under which you would switch it back off.
Six Honest Reasons Pilots Stall, and What Changes
| Task | Traditional | With Yes AI | Notes |
|---|---|---|---|
| The pilot lost its owner when the person who championed it changed roles | It quietly becomes nobody's job and sits untouched in a shared drive | A named owner with the authority to decide, working to a dated go or no-go | The owner has to be senior enough to approve a monthly cost and a change to how a team works. An enthusiastic analyst with no budget is not an owner, however good the pilot was. |
| Nobody agreed what success looked like before the pilot ran | Circular debate about whether the demo was impressive enough to back | A measured baseline and a written pass condition agreed before anything runs | Agree the number first, never after. Once people have seen the output, the target quietly moves to wherever the result happened to land. |
| IT hears about it for the first time when it asks for production access | Security review starts from zero and the go-live slips a full quarter | IT in the room at scoping, with the access model and data flows agreed early | The fastest way to lose two months is to build on a personal login and then ask for a service account, an allow-listed connection and a data flow diagram at the very end. |
| The demo was built on an exported copy of the data | It works beautifully on clean history and falls over on live records | Rebuilt against the live system with real write-back and proper error handling | A copy is always tidier than the live table. It has no half-finished records, no locked rows and nobody editing the same job at the same moment you are. |
| Nobody costed the monthly run rate before the board meeting | Approved on a build price, then questioned hard at the first monthly invoice | A monthly figure in AUD covering usage, licences, hosting and review time | Usage-based costs scale with volume, so we model them on your busiest month rather than your quietest, then put a spend cap and an alert on top. |
| The staff whose work it touches found out at go live | Quiet resistance, private workarounds, and a system nobody uses properly | The people who do that work brought in at the edge-case stage and told plainly what changes | The people who know the exceptions are the ones whose job it touches. Bring them in early and they hand you the test set. Leave them out and they hand you the reasons it will not work. |
Six Things That Go Wrong, Said Plainly
A pilot result is not a production result, and never has been
The pilot ran on a curated sample, in a quiet week, watched by the person who built it, with somebody quietly fixing anything odd before anyone else saw it. Production runs on everything, in the busiest week, unwatched, at two in the morning. Expect the accuracy you measured in the pilot to drop once it meets the full mix of real work. That is not failure, it is the normal gap, and the entire point of testing on real edge cases before you commit is that you know the size of the gap in advance rather than discovering it in front of a customer.
AI should not be the last word on anything hard to reverse
We do not put an automated system in sole charge of a step you would struggle to undo: paying money out, issuing a credit, changing somebody's pay, cancelling a booking, sending a legal or clinical communication, or lodging anything with the ATO. The pattern that works is that it drafts, prepares and routes, and a person approves the irreversible step, with that approval logged. Over time you can widen what runs unattended by tracking how often the human changed the answer, but you widen it on evidence, not on optimism.
The copy-of-the-data demo hides most of the real work
Most of the stalled pilots handed to us were built against an export, and exports are clean. The live system has half-entered records, duplicate customers, jobs edited by two people at once, fields your team repurposed years ago for something else entirely, and rows the automation must not touch. Rebuilding against the live system, with sane handling for a write that fails partway through, is often more work than the original pilot. Budget for it honestly at the start and it is a task. Discover it at go-live and it is the reason the project dies.
Run cost is a monthly commitment and it moves with volume
A build price is a one-off, a run cost is forever, and usage-based pricing rises with your busiest month, not your average one. Before the go decision we cost it per month in AUD across usage, licences, hosting, monitoring and the human review time still required, model it at peak volume, and put a spend cap and an alert in place. Well-built automations get switched off not because they failed but because nobody could tell the board what they cost to keep running.
A rollback nobody has rehearsed is not a rollback
Every production system needs a documented way back to the manual process and a kill switch a non-technical manager can throw without a phone call. Write down who is allowed to stop it, how they do it, what happens to work already in flight, and who has to be told. Then run the drill before go-live, at a quiet time, and time how long it takes. The drill usually exposes something awkward, commonly that half-processed items have nowhere sensible to go. Far better to find that in a rehearsal than in the middle of a real incident.
Sometimes the honest answer is no-go, and that is a good outcome
Not every pilot should ship. If accuracy on real edge cases is not good enough for the consequence of an error, if the run cost is higher than the work it replaces, if the only way in is a fragile screen-scrape of a system the vendor is retiring, or if the volume is simply too low to justify the support burden, then no-go is the right call. Write down the reason, the number that would have to change, and the date to look again. A documented no-go is a decision the business can learn from. A pilot left drifting for a year is the same no-go taken slowly, with none of the learning and all of the cost.
How Yes AI Helps Get It Over the Line
Pilot triage and baseline
We look at what you already have, work out how much of it is genuinely reusable, and measure the current process properly so there is a real baseline to compare against. You come out with an honest read on how far the pilot actually sits from production and exactly what is missing.
Running the gates with your people
We chair the gates: the edge-case test set, the security and access review with IT, the privacy position against the Privacy Act and the Australian Privacy Principles, the integration design, the monitoring and rollback plan, and the run cost in AUD. Each gate gets a written pass condition and a named owner.
Building the production half
We build what a pilot never has: the integration into your live systems, error handling for the cases the demo never saw, logging, alerting, the kill switch and the monthly reporting the owner needs. One Australian team scopes it, builds it, integrates it and supports it afterwards.
Go-live, training and support
We train the team whose work changes, rehearse the rollback before cutover, sit with you through the first weeks and stay on afterwards. If the honest answer at the gate is no-go, we write up the reason, the number that would change it, and when to revisit.
Our Pilot to Production Run
Five steps from a stalled pilot to a decision you can defend. Most Australian businesses reach a documented go or no-go inside two to six weeks, depending on how much production work the pilot is missing and how quickly security and privacy sign-off can be scheduled.
Triage the pilot and measure the baseline
We review what exists, how it was built, what data it ran on and who has touched it, then measure the current manual process on the same work: volume, minutes per item, rework rate and cost per month at your real wage rates. We agree the pass condition in writing and name the person who will make the go or no-go call, before anything else begins.
Set the gates and gather the edge cases
We turn the checklist into your gates, each with a written pass condition, an owner and a date. Then we collect the real edge cases from the people who do the work, which is the step nearly every pilot skips, and build them into a test set that reflects your actual mix rather than the tidy examples the demo was built on.
Security, privacy and integration design
IT comes in properly: a dedicated service account with least-privilege access rights, documented data flows, and the privacy position written against the Privacy Act and the Australian Privacy Principles wherever personal information is involved. In parallel we design the real integration into your systems of record, including who owns each record and what happens when a write fails.
Build, test and rehearse the rollback
We build the production half, run it against the edge-case set, and measure accuracy, failure rate and how often it is confidently wrong. We stand up monitoring and alerting, brief the staff whose job it touches, and rehearse the rollback and kill switch at a quiet time so you know exactly how long a full stop takes and who has to be told.
Go or no-go, then run it
The named decision maker sees the baseline against the result, the edge-case numbers, the security and privacy sign-offs, the monthly run cost in AUD and the support arrangement, and makes a call that gets written down and dated. On a go we cut over, train, watch it closely through the first weeks and support it. On a no-go we document the reason and the trigger to revisit.
Related Reading
AI Implementation Roadmap
The wider sequence a pilot sits inside.
Writing an AI Project Brief
Set the gates before the pilot starts.
AI Incident Response
A production gate you should not skip.
AI Failure Modes
The edge cases a happy-path demo never hits.
AI Change Management
The people half, which is where most pilots die.
AI Readiness Checklist
Whether the business is ready to run anything in production.
FAQ
Get the Pilot Over the Line, or Stop It Honestly
Book a free call and we will run your stalled pilot against these gates, tell you which ones it would fail today, and give you a straight read on whether it is worth shipping. No obligation, no jargon, and a real answer either way.
All discussions held in confidence. Australian-based consultants.