Why your AI pilot has not shipped, and what actually blocks it
The demo was not the hard part
The pattern is consistent enough to be predictable. A team builds something genuinely impressive in a few weeks. It demos well. Everyone in the room agrees it should go live. Then six months pass and it has not, and nobody can point to the meeting where it was cancelled, because it never was.
It is tempting to blame organisational inertia. That is rarely the real cause. Four specific obstacles stop these projects, and all four are technical questions that the pilot was never designed to answer.
One: nobody can describe what it is allowed to do
A pilot usually runs with the permissions of whoever built it. That is fine in a sandbox and fatal in a review, because the first question a reviewer asks is what this thing can reach, and the honest answer is “everything its author could”.
The fix is unglamorous and it is the single highest-value change you can make. The agent gets its own identity. Not a shared service account, not a developer's credentials, not an API key in an environment variable that four systems share. Its own identity, with permissions scoped to the specific resources the specific task needs.
Once that exists, the reviewer's question has a short answer with evidence behind it. Until it exists, every conversation about shipping goes in a circle.
Two: nobody believes the cost projection
Pilot costs are measured over a few hundred runs by one team. Production costs depend on volume, on retries, on how much context each call carries, and on the failure modes nobody has seen yet.
So the projection presented to a finance team is an extrapolation from a tiny sample, and finance teams are right to distrust it. The way out is to narrow the scope until the number can be measured rather than estimated. One process, with a known volume, instrumented so that cost per run is a figure you report rather than a figure you model.
A measured number for one process beats a modelled number for a platform, every time, in every approval conversation.
Three: it is not watched, and everyone knows it
Ask how you would find out if the agent started behaving differently — producing subtly wrong output, taking an unexpected path, calling something it used to leave alone. In most pilots the answer is that a user would eventually complain.
That answer does not survive a review, and it should not. An agent in production needs the same operational treatment as any other production service: structured logging of what it did and why, alerting on the conditions that matter, and a defined evaluation of output quality that runs continuously rather than once before launch.
This is the part most teams skip, because it is ordinary engineering rather than interesting machine learning. It is also the part that determines whether the thing is still running in a year.
Four: nobody has agreed to own it
This is the quietest obstacle and often the decisive one. When the agent does something wrong at 3am, whose phone rings? If the answer is the team that built it, and that team is a project team that disbands, then the operations group is being asked to adopt a system they did not build and cannot debug. They will decline, politely, for months.
Ownership is settled by a named person and a runbook, written before launch, that describes what the agent does, how to tell when it is misbehaving, how to stop it, and what to check first. A system with a runbook can be handed over. A system without one cannot, regardless of how good it is.
What to do instead of another pilot
The instinct after a stalled pilot is to build a better pilot. That repeats the mistake, because a pilot is optimised to demonstrate capability and the obstacles are all about operation.
Build the smallest thing that is genuinely in production instead. One process, scoped narrowly enough that you can enumerate every system it touches. Its own identity with least-privilege permissions. Instrumented so cost and quality are measured. Logged and alerted. A named owner and a runbook.
That sounds like less ambition than the pilot had. It produces something that is actually running, which the pilot did not, and it establishes every pattern the second and third processes will reuse. The second one is much faster, because the hard parts were never about the model.
Have somebody attack it before your reviewers do
One more step, and it changes the tone of the whole approval process.
Before the security review, have someone deliberately try to break the agent — feed it instructions through the data it reads, push it to use permissions beyond its task, try to extract information it should not reveal, see whether it can be steered off purpose. Write down what worked, what you changed, and what now prevents it.
Walking into a review with that document changes the conversation completely. Instead of asking your reviewers to trust a design, you are showing them the attacks and the fixes. Reviewers are far more comfortable approving something that has already been attacked by someone on their side.
That sequence — build it, break it, secure it, then run it — is how we work, and the breaking is the step that makes the rest credible.
Ready to Unlock the Full Power of AWS?
Let’s talk about your cloud goals — no pressure, no hard sells. We’ll audit your setup, suggest improvements, and help you scale smarter.





