Six places AWS cost quietly leaks, and how to find them in an afternoon
Cloud cost problems are rarely architectural
When a bill grows faster than the business, the instinct is to suspect the architecture. Occasionally that is right. Far more often the bill is carrying a dozen small things that nobody owns, each too minor to raise in a stand-up and collectively larger than anyone expects.
You can find most of them in an afternoon with the console and a spreadsheet. Here are the six places we look first, in the order we look.
One: storage that outlived its workload
Start with block storage volumes in an available state. Each one belonged to an instance that no longer exists, and each one bills every hour regardless.
Then look at snapshots. Snapshot sprawl is almost universal, because deleting a snapshot feels risky and keeping one feels free. It is not free. Sort by age, find the ones whose source volume no longer exists, and ask whether anyone would genuinely restore a machine image from two years ago.
Object storage is the same story with a different shape. Check whether lifecycle rules exist at all. Vast numbers of buckets hold every version of every object forever because versioning was enabled, which was sensible, and no expiry was configured, which was not.
Two: non-production running like production
Development and staging environments are usually sized by copying production and then never revisited. They also typically run twenty-four hours a day for a team that works eight.
Two questions. Does non-production need production's instance sizes, and does it need to be awake at 3am on a Sunday? For most teams the honest answers are no and no, and a schedule that stops them overnight and at weekends removes roughly two-thirds of their cost without a single architectural change.
The common objection is that someone occasionally works late. Make the start-up a self-service action rather than keeping everything awake for the exception.
Three: logs kept forever by default
Log retention defaults to never expiring in more places than people realise. Years of debug output from a service that was retired accumulates quietly, and it is charged for both storage and, if anything queries it, for scanning.
Decide retention per log group deliberately. Operational logs for debugging rarely need more than a few weeks. Audit logs needed for compliance need exactly as long as the obligation says and not a decade more. Those are different answers and they should not share a default.
Four: data moving for no reason
Traffic between availability zones is charged, and it is invisible in the architecture diagram. A service in one zone chatting constantly to a database in another generates a line item that nobody can trace back to a design decision, because the design decision was made by a scheduler.
Equally common: service traffic leaving through a managed gateway when a private endpoint would keep it inside the network at a fraction of the cost. Look at your network charges and ask which flows they represent. If nobody can answer, that is the finding.
Five: provisioned capacity that has never been approached
Anything you provision rather than consume on demand is a bet on future load. Bets go stale.
Compare the provisioned figure against the observed peak over a sensible window. A database provisioned for ten times its busiest hour, a stream with far more shards than its throughput needs, a cache sized for a traffic pattern that changed — all of these bill at the bet, not at the use.
The counterpart mistake is real, so be careful: trimming capacity to just above observed peak leaves nothing for a genuine spike. Leave deliberate headroom, and write down why you chose that number.
Six: idle managed resources
Load balancers with no healthy targets. Gateways with no traffic. Clusters with no workloads. Elastic addresses attached to nothing. Each is a small hourly charge attached to something nobody uses, and each exists because deleting infrastructure feels more dangerous than leaving it.
This category is the easiest to clear and the most satisfying, because nothing depends on any of it.
Do the boring thing before the clever thing
There is a strong pull towards the interesting optimisation — a rewrite onto a cheaper runtime, a migration to a different database, an architecture that scales to zero. Those can be right, and they are also months of engineering.
The six items above are an afternoon of looking and a few days of deleting, and in most estates they are worth more than the rewrite. Do them first. Then, if the bill is still wrong, you are looking at a real architectural question rather than at accumulated neglect, and you will be able to see it clearly.
Then make it stick
Cost findings come back. Volumes detach again, snapshots accumulate again, a new log group defaults to forever again.
So end the afternoon by writing down two things: which tag or account identifies each spend, and who looks at this monthly. Without an owner the exercise is a one-off saving. With one it becomes a number that stays where you put it.
Ready to Unlock the Full Power of AWS?
Let’s talk about your cloud goals — no pressure, no hard sells. We’ll audit your setup, suggest improvements, and help you scale smarter.





