The checklist we actually use
AWS cost optimisation checklist
This is the working checklist behind our audit, in the order we run it. It is ordered by expected saving rather than by AWS service, because the first four sections are where nearly all of the money is.
Work through it in order: commitments first because that is the largest single line in most accounts, then idle and oversized compute, then storage and snapshots, then network, then databases and logging. Skip any service that is not in your top five by spend. A first pass on a never-reviewed account takes a day and usually finds 20% to 35% of the monthly bill.
Before you start
- 1Enable Cost Explorer if it is not already on. It takes 24 hours to populate after first enabling, so do this the day before rather than the morning of.
- 2Opt in to Compute Optimizer. Free, and its recommendations take about 12 hours to appear. It does the rightsizing arithmetic for you.
- 3List every region with spend above about a dollar. Cost Explorer grouped by region. Orphans live in regions nobody opens, and an audit of only your main region misses them entirely.
- 4Rank services by six-month spend. The top five services are typically over 90% of the bill. Spend your time proportionally.
1. Commitments, usually the biggest line
| Check | Decision rule | Typical saving |
|---|---|---|
| Savings Plan coverage | Coverage below about 70% of a stable compute baseline is money left on the table | Up to 30% of covered compute |
| Existing commitment utilisation | Utilisation below 95% means you are paying for capacity you do not use | Recover the unused portion |
| Plan type | Compute Savings Plans for anything that might change; EC2 Instance plans only for a fixed family | Flexibility rather than cash |
| Term | One year, no upfront, for a first commitment. Three years only when the baseline has been stable for a year | Avoids committing to capacity you later remove |
| RDS, ElastiCache, OpenSearch, Redshift reservations | Databases are the most stable workload in any account, so reserve them | 30% to 40% off those instances |
| Rightsize before committing | Always. Commit to the baseline that remains after cleanup, not the one you have now | Prevents a three-year mistake |
2. Compute
| Check | Decision rule | Typical saving |
|---|---|---|
| Idle instances | Average CPU under 3% and max under 10% over 14 days, with negligible network | 100% of the instance |
| Oversized instances | Compute Optimizer says over-provisioned with risk Very Low or Low | 20% to 50% per instance |
| Previous-generation types | Any t2, m4, c4, r4 still running | 10% to 20%, better performance |
| Graviton candidates | Interpreted runtimes, managed services and containers with multi-arch images | ~20% |
| Non-production out of hours | Dev, staging and QA running 168 hours a week for a team that works 45 | ~70% of those environments |
| Stopped instances with attached volumes | Stopped for more than a month means snapshot and terminate | The volume cost |
| Empty and oversized EKS clusters | $73 per cluster per month before a single node runs | $73 per cluster |
| Spot for interruptible work | CI runners, batch jobs, anything with a retry | Up to 70% on those workloads |
3. Storage
| Check | Decision rule | Typical saving |
|---|---|---|
| gp2 volumes | Every one of them, no exceptions | Flat 20%, live, no downtime |
| Unattached volumes | State "available" means attached to nothing | 100% of the volume |
| Over-provisioned io1/io2 IOPS | Provisioned IOPS far above observed | Often large |
| Snapshots with no source volume | Source volume gone means nothing will restore from it | 100% |
| No snapshot lifecycle policy | Data Lifecycle Manager is free | Caps a growing line |
| Deregistered AMIs with live snapshots | Deregistering an AMI does not delete its snapshots | 100% |
| S3 buckets with no lifecycle rule | Especially log and backup buckets | Varies, often large |
| S3 versioning with no expiry on non-current versions | Every overwrite kept forever | Varies |
| Incomplete multipart uploads | Invisible in the console, billed as storage | Often surprising |
| EFS with no lifecycle policy | $0.30 per GB-month in Standard | ~90% on cold data |
4. Network
| Check | Decision rule | Typical saving |
|---|---|---|
| S3 and DynamoDB gateway endpoints missing | They are free and remove that traffic from the NAT gateway | $0.045 per GB of that traffic |
| Idle NAT gateways | Any NAT gateway with negligible traffic | $32.85 each per month |
| One NAT gateway per AZ in non-production | Resilience nobody exercises in staging | $65 per environment |
| Unassociated Elastic IPs | Charged at $3.60 a month whether used or not | $3.60 each |
| Public IPv4 on instances that do not need it | Behind a load balancer or VPN | $3.60 each |
| Idle load balancers | No targets, or no requests for 30 days | $16 to $18 each |
| Unused VPN connections and Transit Gateway attachments | $36 and $36.50 a month respectively | Per resource |
| Cross-AZ chatter | $0.01 per GB each way adds up on chatty services | Varies |
| Egress without a CDN | CloudFront both discounts the rate and caches | Varies |
5. Databases and logging
| Check | Decision rule | Typical saving |
|---|---|---|
| RDS Extended Support | $0.10 per vCPU-hour on an out-of-support engine version | Whole fee, on upgrade |
| Multi-AZ on non-production | Doubles the instance cost for failover staging does not need | 50% of those instances |
| RDS gp2 storage | Same 20% as EBS | 20% |
| Idle RDS instances | Near-zero connections over 14 days | 100% |
| Manual snapshots that outlived their instance | They persist after deletion | 100% |
| Log groups set to never expire | The default on every log group ever created | Storage on the excess |
| VPC Flow Logs capturing ALL traffic | Often the largest ingester in the account | Large |
| Debug log levels in production | $0.50 per GB ingested | Proportional |
| Custom metric sprawl | $0.30 per metric per month | Proportional |
| Dashboards beyond the free three | $3 each per month | Small but free to fix |
6. Account hygiene
| Check | Decision rule | Why it matters |
|---|---|---|
| Cost allocation tags active | Above 30% untagged spend means you cannot attribute anything | Enables everything else |
| Cost Anomaly Detection configured | Free, ten minutes to set up | Catches the next problem early |
| Budgets with alerts | At least one at your expected monthly total | Turns surprises into warnings |
| Support plan matched to use | Business support is a percentage of spend | Worth reviewing if tickets are rare |
| Resources in unexpected regions | Anything above a dollar outside your main regions | Where orphans hide |
The order matters more than the list
Every item here is real, but running them in the wrong order wastes effort. Committing before rightsizing locks in capacity you are about to delete. Migrating to Graviton before deleting idle instances means porting things you were going to throw away. Clean up, size correctly, then commit.
Put numbers on it
Is my AWS bill too high?
Benchmark your spend against typical ranges for your team size in 30 seconds.
Open calculator →
gp2 → gp3 savings calculator
A flat 20% off your EBS storage, usually with better performance and zero downtime.
Open calculator →
EC2 idle cost estimator
What your under-5%-CPU instances cost, and what rightsizing the quiet ones saves.
Open calculator →
Savings Plans calculator
On demand against one and three year commitments, and why rightsizing has to come first.
Open calculator →
Frequently asked questions
How long does working through this take?
A focused day for the first four sections on a single-account estate, assuming Cost Explorer and Compute Optimizer are already enabled. The long tail, meaning every region, every log group and every snapshot chain, is where the time goes and where automation earns its keep.
Which checks are genuinely zero-risk?
Setting log retention, adding S3 and DynamoDB gateway endpoints, migrating gp2 to gp3, releasing unassociated Elastic IPs, deleting snapshots whose source volume no longer exists, and buying a one-year no-upfront Savings Plan sized below your baseline. None of those change how anything runs.
What is deliberately not on this list?
Anything that trades reliability for cost without saying so: removing Multi-AZ from production, dropping backup retention below your recovery requirement, or running production on Spot without a fallback. Those are business decisions rather than optimisations, and an audit that presents them as savings is misleading you.
Can I run these checks automatically?
Most of them. About seventy of the checks behind this list are scriptable against a read-only role, which is what our scanner does. The ones that are not, such as whether an instance at 4% CPU is idle or memory-bound, need someone to look at it, which is why the report is human-verified before it goes out.
Keep reading
Prices checked against AWS list rates on 2 August 2026. AWS changes prices; treat every figure here as a close approximation rather than a quote.
£499 fixed. Free scan first. 20%+ found or it’s free.
The scan is read-only — a role you create and delete, no keys shared — and shows your estimated monthly saving before anyone talks about money.