Every AWS tutorial shows you what the service you’re building costs. Lambda is $0.20 per million invocations. Those numbers are real and they are cheap.
What the tutorial leaves out is the cost of the services around the one you’re building. NAT Gateway, cross-AZ data transfer, CloudWatch Logs, API Gateway, forgotten EBS snapshots. None of them appear on the architecture diagram. All of them appear on the bill.
I’ve deliberately not put a total on any of this. What you pay depends entirely on your traffic shape, and a made-up monthly figure would tell you nothing about your own bill. The arithmetic below is worth more than a number: plug in your own volumes.
NAT Gateway
$0.045 per hour, per gateway, plus $0.045 per GB processed. The hourly charge is the one that catches people, because it’s charged for existing — roughly $33 a month before a single byte moves. Run one per availability zone for high availability and you’re paying that three times.
You need it because a Lambda inside a VPC has no route to the internet, or to an AWS service endpoint.
Two ways out:
Take the Lambda out of the VPC. Since 2019, a Lambda outside a VPC can reach most AWS services directly. If yours only talks to DynamoDB and S3, it doesn’t need to be in a VPC at all, and the NAT Gateway disappears with it.
Use VPC endpoints. If you do need the VPC — an RDS instance, typically — gateway endpoints for S3 and DynamoDB are free, and interface endpoints for other services are billed hourly plus about $0.01 per GB rather than $0.045. For traffic that’s mostly hitting AWS services, that’s the bulk of the NAT bill gone.
Cross-AZ data transfer
$0.01 per GB, and it’s charged in each direction. Internet egress is $0.09 per GB after the first 100 GB free per month.
The reason this hides so well is that nothing in your code distinguishes a same-AZ call from a cross-AZ one. Work the arithmetic for your own service graph: three services calling each other 1,000 times a second at 2 KB per request is around 86 GB a day, which at $0.01/GB in each direction is real money for internal chatter that produces no customer value.
The fix is AZ-aware routing — prefer an instance of the callee in your own AZ, fall back to any:
const az = await fetch(
"http://169.254.169.254/latest/meta-data/placement/availability-zone",
).then((r) => r.text());
function getServiceUrl(serviceName: string): string {
const endpoints = serviceDiscovery.get(serviceName);
const sameAZ = endpoints.filter((e) => e.az === az);
const pool = sameAZ.length > 0 ? sameAZ : endpoints;
return pool[Math.floor(Math.random() * pool.length)].url;
}
On EKS you don’t need to write that — topology aware routing does it with an annotation. How much it saves depends entirely on how chatty your services are and how they’re spread across zones, so measure before and after rather than believing a percentage.
The catch worth stating: routing preferentially within an AZ trades some resilience for cost. If that AZ degrades, you’ve concentrated your traffic in it.
CloudWatch Logs
$0.50 per GB ingested, $0.03 per GB per month stored, and the default retention is never expire.
That default is the expensive part. Ingestion is a one-time charge per log line; storage is forever. Multiply your daily ingest by 30 and by $0.50 for the monthly ingestion cost, then note that the stored volume — and its bill — grows every month until you set a retention policy.
Debug logging in production is the usual culprit, because the cost is proportional to volume and debug logging is what makes volume large.
const LOG_LEVEL = process.env.LOG_LEVEL || "warn";
const LEVELS = { debug: 0, info: 1, warn: 2, error: 3 };
function log(level: keyof typeof LEVELS, msg: string, data?: unknown) {
if (LEVELS[level] < LEVELS[LOG_LEVEL as keyof typeof LEVELS]) return;
console.log(JSON.stringify({ level, msg, data, ts: Date.now() }));
}
Three things that help immediately:
warnin production. Turndebugon deliberately, for a window, when you’re actually debugging.- Set a retention policy. Two weeks for dev, whatever your compliance requires for prod. This is a one-line change per log group and it is the single highest-leverage thing in this post.
- Export to S3 for anything long-lived. S3 Standard-IA is $0.0125 per GB per month against CloudWatch’s $0.03, and Athena is a better query engine than Logs Insights for large scans anyway — Insights bills $0.005 per GB scanned, which gets expensive precisely when you have enough logs to need it.
API Gateway
$3.50 per million requests for REST APIs. HTTP APIs are around a third of that and do less.
At low volume this is irrelevant. At high volume it’s worth comparing against an Application Load Balancer, which is billed hourly plus a capacity-unit charge rather than per request — so for a high-request, low-complexity workload, ALB plus Lambda is usually cheaper.
The question to ask is whether you use what you’re paying for. API Gateway gives you usage plans, API keys, request validation, throttling per client. If you’re using none of those, you’re paying per-request for a proxy.
Because the two price on completely different axes, there’s a crossover point where containers on Fargate or ECS get cheaper than per-request serverless. Where it falls depends on your request size, compute duration and access patterns, so I’d rather you model it than take a number from me. The direction is what matters: serverless cost scales linearly with requests; container cost scales sublinearly. Below the crossover serverless is cheaper and you pay nothing at idle. Above it, you’re paying a per-request tax on a machine you could have rented outright.
EBS and the things you forgot
Migrating gp2 volumes to gp3 is free, cheaper per GB, and gives you 3,000 IOPS as a baseline instead of gp2’s 3 IOPS per GB. There is essentially no argument for staying on gp2.
The real waste is snapshots. Delete an EC2 instance and its snapshots remain. Delete a volume and its snapshots remain. They are nobody’s job, they don’t appear in any dashboard you look at, and they bill monthly forever.
Same for RDS: if CPU utilisation sits very low, the instance is oversized, and instance cost roughly halves with each size down. Multi-AZ doubles it again. If you know a database will run for the next twelve months — and you do — there’s no reason to pay on-demand rather than take a Savings Plan.
The actual advice
Put a recurring half hour in your calendar. Open Cost Explorer, group by service, sort by cost descending, and look for three things:
- services you don’t recognise
- services whose cost went up since last month
- services with no traffic that are still billing
That’s it. The expensive item is almost never the one you’d guess, which is the entire reason this post exists.