Writing / 2016

AWS Lambda: When Serverless Makes Sense and When It Doesn't

AWS Lambda fits short, bursty, event-driven work. Five questions to decide, and where the 5-minute limit, cold starts, and steady-load costs bite.

Lambda is useful. The hype around it isn’t.

I’ve been using AWS Lambda on a data-ingestion system for a few specific tasks: processing incoming data events, running lightweight transformations on S3 uploads, firing off notifications when certain thresholds trigger. It works well for those. Really well.

What I haven’t done is rewrite our entire backend in Lambda functions. And if you listen to the conference circuit right now, that apparently makes me a dinosaur.

The hype is exhausting

Lambda has been generally available for about a year. In that time, the narrative has gone from “interesting compute primitive” to “servers are dead, long live serverless.” I’ve watched otherwise reasonable engineers pitch rewriting stable, well-understood services into chains of Lambda functions because it’s the future, or because ops is hard and this seems to make it someone else’s problem.

Ops doesn’t disappear because you stopped managing EC2 instances. It changes shape. You still need to think about permissions, packaging, deployment, monitoring, and failure modes. You just think about them differently. Anyone who tells you Lambda eliminates operational complexity is selling something.

Where Lambda actually earns its keep

Event-driven, short-lived, bursty. That’s the sweet spot.

On that system we use it for processing webhook payloads from data providers. A payload arrives, Lambda transforms it, writes to DynamoDB, done. The function runs for a few hundred milliseconds and costs nearly nothing. Running a dedicated service for that would be waste.

S3 triggers are another clear win. File lands, thumbnail gets generated, metadata gets extracted. No process sitting idle waiting for uploads that might come once an hour or once a second. Lambda scales both directions without intervention.

Glue logic between services. Translating payloads, orchestrating a couple of API calls, reformatting data. The kind of thing that doesn’t justify a full service but still needs to run somewhere. Lambda is perfect for this.

Prototyping. When I’m validating an idea and the last thing I want is to provision infrastructure, Lambda lets me focus on behavior. Ship the function, see if the concept works, decide later if it needs a real service behind it.

Where Lambda breaks down

The five-minute execution limit is real. I’ve seen teams try to break long-running jobs into chains of Lambda invocations with S3 or SQS as coordination layers. Congratulations, you just built a worse version of a worker process. With more failure modes.

Cold starts. If your function hasn’t run recently, the runtime initializes from scratch. That adds latency, sometimes hundreds of milliseconds, sometimes more depending on the runtime and package size. For internal event processing, nobody cares. For a user-facing API with a sub-100ms promise, that’s a problem you can’t hand-wave away.

Statelessness is a feature until it isn’t. Every invocation starts clean. You can store state in DynamoDB or S3, but now you have added network round-trips and consistency concerns to what was supposed to be a simple function. If your logic needs in-memory state or connection pooling, Lambda will fight you the entire way.

Cost math flips at steady load. Lambda pricing is great when traffic is spiky and idle time dominates. When traffic is steady and predictable, dedicated compute is cheaper. I’ve run the numbers for our heavier workloads. Reserved instances win by a wide margin once utilization stays above 30-40%.

Five questions before anything goes on Lambda

Before putting anything on Lambda, I run through five questions:

  1. Can it finish in under two minutes without contortions?
  2. Does it tolerate variable latency on cold starts?
  3. Is the traffic pattern spiky or unpredictable?
  4. Can it run stateless without bolting on coordination infrastructure?
  5. Is it simple enough that debugging through CloudWatch logs won’t make someone quit?

Mostly yes? Lambda. More than one firm no? Traditional service. This isn’t complicated.

The right architecture is boring

The best setup I’ve built combines both. Lambda handles event processing, glue, and bursty workloads. Go services on EC2 handle the steady, latency-sensitive, stateful work. They coexist fine. There’s no ideological conflict.

The “serverless everything” crowd treats this as a binary choice. You pick the tool that fits the constraint. Sometimes that’s Lambda. Sometimes it’s a process on a server you manage. The goal is a system that works, not a system that impresses people at meetups.

References