Glue versus Lambda versus EMR
AWS gives you three popular ways to run a Python data job: Lambda, Glue, and EMR. They look similar from a distance. Up close they bill differently, scale differently, and break differently. Picking the wrong one costs money and weekends.
Decision tree: Lambda, Glue, or EMR
Lambda bills you for milliseconds and tops out at 15 minutes. Glue bills per DPU-second with a one-minute minimum and runs as long as your job needs. EMR bills per cluster-hour whether your job is busy or idle. The shape of the cost curve usually picks the tool for you.
Glue scales by adding workers, not by changing tools. The same code that runs on 2 DPU runs on 20 with a config bump. The breakpoint is around tens of GB or hours of compute. Past that, EMR with autoscaling starts to win on cost. Below it, Glue wins on operational overhead because there is no cluster to babysit.
Matching exercise: Match the workload to the runtime
Loading practice…
Quiz: Quiz
Loading practice…
AI prompt: Try it: classify your existing jobs
Loading practice…