In my last post I covered making a production AWS Glue pipeline run faster by tuning individual jobs. This post is about the bigger change that followed. It had nothing to do with how the jobs ran. It was about how many times we started them.
All costs below use Glue's standard rate of $0.44 per DPU-hour, billed per second with a 1-minute minimum per job run (Glue 2.0+).
The pipeline
It's an Oracle-to-vendor feed with five layers, orchestrated by Step Functions:
| Layer | Glue jobs | Worker config | What it does |
|---|---|---|---|
| Ingestion | 63 | G.1X × 4 (4 DPU) | Oracle → S3, one job per source table |
| Transformation | 63 | G.1X × 4 (4 DPU) | Business transforms |
| Publish | 63 | G.1X × 4 (4 DPU) | Write final output |
| Outbound | 1 | G.2X × 6 (12 DPU) | Package into ~5 GB split files, zip, encrypt |
| SFTP | 1 | Python Shell, 1 DPU | Ship to the vendor |
Step Functions runs the layers in order, and each layer fans its jobs out through a Map state. Running daily, that's 191 Glue job starts every day.
Tuning helped, up to a point
The optimizations from the last post made a real difference:
| Stage | Before (min) | After (min) | Change |
|---|---|---|---|
| Ingestion | 23.38 | 6.38 | −73% |
| Transform | 10.02 | 5.43 | −46% |
| Publish | 25.23 | 5.15 | −80% |
| Outbound | 8.00 | 5.55 | −31% |
| SFTP | 5.02 | 5.54 | noise (bandwidth-bound) |
| End to end | 1h 15m 13s | 29.3 min | −61% |
26M records, 1.3 GB zip.
What that did to cost
Wall-clock time and cost aren't the same thing here. A layer's wall clock is roughly its slowest job, but you pay for every job in the layer. A single minute of wall clock on a 63-job G.1X × 4 layer can bill up to 63 × 4 = 252 DPU-minutes.
The estimates below treat every job in a layer as running for the layer's full duration. That makes them an upper bound: jobs that finish early cost less, so actual spend is lower. But the same method applies to every row, so the comparisons hold.
| Stage | DPU-hours before | DPU-hours after | Cost / run before | Cost / run after |
|---|---|---|---|---|
| Ingestion | 98.2 | 26.8 | $43.21 | $11.79 |
| Transform | 42.1 | 22.8 | $18.52 | $10.04 |
| Publish | 106.0 | 21.6 | $46.63 | $9.52 |
| Outbound | 1.6 | 1.1 | $0.70 | $0.49 |
| SFTP | 1.0 | 0.1 | $0.44 | $0.04 |
| Total | 248.9 | 72.4 | $109.49 | $31.87 |
SFTP ran on G.2X × 6 before being moved to a 1-DPU Python Shell job.
Estimated cost per run fell by about 71%, more than runtime did. That's because nearly all of the cost sits in the three 63-job layers, and those improved the most. Outbound and SFTP are single jobs, so even their slow minutes are cheap.
The run went from over an hour to under 30 minutes. But while going through the run history, I found two days that showed where the rest of the time was going.
The run that processed almost nothing
On two days the source sent almost no data: 5,822 records, a 75 KB file.
Those runs still took 25 and 27 minutes. Ingestion alone took 5 minutes to read a few thousand rows across 63 jobs. Nearly all of that time was spent before any data was touched.
Every Glue job pays the same fixed cost before it does any work:
flowchart LR
subgraph FIXED["Fixed cost: paid on every job start, regardless of data volume"]
A["Provision workers"] --> B["Install libraries"]
B --> C["Start Spark context"]
C --> D["Open JDBC connection"]
end
D --> E["Read and write data<br/>scales with volume"]
That cost doesn't depend on data volume. A job reading 5,000 rows pays the same startup as a job reading 5 million. Glue 2.0+ also bills a 1-minute minimum per job run, so even a very short job costs a full minute of DPU time.
What the empty run cost
| Run | Records | Runtime | DPU-hours | Cost |
|---|---|---|---|---|
| Near-empty day | 5,822 | 25.3 min | 73.3 | $32.23 |
| Normal day | 26M | 1h 15m | 248.9 | $109.49 |
Both runs are pre-tuning, estimated the same way as above.
The near-empty run cost about 29% of a full run while processing 0.02% of the data. That 29% is roughly what we paid every day just to start 191 jobs.
We were paying that startup cost 191 times a day. Tuning had reduced the time spent on actual work, which only made the fixed overhead a bigger share of what was left. Further tuning wasn't going to change that, because the overhead comes from starting the job, not from the work it does.
Removing the overhead: start fewer jobs
The downstream consumer didn't need the feed every day. So instead of making each job start faster, we reduced how often the jobs start.
We made the cadence configurable: daily, weekly, or monthly. A weekly or monthly run processes the whole period's data in one pass through the same pipeline.
While building this, we took a closer look at the transformation layer. The architecture called for it as a separate stage, but in practice its 63 jobs weren't applying any business transformations. They read the ingested data and passed it through to publish unchanged. That meant 63 extra rounds of provisioning, library installs, and Spark startup every run, just to move data from one S3 prefix to another. In the daily run that layer still took 5.4 minutes. So weekly and monthly runs skip it, and each one starts 128 jobs instead of 191.
flowchart LR
SF(["Step Functions"]) --> ING["Ingestion<br/>63 Glue jobs<br/>Oracle → S3"]
ING --> CAD{"Cadence?"}
CAD -->|"Daily"| TR["Transform<br/>63 Glue jobs<br/>pass-through only"]
TR --> PUB["Publish<br/>63 Glue jobs"]
CAD -->|"Weekly / Monthly<br/>skip transform"| PUB
PUB --> OUT["Outbound<br/>1 Glue job<br/>~5 GB splits, zip, encrypt"]
OUT --> SFTP["SFTP<br/>1 Glue job"]
SFTP --> V(["Vendor"])
classDef removed fill:#eeeeee,stroke:#999999,stroke-dasharray:5 5,color:#666666
class TR removed
Here's what that does to the number of job starts:
| Cadence | Jobs / run | Glue job starts / month | Glue job starts / year |
|---|---|---|---|
| Daily | 191 | 5,921 | 69,715 |
| Weekly | 128 | ~567 | ~6,674 |
| Monthly | 128 | 128 | 1,536 |
xychart-beta
title "Glue job starts per month"
x-axis ["Daily", "Weekly", "Monthly"]
y-axis "Job starts" 0 --> 6500
bar [5921, 567, 128]
Weekly means ~90% fewer job starts than daily. Monthly means ~98% fewer. Every one of those removed starts is a round of provisioning, library installs, and Spark startup that no longer happens.
The billing floor of a job start
Because Glue bills at least one minute per job run, every job start has a minimum price, even if the job does no work at all:
| Job type | Minimum billed per start |
|---|---|
| G.1X × 4 (4 DPU) | $0.029 |
| G.2X × 6 (12 DPU) | $0.088 |
| Python Shell (1 DPU) | $0.007 |
Multiply that by the job count and cadence, and you get the least the pipeline can cost just for starting its jobs:
| Cadence | Floor per run | Floor per month | Floor per year |
|---|---|---|---|
| Daily (191 jobs) | $5.64 | $174.80 | $2,058 |
| Weekly (128 jobs) | $3.79 | $16.79 | $198 |
| Monthly (128 jobs) | $3.79 | $3.79 | $45 |
These are lower bounds. Every job runs longer than a minute, so actual cost is higher, but the floor alone shows the pattern. The transform layer's share of the daily floor was $1.85 per run, about $675 a year, just to start 63 jobs that didn't change the data.
What the pass-through layer cost
| Per daily run | Per month (31 runs) | Per year | |
|---|---|---|---|
| Transform layer (63 jobs, no business logic) | $10.04 | $311 | $3,663 |
Post-tuning estimate, same method as above.
What it did to runtime
Dropping the pass-through layer is a win on its own, but I wanted to measure the cadence change by itself. So the baseline here is a daily run without transform: 29.3 − 5.43 ≈ 23.9 minutes. Every gain below comes from running less often, not from removing a layer.
| Cadence | Avg run time | Avg records / run | Runs / month | Pipeline time / month |
|---|---|---|---|---|
| Daily (excl. transform) | ~23.9 min | 21M | 31 | ~12.3 h |
| Weekly | 65.4 min | 110.5M | ~4.4 | ~4.8 h |
| Monthly | 2h 43m | 540M | 1 | ~2.7 h |
Weekly: average of four full 7-day runs (107–114M records, 51–71 min each). Monthly: average of the July and August runs (505M and 575M records). The daily monthly total is an estimate based on the post-tuning run time.
A weekly run processes about 5× the records of a daily run in about 2.7× the time. A monthly run processes about 26× the records in about 6.8× the time. If runtime grew with data volume, those ratios would be about the same. They aren't, because the fixed startup cost is paid once per run no matter how much data the run carries.
xychart-beta
title "Records processed per minute (millions)"
x-axis ["Daily (excl. transform)", "Weekly", "Monthly"]
y-axis "M records / min" 0 --> 4
bar [0.9, 1.7, 3.3]
The jobs, Spark configs, and code were the same at every cadence. The weekly cadence runs about twice as efficiently as daily, and monthly about four times, only because less of each run is spent starting up.
Over a month, weekly cadence cuts total pipeline time by about 61% compared with daily, and monthly cuts it by about 78%.
What it did to cost
Monthly Glue cost by layer at each cadence (daily includes transform, since that's what production ran):
| Layer | Daily / month | Weekly / month | Monthly / month |
|---|---|---|---|
| Ingestion | $365.49 | $204.23 | $67.34 |
| Transform | $311.09 | $0 (skipped) | $0 (skipped) |
| Publish | $295.03 | $114.21 | $94.81 |
| Outbound | $15.14 | $6.73 | $3.83 |
| SFTP | $1.26 | $0.27 | $0.23 |
| Total | $988.01 | $325.44 | $166.21 |
| Per year | $11,633 | $3,832 | $1,995 |
Upper-bound estimates, same method as above. Daily uses the post-tuning run. Weekly uses the average layer times of the four full weeks. Monthly uses the average of the July and August runs.
Compared with daily, weekly cadence cuts estimated Glue cost by about 67% and monthly by about 83%. Nearly all of the savings come from the 63-job layers: ingestion and publish run far fewer times, and transform drops out completely. The single-job layers barely register either way.
The trade-off
The cost is data freshness. With a weekly feed, the vendor can see data up to seven days old. With monthly, up to a month old. That was acceptable for this consumer, which is why the change was an option at all. If your downstream needs daily data, this approach doesn't apply, and reducing per-job overhead means something different, such as consolidating many small table jobs into fewer larger ones.
The daily-cadence totals above are estimated from a 21M-record run. Low-volume days run shorter, so the real saving is somewhat less than the table shows. It's still the largest single improvement we made to this pipeline.
What I'd check on any Glue pipeline
- Find a run with almost no data. Its runtime is your fixed overhead floor. If that floor is a large share of a normal run, job tuning will stop paying off quickly.
- Count job starts, not just runtime. 191 jobs × 1 run a day looks fine on a per-run dashboard. At nearly 70,000 job starts a year, the overhead is easier to see.
- Ask whether the consumer needs the current cadence. Daily is often the default rather than a requirement. The cheapest overhead to remove is on runs nobody needed.
- Look for pass-through layers. A stage can exist in the architecture diagram without changing the data. If a layer's output matches its input, every job in it is pure overhead.
- Sum DPU-hours across parallel jobs, not wall clock. A Map state with 63 jobs finishes when the slowest one does, but you pay for all 63.
The takeaway
Tuning took a daily run from 1h15m to 29 minutes, and that was worth doing. But the change that mattered more was noticing that a lot of each run was startup, not work. Changing the cadence and dropping a pass-through layer took monthly pipeline time from over 12 hours to under 3, cut job starts from 5,921 to 128, and cut estimated monthly Glue cost from about $988 to about $166, without touching any job code.
When a pipeline is built from many small jobs, the fastest job is often the one you never start.
Top comments (0)