DEV Community

Kiran Gunturu
Kiran Gunturu

Posted on

The fastest Glue job is the one you never start: removing per-job overhead from a 191-job pipeline

In my last post I covered making a production AWS Glue pipeline run faster by tuning individual jobs. This post is about the bigger change that followed. It had nothing to do with how the jobs ran. It was about how many times we started them.

All costs below use Glue's standard rate of $0.44 per DPU-hour, billed per second with a 1-minute minimum per job run (Glue 2.0+).

The pipeline

It's an Oracle-to-vendor feed with five layers, orchestrated by Step Functions:

Layer Glue jobs Worker config What it does
Ingestion 63 G.1X × 4 (4 DPU) Oracle → S3, one job per source table
Transformation 63 G.1X × 4 (4 DPU) Business transforms
Publish 63 G.1X × 4 (4 DPU) Write final output
Outbound 1 G.2X × 6 (12 DPU) Package into ~5 GB split files, zip, encrypt
SFTP 1 Python Shell, 1 DPU Ship to the vendor

Step Functions runs the layers in order, and each layer fans its jobs out through a Map state. Running daily, that's 191 Glue job starts every day.

Tuning helped, up to a point

The optimizations from the last post made a real difference:

Stage Before (min) After (min) Change
Ingestion 23.38 6.38 −73%
Transform 10.02 5.43 −46%
Publish 25.23 5.15 −80%
Outbound 8.00 5.55 −31%
SFTP 5.02 5.54 noise (bandwidth-bound)
End to end 1h 15m 13s 29.3 min −61%

26M records, 1.3 GB zip.

What that did to cost

Wall-clock time and cost aren't the same thing here. A layer's wall clock is roughly its slowest job, but you pay for every job in the layer. A single minute of wall clock on a 63-job G.1X × 4 layer can bill up to 63 × 4 = 252 DPU-minutes.

The estimates below treat every job in a layer as running for the layer's full duration. That makes them an upper bound: jobs that finish early cost less, so actual spend is lower. But the same method applies to every row, so the comparisons hold.

Stage DPU-hours before DPU-hours after Cost / run before Cost / run after
Ingestion 98.2 26.8 $43.21 $11.79
Transform 42.1 22.8 $18.52 $10.04
Publish 106.0 21.6 $46.63 $9.52
Outbound 1.6 1.1 $0.70 $0.49
SFTP 1.0 0.1 $0.44 $0.04
Total 248.9 72.4 $109.49 $31.87

SFTP ran on G.2X × 6 before being moved to a 1-DPU Python Shell job.

Estimated cost per run fell by about 71%, more than runtime did. That's because nearly all of the cost sits in the three 63-job layers, and those improved the most. Outbound and SFTP are single jobs, so even their slow minutes are cheap.

The run went from over an hour to under 30 minutes. But while going through the run history, I found two days that showed where the rest of the time was going.

The run that processed almost nothing

On two days the source sent almost no data: 5,822 records, a 75 KB file.

Those runs still took 25 and 27 minutes. Ingestion alone took 5 minutes to read a few thousand rows across 63 jobs. Nearly all of that time was spent before any data was touched.

Every Glue job pays the same fixed cost before it does any work:

flowchart LR
    subgraph FIXED["Fixed cost: paid on every job start, regardless of data volume"]
        A["Provision workers"] --> B["Install libraries"]
        B --> C["Start Spark context"]
        C --> D["Open JDBC connection"]
    end
    D --> E["Read and write data<br/>scales with volume"]

That cost doesn't depend on data volume. A job reading 5,000 rows pays the same startup as a job reading 5 million. Glue 2.0+ also bills a 1-minute minimum per job run, so even a very short job costs a full minute of DPU time.

What the empty run cost

Run Records Runtime DPU-hours Cost
Near-empty day 5,822 25.3 min 73.3 $32.23
Normal day 26M 1h 15m 248.9 $109.49

Both runs are pre-tuning, estimated the same way as above.

The near-empty run cost about 29% of a full run while processing 0.02% of the data. That 29% is roughly what we paid every day just to start 191 jobs.

We were paying that startup cost 191 times a day. Tuning had reduced the time spent on actual work, which only made the fixed overhead a bigger share of what was left. Further tuning wasn't going to change that, because the overhead comes from starting the job, not from the work it does.

Removing the overhead: start fewer jobs

The downstream consumer didn't need the feed every day. So instead of making each job start faster, we reduced how often the jobs start.

We made the cadence configurable: daily, weekly, or monthly. A weekly or monthly run processes the whole period's data in one pass through the same pipeline.

While building this, we took a closer look at the transformation layer. The architecture called for it as a separate stage, but in practice its 63 jobs weren't applying any business transformations. They read the ingested data and passed it through to publish unchanged. That meant 63 extra rounds of provisioning, library installs, and Spark startup every run, just to move data from one S3 prefix to another. In the daily run that layer still took 5.4 minutes. So weekly and monthly runs skip it, and each one starts 128 jobs instead of 191.

flowchart LR
    SF(["Step Functions"]) --> ING["Ingestion<br/>63 Glue jobs<br/>Oracle → S3"]
    ING --> CAD{"Cadence?"}
    CAD -->|"Daily"| TR["Transform<br/>63 Glue jobs<br/>pass-through only"]
    TR --> PUB["Publish<br/>63 Glue jobs"]
    CAD -->|"Weekly / Monthly<br/>skip transform"| PUB
    PUB --> OUT["Outbound<br/>1 Glue job<br/>~5 GB splits, zip, encrypt"]
    OUT --> SFTP["SFTP<br/>1 Glue job"]
    SFTP --> V(["Vendor"])

    classDef removed fill:#eeeeee,stroke:#999999,stroke-dasharray:5 5,color:#666666
    class TR removed

Here's what that does to the number of job starts:

Cadence Jobs / run Glue job starts / month Glue job starts / year
Daily 191 5,921 69,715
Weekly 128 ~567 ~6,674
Monthly 128 128 1,536
xychart-beta
    title "Glue job starts per month"
    x-axis ["Daily", "Weekly", "Monthly"]
    y-axis "Job starts" 0 --> 6500
    bar [5921, 567, 128]

Weekly means ~90% fewer job starts than daily. Monthly means ~98% fewer. Every one of those removed starts is a round of provisioning, library installs, and Spark startup that no longer happens.

The billing floor of a job start

Because Glue bills at least one minute per job run, every job start has a minimum price, even if the job does no work at all:

Job type Minimum billed per start
G.1X × 4 (4 DPU) $0.029
G.2X × 6 (12 DPU) $0.088
Python Shell (1 DPU) $0.007

Multiply that by the job count and cadence, and you get the least the pipeline can cost just for starting its jobs:

Cadence Floor per run Floor per month Floor per year
Daily (191 jobs) $5.64 $174.80 $2,058
Weekly (128 jobs) $3.79 $16.79 $198
Monthly (128 jobs) $3.79 $3.79 $45

These are lower bounds. Every job runs longer than a minute, so actual cost is higher, but the floor alone shows the pattern. The transform layer's share of the daily floor was $1.85 per run, about $675 a year, just to start 63 jobs that didn't change the data.

What the pass-through layer cost

Per daily run Per month (31 runs) Per year
Transform layer (63 jobs, no business logic) $10.04 $311 $3,663

Post-tuning estimate, same method as above.

What it did to runtime

Dropping the pass-through layer is a win on its own, but I wanted to measure the cadence change by itself. So the baseline here is a daily run without transform: 29.3 − 5.43 ≈ 23.9 minutes. Every gain below comes from running less often, not from removing a layer.

Cadence Avg run time Avg records / run Runs / month Pipeline time / month
Daily (excl. transform) ~23.9 min 21M 31 ~12.3 h
Weekly 65.4 min 110.5M ~4.4 ~4.8 h
Monthly 2h 43m 540M 1 ~2.7 h

Weekly: average of four full 7-day runs (107–114M records, 51–71 min each). Monthly: average of the July and August runs (505M and 575M records). The daily monthly total is an estimate based on the post-tuning run time.

A weekly run processes about 5× the records of a daily run in about 2.7× the time. A monthly run processes about 26× the records in about 6.8× the time. If runtime grew with data volume, those ratios would be about the same. They aren't, because the fixed startup cost is paid once per run no matter how much data the run carries.

xychart-beta
    title "Records processed per minute (millions)"
    x-axis ["Daily (excl. transform)", "Weekly", "Monthly"]
    y-axis "M records / min" 0 --> 4
    bar [0.9, 1.7, 3.3]

The jobs, Spark configs, and code were the same at every cadence. The weekly cadence runs about twice as efficiently as daily, and monthly about four times, only because less of each run is spent starting up.

Over a month, weekly cadence cuts total pipeline time by about 61% compared with daily, and monthly cuts it by about 78%.

What it did to cost

Monthly Glue cost by layer at each cadence (daily includes transform, since that's what production ran):

Layer Daily / month Weekly / month Monthly / month
Ingestion $365.49 $204.23 $67.34
Transform $311.09 $0 (skipped) $0 (skipped)
Publish $295.03 $114.21 $94.81
Outbound $15.14 $6.73 $3.83
SFTP $1.26 $0.27 $0.23
Total $988.01 $325.44 $166.21
Per year $11,633 $3,832 $1,995

Upper-bound estimates, same method as above. Daily uses the post-tuning run. Weekly uses the average layer times of the four full weeks. Monthly uses the average of the July and August runs.

Compared with daily, weekly cadence cuts estimated Glue cost by about 67% and monthly by about 83%. Nearly all of the savings come from the 63-job layers: ingestion and publish run far fewer times, and transform drops out completely. The single-job layers barely register either way.

The trade-off

The cost is data freshness. With a weekly feed, the vendor can see data up to seven days old. With monthly, up to a month old. That was acceptable for this consumer, which is why the change was an option at all. If your downstream needs daily data, this approach doesn't apply, and reducing per-job overhead means something different, such as consolidating many small table jobs into fewer larger ones.

The daily-cadence totals above are estimated from a 21M-record run. Low-volume days run shorter, so the real saving is somewhat less than the table shows. It's still the largest single improvement we made to this pipeline.

What I'd check on any Glue pipeline

  1. Find a run with almost no data. Its runtime is your fixed overhead floor. If that floor is a large share of a normal run, job tuning will stop paying off quickly.
  2. Count job starts, not just runtime. 191 jobs × 1 run a day looks fine on a per-run dashboard. At nearly 70,000 job starts a year, the overhead is easier to see.
  3. Ask whether the consumer needs the current cadence. Daily is often the default rather than a requirement. The cheapest overhead to remove is on runs nobody needed.
  4. Look for pass-through layers. A stage can exist in the architecture diagram without changing the data. If a layer's output matches its input, every job in it is pure overhead.
  5. Sum DPU-hours across parallel jobs, not wall clock. A Map state with 63 jobs finishes when the slowest one does, but you pay for all 63.

The takeaway

Tuning took a daily run from 1h15m to 29 minutes, and that was worth doing. But the change that mattered more was noticing that a lot of each run was startup, not work. Changing the cadence and dropping a pass-through layer took monthly pipeline time from over 12 hours to under 3, cut job starts from 5,921 to 128, and cut estimated monthly Glue cost from about $988 to about $166, without touching any job code.

When a pipeline is built from many small jobs, the fastest job is often the one you never start.

Top comments (0)