Disclosure: I have no affiliation with Blacksmith. Nobody asked me to write this and nobody paid for it. I installed their GitHub App like anyone e...
For further actions, you may consider blocking this person and/or reporting abuse
Timing the work with date inside the job instead of the duration GitHub reports is what makes these numbers usable, and it is also what bounds them. Queue time and runner boot sit outside that window, and on the hosted side that is where a fair amount of the waiting lives. If you still have the runs, the gap between the job's created_at and your first date line would give you the other half of what a developer actually feels.
The other thing ten trials can hide is cache warmth. If both sides were pulling the same dependencies across their ten runs, trial 10 is not measuring the same job as trial 1, and part of the tighter variance on one side can be a warmer cache rather than a faster machine. Interleaving the two runners rather than running ten of each in a block would separate those.
The org-only wall is the detail I would not have found without hitting it. Queued with no error is the worst failure mode a CI product can ship.
Good questions, and I can actually answer the queue-time one with real numbers since the runs are still on GitHub.
First, the trial structure: it's not 10 separate workflow runs per side, it's one workflow_dispatch that kicks off two jobs in parallel, github-hosted and blacksmith, and each job internally loops 10 times inside a single script (run-trials.sh) after one throwaway warmup call. So queue and boot only happen once per side per benchmark run, not 10 times, and there's no dependency install step at all, no setup-node, no npm ci, just a local node server.js and node client.js talking to each other on the same runner. That actually removes your cache-warmth concern almost entirely, since there's nothing being pulled across trials for a cache to warm. The one thing that could still drift trial to trial is JIT/filesystem warmth within the loop itself, and the single warmup call before it starts is the only mitigation for that, so a fair critique would be that one warmup call might not fully flatten that curve.
On queue time, I pulled the job-level timestamps from the Actions API for the two clean runs (after moving the repo into an org): created_at to started_at gap was 2 to 3 seconds for github-hosted and 8 to 9 seconds for blacksmith. So in this small sample, Blacksmith's queue and boot time was actually longer, not shorter, than GitHub-hosted's, which eats into the wall-clock win somewhat even though the work itself still finished faster once it started. That's a genuinely different picture than "faster all the way through," and it's a small sample (2 runs), so I wouldn't generalize it hard, but it's real data, not inference.
On the org-only wall, you're right that it's the worst kind of failure, and the API confirms just how bad: the three runs before I moved the repo into an org show blacksmith job durations of 323, 364, and 719 seconds against github-hosted's steady 5 to 7 seconds on the same trigger, sitting there the whole time with no error surfaced.
Appreciate the actual methodology pushback, this is the kind of comment that makes we wanna go check the raw data instead of trusting my own summary of it.
That structure answers the question and moves where the uncertainty sits rather than removing it. Ten trials inside one job are ten samples of one runner instance, not ten samples of a runner. The spread you report is within-instance variance. What a developer feels across a week is mostly between-instance: a noisy neighbour, a different host, one slow box. This design measures none of that, and two clean runs is two samples of it. Running the same dispatch on a few separate days and reporting the between-run spread next to the within-run spread would close it, and it costs only calendar time.
The queue numbers are the more interesting half, and they belong in the article rather than in a comment. A six second penalty against a three to four times speedup gives a crossover you can state exactly: the hosted job has to be doing more than about eight seconds of real work before Blacksmith is ahead on the clock. That is w > (q_b - q_h) / (1 - 1/s), with your 8 to 9 against 2 to 3 seconds and s of 3.5. Below it the queue eats the win, well above it the queue is noise. For a reader deciding whether to switch, that sentence is worth more than the multiplier.
The 323, 364 and 719 second spread is the part that makes the org wall dangerous. A flat hang reads as broken and gets reported. A varying one reads as a slow machine, and people just wait.
The within vs between-instance point is fair, and yeah, it's cheap to just close, I'll run the same dispatch on a few more days and post the between-run spread next to what I've already got.
That crossover formula is genuinely better than what I wrote. "8 seconds of real work before Blacksmith wins on the clock" says something a reader can actually use, the multiplier alone doesn't. Might go add that to the post directly.
And the 323/364/719 thing, hadn't thought about it that way, but you're right, a hang that varies just looks like a slow day instead of a bug, so nobody flags it. Kind of a rough one.
Faster runners only save money if the price per minute doesn't eat the gain, a 3x speedup at more than 3x the rate is a wash on cost even though it wins on wall time
Curious whether you compared cost per job as well as duration across the two runners
Fair challenge, and no, I didn't measure billed cost directly, the test ran entirely inside Blacksmith's free 3,000 minutes/month, so I never got an invoice from either side. That said, your "more than 3x the rate" assumption doesn't hold based on what's currently published: GitHub's own docs list the standard Linux 2-core runner (what ubuntu-latest actually is) at $0.006/min, and Blacksmith's pricing page states $0.004/min for the 4vCPU runner I tested, so Blacksmith is cheaper per minute, not pricier, despite being the bigger box.
Run the numbers through: if GitHub-hosted takes X minutes at $0.006, Blacksmith takes roughly X/3.17 to X/3.77 minutes at $0.004. That works out to Blacksmith costing about 18 to 21% of the GitHub-hosted job, not a wash, a real cost win layered on top of the wall-time win.
Caveat since I want to be straight about it: that's arithmetic on two providers' listed prices, not a reconciled bill from either one, and I noted in the piece I can't independently confirm Blacksmith's stated rate is what actually gets charged. If you've seen a real invoice that tells a different story, I'd genuinely want to know, since that would be a more interesting finding than the one I have.
Fair correction, my 3x rate premise was wrong for the numbers you quoted. At $0.004 a minute against $0.006, and 3.17 to 3.77 times faster, the cost per job comes out around 4.7 to 5.7 times lower on Blacksmith, on top of the speed
That's a much stronger result than the timings alone, and it deserves a line in the post
One thing left to check is whether the free minutes hide any queue or startup time the invoice would show later
Did you see any startup or queue delay on the Blacksmith side once the repo was in the organization
Once it was actually in an org with a working image tag, no real queue lag, unqueued in seconds pretty much every time. The weird stuff (personal-repo block, one job that got cancelled around 5 min in) happened earlier and didn't look like queueing, more like setup/account stuff. Didn't see it again after that.
But honestly, the numbers I posted don't cover this at all. I timed with date wrapped right around the node call, so the clock only starts once the job's already running. Queue and boot time on either side is a total blind spot in what I measured, same thing Hayrullah's poking at in the other thread.
So, no invoice surprise that I noticed, but I genuinely don't know how the two compare on steady-state queue time, never tested it. Diffing the job's created_at against my first date line would actually answer that, probably the right next thing to check.
Diffing created_at against your first timed line is exactly the right check, that answers the real question your numbers couldn't
One thing worth separating out once you do that comparison: does your 3-4x number assume a pre-warmed runner pool on both sides, or could part of the gap be image pull / cold boot rather than raw execution speed. Self-hosted runners often skip GitHub's queue entirely but can still eat boot time if the image isn't already warm
Good question, and I can actually answer this one clean: no, cold boot and image pull aren't in the 3-4x number at all. Same reason queue time isn't, I wrapped date right around the node invocation, so the clock starts after the runner's already up and serving the job, not from job trigger. Whatever pull/boot cost either side pays happens entirely before my timer starts.
So the two things are separate. Queue-to-start timestamp diff will tell us about getting the runner ready. The 3-4x is purely "once both sides are hot and running the exact same script, who finishes first," nothing about image warmth baked in either direction.
Rip to my company's AWS/Runner budget, but hey, at least the builds are blazing fast 😂.
hahaha yeah, 0 to 100 real quick 😁😁