๐ฉ The part of CI/CD nobody enjoys
You want a GitHub Actions job to rsync a build onto an EC2 box and restart a service. Simple, right?
Then reality shows up:
- ๐ The key. You generate an SSH key pair, paste the private half into
secrets.SSH_PRIVATE_KEY, and append the public half to~/.ssh/authorized_keyson the instance. That key now lives forever. It never rotates. Anyone who can read repo secrets โ or any action youuses:that decides to be clever โ has shell on production. - ๐ Port 22. GitHub-hosted runners come from a huge, changing pool of egress IPs. So you either open
22/tcpto0.0.0.0/0(๐), or you write a scheduled job that pulls GitHub's meta API and rewrites your security group ingress rules โ dozens of CIDRs, churning weekly, on every instance you deploy to. - ๐ฐ The bastion. The "proper" fix. Now you have an extra instance to patch, monitor, pay for, and whose own key you also have to manage. Congratulations, the problem has a second copy of itself.
- ๐ง The leftovers. A public IP on a box that has no business having one. A key on an ex-employee's laptop. A
known_hostsprompt that hangs a job at 2am.
Every one of these is accepted as "just how deploys work." It isn't, anymore.
โก Enter AWS Systems Manager Session Manager
Session Manager flips the direction of the connection. ๐
The SSM Agent on your instance makes an outbound HTTPS connection to AWS and holds it open. When you want in, you ask the SSM API for a session, and AWS brokers the two ends together over that existing channel.
Read that again, because everything good follows from it:
- ๐ซ Zero inbound rules. Security group ingress can be completely empty. Port 22 closed. To everyone. Forever.
- ๐ณ๏ธ No public IP, no bastion. Private subnet instances work identically. Behind a NAT gateway, or with no internet at all if you add the three VPC interface endpoints.
- ๐ชช IAM is the auth layer. Access is an IAM policy, not a file on a disk. Revoke a role and access dies instantly โ no hunting for
authorized_keysentries. - ๐ CloudTrail sees every session start, attributed to the identity that opened it.
And the underrated trick: the AWS-StartSSHSession document turns that broker into a raw byte tunnel, which OpenSSH will happily use as a ProxyCommand. Meaning SSH itself still runs โ end-to-end encrypted, host keys and all โ it just stops caring about routing.
Pair it with EC2 Instance Connect (SendSSHPublicKey) and the last piece falls over too: you push a freshly generated public key into instance metadata, where sshd picks it up for 60 seconds and then forgets it. ๐ฅ An SSH key with a one-minute shelf life. Nothing to rotate, nothing to leak.
That's the whole idea. The annoying part is the boilerplate: generate a key, push it, write a ProxyCommand block into ~/.ssh/config, get the host-key checking right, and tear it all down afterwards.
So I packaged it. ๐ฆ
๐ The action
ankurk91/setup-ssh-over-ssm-action does the setup and the cleanup. It doesn't wrap ssh โ it configures the runner, then gets out of your way.
name: Deploy
on:
push:
branches: [ main ]
permissions:
id-token: write
contents: read
jobs:
deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_IAM_ROLE_ARN }}
aws-region: us-east-1
- uses: ankurk91/setup-ssh-over-ssm-action@v1
with:
instance-id: ${{ vars.EC2_INSTANCE_ID }}
os-user: ubuntu
- run: ssh ssm-target 'uptime'
No secrets. No key. No security group rule. ๐ OIDC gets short-lived AWS credentials, the action does the rest.
๐งฉ What it actually wrote
That middle step generates an ephemeral ed25519 key, pushes it via EC2 Instance Connect, and drops a fenced block at the top of ~/.ssh/config:
Host ssm-target
HostName i-0123456789abcdef0
User ubuntu
Port 22
IdentityFile "/home/runner/.ssh/ssm-ssm-target-i-0123456789abcdef0-9f3c1a2b"
IdentitiesOnly yes
StrictHostKeyChecking accept-new
UserKnownHostsFile "/home/runner/.ssh/ssm-ssm-target-i-0123456789abcdef0-9f3c1a2b.known_hosts"
ServerAliveInterval 30
ControlMaster auto
ControlPath "/home/runner/.ssh/ssm-9f3c1a2b.sock"
ControlPersist 1h
ProxyCommand sh -c "aws ssm start-session --target %h --document-name AWS-StartSSHSession \
--parameters 'portNumber=%p' --region us-east-1 \
--reason 'setup-ssh-over-ssm-action/18273645/1/9f3c1a2b'"
Match all
Two details worth pointing at. The --reason stamp carries the workflow run id, so every tunnel is labelled in the Session Manager console and in CloudTrail โ and the post step can find exactly its own sessions later. And Match all closes the stanza, because the block goes in first in the file and must not swallow any global directives you keep above your own Host lines.
Everything else is ordinary SSH config, which is the point: anything that speaks SSH just works, unmodified ๐ ๏ธ
- run: rsync -az --delete ./build/ ssm-target:/var/www/app/current/
- run: scp ./config/production.env ssm-target:/srv/app/.env
- run: ssh ssm-target 'cd /srv/app && ./bin/migrate --no-interaction'
- run: ssh ssm-target 'sudo systemctl reload nginx'
- run: ansible-playbook -i inventory.yml site.yml # ansible_host: ssm-target
Name the alias whatever reads well in your pipeline:
- uses: ankurk91/setup-ssh-over-ssm-action@v1
with:
instance-id: i-0123456789abcdef0
host-alias: app-server
os-user: ec2-user # Amazon Linux
Bring your own key instead of the ephemeral one โ it must have no passphrase, and it's the required route for hybrid mi- managed nodes, which EC2 Instance Connect doesn't support:
- uses: ankurk91/setup-ssh-over-ssm-action@v1
with:
instance-id: mi-0123456789abcdef0
private-key: ${{ secrets.SSH_PRIVATE_KEY }}
๐งน And it cleans up after itself
A post step running on always() removes the config block, deletes the key material, closes the multiplexed master connection, and terminates the SSM sessions carrying this run's marker. Nothing survives the job. ๐ซง
โ What you need
On the runner โ AWS CLI v2 (recent enough to accept aws ssm start-session --reason) and the Session Manager plugin. Both are preinstalled on GitHub-hosted Ubuntu runners, so: nothing. On self-hosted, add install-aws-cli-action and install-session-manager-plugin-action.
On the instance โ Linux, sshd running (the port can stay firewalled shut), SSM Agent โฅ 2.3.672.0, AmazonSSMManagedInstanceCore on the instance profile, and the ec2-instance-connect package (preinstalled on AL2023 standard and Ubuntu 20.04+).
On the runner role โ ssm:StartSession, ssm:DescribeInstanceInformation, ssm:DescribeSessions, ssm:TerminateSession and ec2-instance-connect:SendSSHPublicKey. The ready-to-paste policy, plus the mistakes that cause most failures, are in docs/IAM.md. ๐ The big one: scoping ssm:TerminateSession with ${aws:username} silently does not work under OIDC federation โ use a tag condition instead.
โ ๏ธ Caveats
Sharp edges, up front:
- ๐ No command-level audit. The tunnel is opaque to AWS, so session logging captures nothing readable and CloudTrail notes only that a session opened. If you need a record of what ran, reach for SSM Run Command instead.
- ๐ It's a relay, not a pipe. Tens of MB move fine; a multi-gig artifact is miserable. Ship those through S3.
- ๐ You still own
sshd. The door moved, it didn't disappear โ keep patching it. - โฑ๏ธ Keys expire after a minute. Only the first connection has to beat the clock (multiplexing carries the rest, for an hour of
ControlPersist), so put the action right before the steps that use it. Long gaps in the pipeline? Pass your ownprivate-key. - ๐ท๏ธ Give overlapping jobs their own
host-alias. Keys, sockets and known_hosts files are named per run, but the config block is keyed on the alias โ so two jobs sharing an alias and aHOMEon one self-hosted runner will tread on each other. - ๐ง Linux only, runner and instance both.
Session cleanup, happily, is not on this list: each run stamps its tunnels with a unique --reason marker and the post step terminates only those, so concurrent jobs sharing a role, a runner and an instance never cut off each other. The long version of all of this lives in docs/Caveats.md.
๐ฌ Wrapping up
Long-lived SSH keys in CI secrets and IP-whitelisted port 22 are habits from before Session Manager existed. Swapping them out costs you one step in a workflow file and an IAM policy โ and your deploy commands don't change at all.
๐ Links
- ๐ฆ The action: github.com/ankurk91/setup-ssh-over-ssm-action
- ๐ IAM policies and gotchas
- โ ๏ธ Full caveats
- ๐ AWS: allow SSH connections through Session Manager
- ๐ AWS: install the Session Manager plugin
- ๐๏ธ AWS API:
SendSSHPublicKey - ๐งฐ
install-aws-cli-actionยทinstall-session-manager-plugin-action
โญ If this saved you from writing another security-group-updating cron job, star the repo โ it's the signal that tells me which parts to keep building, and it helps the next person find it instead of pasting a private key into a secret.
Issues and PRs welcome. ๐ Closed port 22 already? Tell me how it went in the comments. ๐ฌ
Top comments (4)
Killing the long-lived key and the open port 22 in one move is the right call, and SSM also removes the bastion instance that usually sits there billing around the clock and needing patches
What tends to get missed afterwards is the VPC endpoints for ssm, ssmmessages and ec2messages, each one bills hourly per AZ, so three endpoints across three AZs is a fixed monthly line before a single deploy runs
Curious whether you went with the endpoints or routed through a NAT for this, since that choice is where most of the cost of this pattern ends up
Thanks. Good point on the endpoint line item. A few notes from building this.
The action is indifferent to NAT vs endpoints. Runner-side calls
(SendSSHPublicKey, StartSession) hit public AWS APIs. The EIC key lands
in IMDS. Only the instance's SSM Agent needs a private path out on 443.
For no-NAT subnets I use interface endpoints, but it's two, not three.
SSM Agent >= 3.3.40.0 prefers ssmmessages over ec2messages, and AWS is
retiring ec2messages. Newer regions never had it. Upgrade the agent,
confirm the endpoint SG allows 443 from the instance SG (otherwise the
agent silently falls back), then drop ec2messages. Add the free S3
gateway endpoint for agent updates.
On cost:
So "endpoints vs NAT" is really decided by whether the box needs
internet egress for its own reasons.
How does cleanup behave when GitHub terminates the runner before the post step runs, leaving the SSM session and SSH config behind?
Fair point, but it depends on how the job dies.
On a normal cancel the post step still runs: GitHub re-evaluates
ifconditions,always()stays true, and cleanup gets a 5-minute budget. State is saved at the very start of the main
step, so even a cancel mid-run leaves the post step knowing what to clean. It's genuinely
skipped in three cases: a force-cancel via the REST API (it bypasses the conditions that
would let cleanup run), cleanup overrunning the 5 minutes, or the runner just vanishing
(spot reclaim, crash).
When that happens:
GitHub-hosted: nothing to clean. The VM is destroyed and never reused, so the config
block, key and control socket die with it.
Self-hosted: the
~/.ssh/configblock and key files survive โ but inert. The ephemeralkey sits in instance metadata for 60 seconds and never reaches
authorized_keys, so theleftover private key opens nothing. Key material is named per run, so orphans accumulate
instead of colliding, and the next run on the same
host-aliasoverwrites the block andwarns.
find ~/.ssh -name 'ssm-*' -mtime +1 -deleteon a cron mops up. If you pass your ownprivate-key, that leftover is a real credential โ clean it up.The SSM session, either way: it ends itself at the Session Manager idle timeout (20 min
default, 1โ60 configurable); set Maximum session duration for a hard cap. A stray session is
not an access path โ the tunnel still lands on
sshd, which wants the expired key. Everysession carries
setup-ssh-over-ssm-action/<run id>/<attempt>/<token>in itsReason, soorphans are easy to list and trace back to the run.