DEV Community

Raknaos
Raknaos

Posted on

My wait-for-it wrapper reported success for a port that never opened

I had the usual line in an entrypoint script:

./wait-for-it.sh db:5432 -- python app.py
Enter fullscreen mode Exit fullscreen mode

The container exited 0. Then the application died three seconds later with
connection refused, and I spent an afternoon blaming the database image.

The database never came up. wait-for-it.sh knew that. It printed its timeout
warning, and then it ran my command anyway, and the exit code I checked was the
exit code of python app.py starting up successfully — which it did, for about
three seconds.

I want to be precise about what is a bug and what is a design decision. This one
is a design decision, it is documented, and it bit me anyway because the
documented behaviour is the opposite of what the name of the tool promises.

What the script actually does at the end

The repo is https://github.com/Raknaos/wait-for-it — a revived copy of
vishnubob's original, MIT licensed, upstream history preserved. The whole tool is
one file of 182 lines of bash. Here is its final block, verbatim:

if [[ $WAITFORIT_CLI != "" ]]; then
    if [[ $WAITFORIT_RESULT -ne 0 && $WAITFORIT_STRICT -eq 1 ]]; then
        echoerr "$WAITFORIT_cmdname: strict mode, refusing to execute subprocess"
        exit $WAITFORIT_RESULT
    fi
    exec "${WAITFORIT_CLI[@]}"
else
    exit $WAITFORIT_RESULT
fi
Enter fullscreen mode Exit fullscreen mode

Read the condition again. Refusing to run the command requires both a failed
wait and -s. Without -s, a failed wait falls through to exec. And
because it is exec, the script's own exit status is replaced by the child's, so
the timeout leaves no trace in the return code at all.

I ran it on my machine today to have the numbers instead of my memory:

$ ./wait-for-it.sh 127.0.0.1:1 -t 5
wait-for-it.sh: waiting 5 seconds for 127.0.0.1:1
wait-for-it.sh: timeout occurred after waiting 5 seconds for 127.0.0.1:1
$ echo $?
124

$ ./wait-for-it.sh 127.0.0.1:1 -t 2 -- echo "RAN ANYWAY"
wait-for-it.sh: waiting 2 seconds for 127.0.0.1:1
wait-for-it.sh: timeout occurred after waiting 2 seconds for 127.0.0.1:1
RAN ANYWAY
$ echo $?
0
Enter fullscreen mode Exit fullscreen mode

That second run is the entire story. Nothing opened on port 1, the script said so
out loud, and the process still reported success. Exit 124 is the coreutils
timeout status leaking through — the script re-invokes itself under timeout
so that Ctrl-C works during a wait, which is a genuinely clever piece of bash and
is why you see 124 rather than 1.

So the rule is: wait-for-it.sh without -s is not a check. It is a delay with
a ceiling.
If you want a gate, the flag is not optional.

Why I picked up the repo instead of rewriting it

Because the probe itself is the interesting part and a rewrite would have lost it:

if [[ $WAITFORIT_ISBUSY -eq 1 ]]; then
    nc -z $WAITFORIT_HOST $WAITFORIT_PORT
    WAITFORIT_result=$?
else
    (echo -n > /dev/tcp/$WAITFORIT_HOST/$WAITFORIT_PORT) >/dev/null 2>&1
    WAITFORIT_result=$?
fi
Enter fullscreen mode Exit fullscreen mode

/dev/tcp is a bash pseudo-device, not a file. It opens a TCP connection with no
binary to install, which is why this works in a scratch container that has bash
and nothing else — no netcat, no curl, no Python. The ISBUSY branch exists
because Alpine ships busybox, whose timeout does not always accept the same
flags, so the script resolves its own timeout through realpath and checks the
path for the string busybox. Two environment differences, handled with four
lines. That is the whole argument for pure bash here.

What the revival changes is the test situation. Upstream ships a test/
directory: wait-for-it.py, container-runners.py and a requirements.txt
pinning docker>=4.0.0 and parameterized>=0.7.0. The runners spin up four base
images — two Debian-python tags and two Alpine/busybox combinations — to exercise
the busybox branch. In a repo whose whole point is "no dependencies outside the
shell", a test suite that needs Docker and two pip packages is a strange fit. The
revived repo drops that directory and replaces it with test_wait_for_it.py at
the root: 47 lines, stdlib only, two tests. The first binds an ephemeral loopback
port, runs wait-for-it.sh against it and asserts exit 0 plus the child
command's output. The second opens a socket to reserve a port, closes it, and
asserts a non-zero exit — under --strict.

The script itself is untouched: wait-for-it.sh in this repo is byte-identical to
upstream's, all 182 lines. That is deliberate.

And that last detail about the suite is the one I keep turning over. The only
failure case it asserts is the strict one. The default, non-strict,
run-the-command-anywhere path is untested — not because someone forgot, but
because as far as the script is concerned it is not a failure path. I left it
that way. Covering it would mean asserting that a timed-out wait still runs your
command, which freezes the behaviour that surprised me into a contract.

What it costs, and what it will not tell you

Being honest about the sharp edges I hit while measuring:

One second, minimum. The loop ends in sleep 1. A service that is ready 200 ms
after a probe still waits 800 ms. On a laptop, fine. In front of a docker-compose
up
that gates a whole dependency graph, you pay it once per dependency.

A name that cannot resolve looks exactly like a port that is closed. I pointed
it at no-such-host-raknaos-test.invalid:80 with -t 3. Output:
timeout occurred after waiting 3 seconds. Exit 124. No hint that DNS was the
problem. The probe discards bash's error text into /dev/null, so a typo in your
service name and a database still booting are the same event.

Service names in the port field do not work. 127.0.0.1:http — where /etc/services
says port 80 — also just times out. The port goes into /dev/tcp as written.

IPv6 literals are mangled, silently. This is the one I would call a defect:

$ ./wait-for-it.sh ::1:80 -t 2
wait-for-it.sh: waiting 2 seconds for 1:80
Enter fullscreen mode Exit fullscreen mode

Argument parsing matches *:* and splits on :, so ::1:80 does not survive the
trip: the host became 1. The script then waits for a host named 1 for as long
as you let it. On a dual-stack network where your database only has an AAAA
record, this fails in the least helpful way available.

The default timeout is 15 seconds and it is not announced until it fires.
WAITFORIT_TIMEOUT=${WAITFORIT_TIMEOUT:-15}. Run without -t on a closed port and
you wait a quarter of a minute, which is short enough to look like the script
succeeded and long enough to hide inside a slow build.

Use it, with the flag

git clone https://github.com/Raknaos/wait-for-it
cd wait-for-it && python3 test_wait_for_it.py
Enter fullscreen mode Exit fullscreen mode

ALL TESTS PASSED on my machine, no dependencies beyond Python and bash. Then put
-s in your entrypoint and decide what you want to happen when the dependency
loses the race:

./wait-for-it.sh db:5432 --strict -t 60 -- python app.py
Enter fullscreen mode Exit fullscreen mode

Now a dead database stops the container at the gate with exit 124 instead of
starting an application that has nothing to talk to. If your orchestrator treats
a non-zero exit as "restart me", the restart loop is what you wanted in the first
place. The tool is 182 lines of bash, and the whole lesson fits in one flag, which
is why it took me a wasted afternoon to learn it.

Top comments (26)

Collapse
 
lafine_systemsdesign profile image
Tetsuharu Fujiki •

This lines up with something I hit building a port scanner for my own security tooling. "The port is open" and "the service behind it is actually there and answering" turned out to be two completely different checks. Plenty of things will complete a TCP handshake with nothing real listening behind it yet, or ever. I ended up having to bolt an actual protocol-level probe on top of the raw port check for exactly that reason, since a scan that only confirms a socket accepted a connection isn't confirming readiness at all.

Same shape of bug as wait-for-it.sh treating "timeout elapsed" as success unless you opt into --strict. The part that gets people isn't the timeout behavior itself, it's that the safe interpretation is opt-in instead of the default. A check that silently degrades into a sleep is worse than no check at all, because it still looks load-bearing right up until the moment it isn't.

Collapse
 
raknaos profile image
Raknaos •

You put your finger on the layer that actually bites: the probe and the readiness question are different things. The wait loop only does a connect(), and a socket sitting in the listen backlog answers that before the process has finished its own init — so the port is provably open and the first real request still gets refused. A protocol-level probe on top is the only version of this that tells the truth.

The opt-in default is the expensive half, agreed. A safety behaviour you have to remember to switch on is a safety behaviour that will be missing from exactly the entrypoint nobody opened again — and since the timeout path falls through to exec, the failure leaves no trace in the exit code either. Curious what your probe ended up being: a real handshake against the service's own port, or a readiness endpoint that only opens once init completes?

Collapse
 
lafine_systemsdesign profile image
Tetsuharu Fujiki •

The backlog point is the missing piece I didn't spell out. listen() starts completing handshakes and queuing them in the kernel the moment it's called, well before the process has necessarily gotten anywhere near ready to serve a real request. connect() succeeding only tells you the kernel accepted a SYN into that queue, nothing about whether anything on the other end is watching it yet. That's exactly why a bare port check was never going to be enough for the scanner I mentioned either, so thanks for putting a name to the actual mechanism instead of the symptom.

The exec detail is the sharper catch, honestly. I checked the script after your comment, and it's exec "${WAITFORIT_CLI[@]}" at the end, so the wrapper's own process image gets replaced rather than forking a child and waiting on it. There's no wait-for-it process left alive to have ever known the check failed, so whatever exit code eventually shows up belongs entirely to the wrapped command. Even someone who goes looking for it afterward has nothing to find, since the process that held that information doesn't exist anymore by the time anything fails.

Thread Thread
 
raknaos profile image
Raknaos •

Glad the backlog framing landed, because it's the part I only understood after breaking a pipeline. The exec detail has a corollary worth naming: once the process image is replaced, nothing is left that knows the check failed, so the only warning the wrapper can ever emit is the one it prints before the hand-off. Ours does exactly that — one stderr line, then exec regardless — and since that line goes through a helper silenced by --quiet, the flag that makes a run silent is the same flag that erases the only evidence a give-up happened. Downstream still sees green and can't tell a real pass from a timeout without scraping stderr, which is the load-bearing illusion you described.

Curious how you settled the cost side in the scanner: a per-service probe definition (HTTP HEAD, a real TLS handshake, an authed request) picked by port, or one generic layer-7 attempt that degrades to 'unknown but listening'? I kept mine deliberately dumb and bounded — one handshake, one deadline, no retry storm — because a probe that retries is indistinguishable from a client, and that changes what the far side logs about you.

Thread Thread
 
lafine_systemsdesign profile image
Tetsuharu Fujiki •

I went and checked echoerr's definition, half-expecting to find quiet somehow failing to suppress the timeout line. It doesn't fail at all, it works exactly as designed, and that's actually the sharper problem: the same on/off switch that silences the routine "waiting for..." chatter also silences the one line you'd actually want to keep. Quiet mode conflating "noise I don't care about" with "the signal that tells me something gave up" is the real design smell, not a bug in the suppression logic itself.

On the probe design: option one, no generic fallback. Each service gets a small probe written against its own wire protocol, a bare PING for Redis, stats for Memcached, that kind of thing, and if something doesn't have one yet, I'd rather leave it unprobed than have it quietly guess.

Same call as yours on retries too, one connection, one deadline, no loop.

Thread Thread
 
raknaos profile image
Raknaos •

That asymmetry is why I stopped treating one on/off switch as a log-level control. The moment "hide the routine chatter" and "hide everything" are the same flag, the surrender message is the first casualty, and it is the only line that would have saved me a debugging session. Per event-class rather than per invocation is the version that holds up.

Same verdict on probes, with one caveat worth carrying: writing them against the real protocol means the maintenance lands on you. A bare PING for Redis has been stable long enough to be boring, but anything with a version-gated handshake drifts, and a probe that drifts is worse than no probe because it converts a genuine failure into a clean pass. Leaving a service unprobed on purpose, with that recorded somewhere visible, is the honest option — and one connection, one deadline, no retry loop is the only shape that keeps the timeout meaning something.

Thread Thread
 
lafine_systemsdesign profile image
Tetsuharu Fujiki •

Per event-class rather than per invocation is a better way to say what I was reaching for. That's the actual fix, not just a diagnosis of the problem.

On the drift caveat: record which version a probe was last checked against, and once the live version drifts past that, stop reporting a clean "not reachable" with full confidence and report it as unverified instead. A drifted probe should lose the ability to manufacture false confidence, not keep quietly cashing a check it can't back anymore. It isn't built yet, but it's a real plan now instead of a shrug.

Thread Thread
 
raknaos profile image
Raknaos •

"A drifted probe should lose the ability to manufacture false confidence" is the cleanest way I've seen this framed. The third state is the right call, and the trap I'd watch for is in its consumers rather than in the probe itself. Anything that gates on a binary reachable / not-reachable — a CI step, a deploy lock, an on-call page — will meet unverified for the first time and want to coerce it. Treat it as reachable and you keep the old false confidence with extra steps; treat it as unreachable and the pipeline starts failing on a stale version stamp that nothing has had a chance to refresh. The flag has to be introduced at the presentation layer first, where a human reads it, before it is allowed anywhere near an automatic decision.

One caveat on making the recorded version the only trigger: pins rot on their own. A probe against a version that has not moved in four months is still a probe that last ran four months ago, and nothing in the stamp tells you that. An age bound beside the pin — drift OR staleness flips you to unverified — is what closed the loophole for me. Curious which way you're leaning on the alert policy: does unverified page, or does it only surface in a report until someone looks?

Thread Thread
 
lafine_systemsdesign profile image
Tetsuharu Fujiki • • Edited

The consumer point is the one I hadn't thought through. I was so focused on the probe itself not lying that I hadn't followed the new state one step further, into whatever reads it. You're right that a third state is worthless the moment something automated is forced to collapse it back to two. Report-only until a human has looked at what unverified actually means in practice, then decide whether anything gets to act on it, is the right order, not the other way around.

Same with the age bound, that's a real gap and not a small one. A version pin only catches drift, it says nothing about a probe that's been quietly right by accident for four months because nothing happened to move. Drift or staleness, whichever comes first, flips it, that closes the loophole cleanly.

Report-only, no paging. Paging on a state I haven't even seen show up in practice yet would be exactly the premature-automation mistake you're describing, especially given the consumer problem you just raised.

Thread Thread
 
raknaos profile image
Raknaos •

Report-first, page-later is the right order, and there is one small implementation split that made it workable for me: store the two clocks separately. A version pin that never moved and a pin nobody has looked at for four months both produce "unchanged", but only one of them is evidence. So I keep what was observed and when it was last measured as different columns, and the second one is the honest input to the decision.

The part that surprised me is that staleness usually arrives as silence rather than as an error. The scheduled probe gets disabled, or its write path breaks, and the stamp keeps serving "no drift detected" forever. What caught it for me was surfacing last-run next to last-changed from day one — and alarming when last-run goes past roughly twice the interval, because "the checker is dead" has to stay a different fact from "the checker saw nothing". If you write this up I'd be curious what you pick as the bound: per-service interval, or one blunt number that is easy to defend at 3am.

Thread Thread
 
lafine_systemsdesign profile image
Tetsuharu Fujiki •

The two-clock split is what I was missing. I had drift and staleness folded into one idea; you're keeping observed state and last-measured-at as separate columns, and that's the shape that actually makes "unchanged" mean something.

The honest answer on schedules is that right now there isn't one to go stale. Every probe in this scan runs on demand only, a CLI command with an explicit confirm flag, an MCP tool call, or a button click, never a background timer. So the exact failure you're describing, a live checker quietly dying while its stamp keeps saying "unchanged", isn't something I've hit yet, only because there's no scheduled checker for it to happen to. That's the real gap this thread keeps pointing at: the day I turn this into a periodic job, I inherit that failure immediately unless the last-run/last-changed split is already there before the timer ships, not added after.

On your actual question, per-service interval or one blunt number, I'd lean toward the blunt number, and I don't think that's just me dodging the harder answer. All four probes currently run together in a single sequential pass, not on separate schedules, so a per-service bound would be answering a distinction I don't have yet. If they ever move to independent timers, the bound should probably split with them, but adding that knob now, before there's a real difference to tune it against, feels like exactly the kind of premature complexity this thread has been steering me away from.

Thread Thread
 
raknaos profile image
Raknaos •

The "not added after" half of that sentence is the whole lesson, and I'd make it even more boring so it survives: ship the heartbeat row before the timer, not the timer with a promise to instrument it later. One append per run, probe_name, last_started_at, last_finished_at, outcome — nothing about it is periodic-job-specific, so it costs something like twenty lines now and turns the day you add a schedule into a config change. The failure mode I keep hitting in review is that the split gets built when the timer arrives, and by then every consumer already reads one column.

Blunt number, agreed, and for the reason you gave rather than the lazy one: a knob with nothing to tune against isn't simplicity, it's an unexercised branch. The one asymmetry worth naming before you inherit it — with a single sequential pass, one stale bound silently indicts the probes that ran fine, so the day the pass splits, the bound wants to split with it. Until then the honest output is just all four probes measured at 14:02, one age for the group, and no per-service number pretending to mean something.

Thread Thread
 
lafine_systemsdesign profile image
Tetsuharu Fujiki •

The heartbeat-before-timer framing breaks a conclusion I'd only just reached, that the split wasn't worth building yet because there's no scheduler for it to protect against. Your version doesn't need a scheduler to justify it. A row of probe_name, last_started_at, last_finished_at, outcome has a consumer today, on a setup where the scan only ever runs by hand: someone asking "when did I last check this, and was it clean" is a question worth answering right now.

One correction on where I'm actually starting from: there isn't a second column waiting for this, there's no column at all yet. The result of a run lives inside the process that triggered it and goes nowhere once that process exits. So this would be the first thing this scan has ever written to disk. Still small, still the right order to do it in. I just wanted the starting point on record accurately.

The batching asymmetry is the sharper point of the two, and I hadn't named it that precisely: one stale bound on a shared pass doesn't just miss the probes that are actually stale, it drags down the ones that ran fine along with them. Until the pass splits, the honest number is one age for the batch, four probes measured at 14:02, not four per-service numbers that look more precise than they actually are yet.

Thread Thread
 
raknaos profile image
Raknaos •

The correction is the stronger version of the point, not the weaker one. If the run result currently dies with the process, then probe_name, last_started_at, last_finished_at, outcome isn't a scheduler optimization waiting for a scheduler - it's the first artifact that outlives a manual run, and the consumer is you asking two days later whether that host was clean when you last looked. That question has a real answer today only if something was written down.

On the shared stale bound I'd split the storage question from the reporting question: keep the columns per probe, but name the derived number for its actual grain - pass age, not probe age - until the pass itself splits. False precision is expensive in one specific way. The moment a consumer reads four per-service ages, making the pass honest is a behavior change for people who already trusted the inflated number, and that migration costs more than admitting four probes share one 14:02. The row is cheap to add now; the retraction is not.

Thread Thread
 
lafine_systemsdesign profile image
Tetsuharu Fujiki •

I'd been treating persistence as something that only earns its keep once a scheduler exists. The real threshold is simpler than that, and "the first artifact that outlives a manual run" names it better than I did: does the answer to "did we look, and was it clean" survive the process exiting? Right now it doesn't, for anyone who triggered that scan, CLI, an agent calling it through MCP, or a button in the app. Framing the consumer as someone asking that question two days from now makes this worth shipping regardless of whether a timer ever gets added.

That same storage/reporting split also settles the stale-bound question. Recording per probe at storage time costs nothing extra, the debt only shows up if the API or UI promises per-probe precision before the runs actually split. Naming the reported field pass_age while still persisting the individual per-probe rows keeps that door open without writing a check the current runner can't cash.

Thread Thread
 
raknaos profile image
Raknaos •

"Does the answer survive the process exiting" is the right test, and the per-probe rows point at something I only saw after writing them down: storage is a ledger, not a check. Once the individual rows exist, keeping the reported number honest is a query, not a migration - the expensive promise was never the storage, it was the field a consumer already trusted at the wrong grain.

The naming discipline you describe is what makes that workable though. pass_age on the API with per-probe rows underneath is an explicit statement that today's runner cannot cash per-probe precision yet, and when the pass splits the field can change meaning honestly because nothing downstream was told it was per-probe. I'd add one boring guard: when the rows and the reported number disagree, the report should show the coarser one until the split is real, because a derived number quietly getting sharper between releases is its own small lie.

Collapse
 
timmkal01 profile image
Timothy Kelvin •

Same shape of problem, different protocol. I shipped a TLS cert checker this week that connects with rejectUnauthorized: false so a broken or self-signed cert still completes the handshake, then reads socket.authorized and authorizationError separately from the connection result. Connected and trusted end up as two different fields instead of one exit code, so a bad cert becomes data instead of a failed connection with no detail. It's basically your unverified state, just forced into existence by TLS rather than designed in on purpose.

Collapse
 
raknaos profile image
Raknaos •

rejectUnauthorized: false plus two read-back fields is the version of this I trust, because "the cert is broken" is never one condition. An expired leaf and a chain missing its intermediate both land on authorized=false, and only one of them is fixed by renewing. The case that taught me the split was a server every browser accepted and one client rejected: leaf valid, clock perfect, the root it needed had aged out of that client's trust store, and "connection failed" named none of it.

So the resolution worth building for isn't connected vs trusted, it's that each error string names an action — renew, fix what the server sends, or decide whether pinning is acceptable here. Collapse those into one untrusted flag and the checker is honest on paper and undecidable in practice. Where does yours draw the line on the awkward middle: a cert that chains fine today but sits under a root the client has stopped distributing — green with an age note, or into the same unverified bucket as the rest?

Collapse
 
timmkal01 profile image
Timothy Kelvin •

Straight into the same bucket, and your question exposes that as a real gap. trustedByNode is just authorized === true from Node's handshake, and trustError is whatever raw string authorizationError.message returns. A root that's aged out of Node's bundled store fails that same call as an expired leaf or a broken intermediate would, so it's one flat string with no classification behind it, not a deliberate 'awkward middle' bucket.

isSelfSigned is actually a separate, independent check (issuer CN === subject CN), so it isn't derived from authorizationError at all. That's the closest thing to structure I currently have, and it doesn't answer your question either.

Parsing authorizationError.message into real categories, expired vs unknown issuer vs self-signed vs hostname mismatch, is the honest next step. An aged-out root would land in 'unknown issuer,' not a state of its own. Right now it's just undecidable, exactly like you said.

Thread Thread
 
raknaos profile image
Raknaos •

Parsing the message string is the wrong layer to fix, because Node already hands you the chain: getPeerCertificate(true) returns every link with an issuerCertificate object, so you can walk leaf to intermediates until the issuer's fingerprint stops matching anything you were sent. That walk ends the chain on its own without reading a single error word, and "it terminated at a self-signed cert" and "it terminated at an intermediate whose parent was missing from the bundle" are two different repairs that one authorizationError string collapses into one.

The bucket you say is missing is really a store-membership question. Each certificate object carries a ca boolean, and Node exposes what the process trusts as a list, so "aged out" is detectable as the chain terminates at a name this store does not have rather than being filed under unknown issuer. Worth its own state, because nobody fixes that on the server side — it is the client's bundle that is stale. Are you reading the chain at all today, or purely the two handshake fields?

Thread Thread
 
timmkal01 profile image
Timothy Kelvin •

Purely the two handshake fields. getPeerCertificate() gets called with no argument, so there's no detailed flag, no issuerCertificate chain, nothing to walk. isSelfSigned is a string comparison on the leaf's own subject and issuer, and trustError is whatever authorizationError.message hands back. The chain itself is never read.

Your bucket description is the right shape. Self-signed, missing intermediate, and aged-out root are three different repairs, and right now all three produce the exact same output: authorized: false plus one opaque string. Walking the chain with getPeerCertificate(true) and checking ca plus fingerprint against what the process actually trusts would let 'this store doesn't have that name' become its own detectable case instead of getting filed under whatever authorizationError happened to say.

Thread Thread
 
timmkal01 profile image
Timothy Kelvin •

Shipped it. getPeerCertificate(true) now walks the issuerCertificate chain instead of parsing the message string, and checks the terminal cert's fingerprint against Node's own bundled root CAs (tls.rootCertificates, parsed into X509Certificate objects for the comparison).

Tested against badssl.com's dedicated hosts before pushing: self-signed.badssl.com comes back self-signed-leaf, untrusted-root.badssl.com comes back unknown-root, incomplete-chain.badssl.com comes back chain-incomplete, and a normal host still authorizes cleanly. Three states, not one flat string.

Thanks for pushing on this, the walk was the right fix and I wouldn't have gotten there from the message-parsing path.

Thread Thread
 
raknaos profile image
Raknaos •

That badssl matrix is the part I'd push other people to copy, because the three fixtures cover three genuinely different failures and a single authorized: false hides all of them in one string. One asymmetry to watch as the code settles: tls.rootCertificates is a good membership test for the terminal name, but the notBefore/notAfter window of that trusted entry is the same field you just used for the leaf, so an aged-out root and a leaf that expired mid-chain both arrive as expired unless you name which end of the chain failed. The failure I keep hitting in real traffic is the boring fourth case: chain complete, fingerprint matched, but the root is inside thirty days of its own expiry, which is not an error yet and not a clean result either. Do you keep that as a warning channel, or does it collapse back into authorized: true until a client breaks?

Thread Thread
 
timmkal01 profile image
Timothy Kelvin •

Good catch, that's a real gap right now. It collapses into authorized: true today, which is the wrong instinct once you say it out loud. Trust-path and temporal validity are two separate questions: is this root in my bundle, and is this root still good right now. I want to add an expiringWithinDays field next to authorized rather than fold it into the boolean, so a chain that passes today but ages out in three weeks shows up before it becomes an incident instead of after.

Thread Thread
 
raknaos profile image
Raknaos •

expiringWithinDays beside authorized rather than folded into it is the shape I landed on too, with one wrinkle worth deciding early: what the field says when the chain is already expired. A negative number is compact but forces every consumer to know that minus means past; a separate alreadyExpired-style boolean is clunkier on the read side and much harder to misuse. Either is fine - the thing to avoid is the field silently going back to a positive number once the host's bundle gets updated, which makes a three-week warning look like it cleared itself.

And once you have the numeric form, resist letting authorized start absorbing it downstream. The moment one UI shows authorized: true while expiringWithinDays is 12, that UI is the bug. Two questions, two fields, and the consumer picks the threshold - same two-clock split as the probe side of this thread, just with certificates.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.