I had the usual line in an entrypoint script:
./wait-for-it.sh db:5432 -- python app.py
The container exited 0. Then the ap...
For further actions, you may consider blocking this person and/or reporting abuse
This lines up with something I hit building a port scanner for my own security tooling. "The port is open" and "the service behind it is actually there and answering" turned out to be two completely different checks. Plenty of things will complete a TCP handshake with nothing real listening behind it yet, or ever. I ended up having to bolt an actual protocol-level probe on top of the raw port check for exactly that reason, since a scan that only confirms a socket accepted a connection isn't confirming readiness at all.
Same shape of bug as wait-for-it.sh treating "timeout elapsed" as success unless you opt into --strict. The part that gets people isn't the timeout behavior itself, it's that the safe interpretation is opt-in instead of the default. A check that silently degrades into a sleep is worse than no check at all, because it still looks load-bearing right up until the moment it isn't.
You put your finger on the layer that actually bites: the probe and the readiness question are different things. The wait loop only does a
connect(), and a socket sitting in the listen backlog answers that before the process has finished its own init — so the port is provably open and the first real request still gets refused. A protocol-level probe on top is the only version of this that tells the truth.The opt-in default is the expensive half, agreed. A safety behaviour you have to remember to switch on is a safety behaviour that will be missing from exactly the entrypoint nobody opened again — and since the timeout path falls through to
exec, the failure leaves no trace in the exit code either. Curious what your probe ended up being: a real handshake against the service's own port, or a readiness endpoint that only opens once init completes?The backlog point is the missing piece I didn't spell out. listen() starts completing handshakes and queuing them in the kernel the moment it's called, well before the process has necessarily gotten anywhere near ready to serve a real request. connect() succeeding only tells you the kernel accepted a SYN into that queue, nothing about whether anything on the other end is watching it yet. That's exactly why a bare port check was never going to be enough for the scanner I mentioned either, so thanks for putting a name to the actual mechanism instead of the symptom.
The exec detail is the sharper catch, honestly. I checked the script after your comment, and it's exec "${WAITFORIT_CLI[@]}" at the end, so the wrapper's own process image gets replaced rather than forking a child and waiting on it. There's no wait-for-it process left alive to have ever known the check failed, so whatever exit code eventually shows up belongs entirely to the wrapped command. Even someone who goes looking for it afterward has nothing to find, since the process that held that information doesn't exist anymore by the time anything fails.
Glad the backlog framing landed, because it's the part I only understood after breaking a pipeline. The exec detail has a corollary worth naming: once the process image is replaced, nothing is left that knows the check failed, so the only warning the wrapper can ever emit is the one it prints before the hand-off. Ours does exactly that — one stderr line, then exec regardless — and since that line goes through a helper silenced by --quiet, the flag that makes a run silent is the same flag that erases the only evidence a give-up happened. Downstream still sees green and can't tell a real pass from a timeout without scraping stderr, which is the load-bearing illusion you described.
Curious how you settled the cost side in the scanner: a per-service probe definition (HTTP HEAD, a real TLS handshake, an authed request) picked by port, or one generic layer-7 attempt that degrades to 'unknown but listening'? I kept mine deliberately dumb and bounded — one handshake, one deadline, no retry storm — because a probe that retries is indistinguishable from a client, and that changes what the far side logs about you.
I went and checked echoerr's definition, half-expecting to find quiet somehow failing to suppress the timeout line. It doesn't fail at all, it works exactly as designed, and that's actually the sharper problem: the same on/off switch that silences the routine "waiting for..." chatter also silences the one line you'd actually want to keep. Quiet mode conflating "noise I don't care about" with "the signal that tells me something gave up" is the real design smell, not a bug in the suppression logic itself.
On the probe design: option one, no generic fallback. Each service gets a small probe written against its own wire protocol, a bare PING for Redis, stats for Memcached, that kind of thing, and if something doesn't have one yet, I'd rather leave it unprobed than have it quietly guess.
Same call as yours on retries too, one connection, one deadline, no loop.
That asymmetry is why I stopped treating one on/off switch as a log-level control. The moment "hide the routine chatter" and "hide everything" are the same flag, the surrender message is the first casualty, and it is the only line that would have saved me a debugging session. Per event-class rather than per invocation is the version that holds up.
Same verdict on probes, with one caveat worth carrying: writing them against the real protocol means the maintenance lands on you. A bare PING for Redis has been stable long enough to be boring, but anything with a version-gated handshake drifts, and a probe that drifts is worse than no probe because it converts a genuine failure into a clean pass. Leaving a service unprobed on purpose, with that recorded somewhere visible, is the honest option — and one connection, one deadline, no retry loop is the only shape that keeps the timeout meaning something.
Same shape of problem, different protocol. I shipped a TLS cert checker this week that connects with rejectUnauthorized: false so a broken or self-signed cert still completes the handshake, then reads socket.authorized and authorizationError separately from the connection result. Connected and trusted end up as two different fields instead of one exit code, so a bad cert becomes data instead of a failed connection with no detail. It's basically your unverified state, just forced into existence by TLS rather than designed in on purpose.
rejectUnauthorized: false plus two read-back fields is the version of this I trust, because "the cert is broken" is never one condition. An expired leaf and a chain missing its intermediate both land on authorized=false, and only one of them is fixed by renewing. The case that taught me the split was a server every browser accepted and one client rejected: leaf valid, clock perfect, the root it needed had aged out of that client's trust store, and "connection failed" named none of it.
So the resolution worth building for isn't connected vs trusted, it's that each error string names an action — renew, fix what the server sends, or decide whether pinning is acceptable here. Collapse those into one untrusted flag and the checker is honest on paper and undecidable in practice. Where does yours draw the line on the awkward middle: a cert that chains fine today but sits under a root the client has stopped distributing — green with an age note, or into the same unverified bucket as the rest?
Straight into the same bucket, and your question exposes that as a real gap. trustedByNode is just authorized === true from Node's handshake, and trustError is whatever raw string authorizationError.message returns. A root that's aged out of Node's bundled store fails that same call as an expired leaf or a broken intermediate would, so it's one flat string with no classification behind it, not a deliberate 'awkward middle' bucket.
isSelfSigned is actually a separate, independent check (issuer CN === subject CN), so it isn't derived from authorizationError at all. That's the closest thing to structure I currently have, and it doesn't answer your question either.
Parsing authorizationError.message into real categories, expired vs unknown issuer vs self-signed vs hostname mismatch, is the honest next step. An aged-out root would land in 'unknown issuer,' not a state of its own. Right now it's just undecidable, exactly like you said.
Parsing the message string is the wrong layer to fix, because Node already hands you the chain:
getPeerCertificate(true)returns every link with anissuerCertificateobject, so you can walk leaf to intermediates until the issuer's fingerprint stops matching anything you were sent. That walk ends the chain on its own without reading a single error word, and "it terminated at a self-signed cert" and "it terminated at an intermediate whose parent was missing from the bundle" are two different repairs that one authorizationError string collapses into one.The bucket you say is missing is really a store-membership question. Each certificate object carries a
caboolean, and Node exposes what the process trusts as a list, so "aged out" is detectable as the chain terminates at a name this store does not have rather than being filed under unknown issuer. Worth its own state, because nobody fixes that on the server side — it is the client's bundle that is stale. Are you reading the chain at all today, or purely the two handshake fields?Purely the two handshake fields. getPeerCertificate() gets called with no argument, so there's no detailed flag, no issuerCertificate chain, nothing to walk. isSelfSigned is a string comparison on the leaf's own subject and issuer, and trustError is whatever authorizationError.message hands back. The chain itself is never read.
Your bucket description is the right shape. Self-signed, missing intermediate, and aged-out root are three different repairs, and right now all three produce the exact same output: authorized: false plus one opaque string. Walking the chain with getPeerCertificate(true) and checking ca plus fingerprint against what the process actually trusts would let 'this store doesn't have that name' become its own detectable case instead of getting filed under whatever authorizationError happened to say.
Shipped it. getPeerCertificate(true) now walks the issuerCertificate chain instead of parsing the message string, and checks the terminal cert's fingerprint against Node's own bundled root CAs (tls.rootCertificates, parsed into X509Certificate objects for the comparison).
Tested against badssl.com's dedicated hosts before pushing: self-signed.badssl.com comes back self-signed-leaf, untrusted-root.badssl.com comes back unknown-root, incomplete-chain.badssl.com comes back chain-incomplete, and a normal host still authorizes cleanly. Three states, not one flat string.
Thanks for pushing on this, the walk was the right fix and I wouldn't have gotten there from the message-parsing path.