A command returns nothing. No message, no exit code, no log line, no stack. You wait, then you press ctrl-C, and the only thing you have afterwards is your own memory that it took a while.
This is a different category from a failure, and the difference is not severity. An error hands you a string, a code, a file and a line, and usually a timestamp. A hang hands you nothing at all, and everything built on top of it inherits the nothing: your shell waits, your editor waits, your CI job waits until somebody's limit kills it, and whatever it prints then is about the limit rather than the cause.
The system's own word for it is "slow"
man 2 read on any Unix is unusually blunt about where this comes from:
[EINTR] A read from a slow device was interrupted before any data arrived by the delivery of a signal.
"Slow device" is the whole category in two words. A local SSD is not a slow device. A network filesystem is. So is a FUSE mount, a stalled VPN route, and any cloud file provider that has to go and fetch the bytes before it can give them to you. Every one of those can turn an ordinary open() or read() into a call with no upper bound on how long it takes, and man 2 open says the same thing about opening the file in the first place.
Nothing in your program is wrong. It asked for a file. The answer has not come back.
The error you eventually see names the wrong layer
Here is the part that costs the afternoon. When a stalled read does surface an error, the message describes what the process was doing, not what was wrong.
A runtime that cannot finish loading a script reports that it could not load that script, and names it. So you go and look at the script. It is present. It is readable. You read it yourself in another window and it comes back instantly. You check permissions, you check the interpreter, you check the path, and all of it is fine, because none of it was ever the problem. Three layers down, one file the process touched before it got to yours lives on something that was not answering.
The message is accurate and it points at the wrong thing, which is worse than being wrong, because you can spend hours confirming that an innocent file is innocent.
And every timeout relabels it again
The obvious defence makes the diagnosis harder.
Wrap the call in timeout 30 and you get exit 124, which reads as "this tool is slow". Put that in CI and the job fails with something generic, and the team writes down "flaky". Add a retry and it passes sometimes, which upgrades "flaky" to "known flaky" and buys the real cause another month of life. Each layer swaps the symptom for a label about itself.
What to do instead of debugging the thing that hung
Do not investigate the command. Investigate the syscall it is sitting in.
# Is it blocked, or just busy? U means uninterruptible wait.
ps -o pid,stat,etime,command -p "$PID"
# What has it got open? The culprit is usually the last entry.
lsof -p "$PID" | tail -20
# macOS: a stack for the blocked thread, which names the syscall.
sample "$PID" 3 -file /tmp/hang.txt && grep -m5 -i 'read\|open\|stat' /tmp/hang.txt
# Linux: the same answer, cheaper.
cat /proc/"$PID"/stack; cat /proc/"$PID"/wchan
man ps defines the one that matters: U marks "a process in uninterruptible wait". A process in that state is not slow, it is waiting on something that has not answered, and no amount of reading its source will tell you what.
Then the control experiment, which is the step people skip: run the same operation with one input removed. Point the tool at a different config file, a different working directory, a different mount. If it returns instantly, the input you removed is the slow device, and you have the answer in one try instead of twenty.
Why this one is worth writing down
Because a hang destroys its own evidence. Kill the process and there is no stack, no log line, no exit code, no timestamp: nothing to attach to a ticket and nothing for anybody else to read later. The next person gets "it froze sometimes" and starts from zero.
That is also why "it hung" and "it was slow" are the least actionable things a user can tell you. They are reports about an event that left no artefact, and you cannot go back and collect one, because the only moment the state existed was while somebody was staring at a spinner and deciding whether to wait.
Whatever you want to know about a stall, you have to take it while the stall is still happening.
Top comments (0)