DEV Community

GAGAN R
GAGAN R

Posted on

Why your NestJS Temporal worker ignores SIGTERM

I sent SIGTERM to a Temporal worker. The worker shut down. The process kept running for an hour.

This was in the nestjs-exchange-rates sample in temporalio/samples-typescript, but the same shape of bug will show up in any NestJS app that creates a Temporal worker and also calls app.listen().

What I expected to find

There was an open issue on the repo, #377, filed in July 2024. It said the NestJS sample calls await this.worker.close(), and close() doesn't exist on Worker from @temporalio/worker. Looked like a one line fix. Change close() to shutdown(), wire it to a Nest lifecycle hook, done.

So I ran the sample and hit Ctrl+C to watch it fail.

It didn't fail.

Worker state changed { state: 'STOPPING' }
Worker state changed { state: 'DRAINING' }
Worker state changed { state: 'DRAINED' }
Worker state changed { state: 'STOPPED' }
Enter fullscreen mode Exit fullscreen mode

That is the complete shutdown sequence. Polling stops, in flight tasks drain, cached workflows evict, and run() resolves. Nothing was broken.

Temporal shuts the worker down for you

The reason is in the SDK. When you create a Runtime it installs its own signal handlers. The default shutdownSignals is ['SIGINT', 'SIGTERM', 'SIGQUIT', 'SIGUSR2'], and any of those calls shutdown() on every worker in the process.

So the broken close() method in the sample never runs. It isn't wired to onModuleDestroy or onApplicationShutdown, and main.ts never calls app.enableShutdownHooks(), so Nest wouldn't call it even if it were. Nothing calls it, which is why nobody noticed it calls a method that doesn't exist on Worker in 1.24.0.

Dead code. Worth deleting, but that wasn't the bug.

SIGTERM is different

Ctrl+C is not the test. When you press Ctrl+C in a terminal, SIGINT goes to the entire foreground process group. That includes the nest CLI wrapper that spawned your worker. The wrapper dies, the child goes with it, and your prompt comes back. Everything looks clean.

Containers don't do that. Kubernetes sends SIGTERM to one process.

So I did the same thing by hand:

$ lsof -t -i :3001
34333
$ kill -TERM 34333
Enter fullscreen mode Exit fullscreen mode

The worker logged all four states and reached STOPPED. Then:

$ lsof -i :3001
node 34333 gaganr 31u IPv6 ... TCP *:redwood-broker (LISTEN)
$ ps -p 34333
34333 ttys003 0:01.34 node --enable-source-maps .../dist/apps/worker/.../main
Enter fullscreen mode Exit fullscreen mode

Still there. I checked again an hour later and it was still there, still holding the port. The only thing that moved it was SIGKILL. I ran the whole thing again on a fresh process and got the same result.

Why it hangs

Two lifecycles, no wire between them.

Temporal's signal handler stops the worker. That is the only thing it does. It has no idea a NestJS application exists.

Meanwhile apps/worker/src/main.ts has this:

const app = await NestFactory.create(ExchangeRatesWorkerModule);
await app.listen(3001);
Enter fullscreen mode Exit fullscreen mode

The HTTP server on 3001 is still bound. Node keeps the event loop alive as long as a server is listening, and nobody ever calls app.close(). So the worker is dead, the app is alive, and the process sits there.

There was a third thing wrong, in the provider factory:

worker.run();
console.log('Started worker!');
return worker;
Enter fullscreen mode Exit fullscreen mode

run() returns a promise that resolves once the worker has finished shutting down. The sample threw it away. Nothing to await during shutdown, and if the worker ever failed at runtime the rejection went nowhere.

In Kubernetes this costs you the full termination grace period on every deploy. The pod gets SIGTERM, ignores it, waits for SIGKILL. Default is 30 seconds.

The fix

Let Temporal keep its signal handling. Give the run() promise somewhere to live, and close the Nest app when it resolves.

The provider factory stops starting the worker:

return await Worker.create({
  taskQueue,
  ...workflowOption,
  activities,
});
Enter fullscreen mode Exit fullscreen mode

The service exposes it:

@Injectable()
export class ExchangeRatesWorkerService {
  constructor(@Inject('EXCHANGE_RATES_WORKER') private worker: Worker) {}

  run(): Promise<void> {
    return this.worker.run();
  }
}
Enter fullscreen mode Exit fullscreen mode

And main.ts owns the lifecycle:

async function bootstrap() {
  const app = await NestFactory.create(ExchangeRatesWorkerModule);
  await app.listen(3001);

  try {
    await app.get(ExchangeRatesWorkerService).run();
  } finally {
    await app.close();
  }
}
Enter fullscreen mode Exit fullscreen mode

run() blocks until the SDK's signal handler shuts the worker down. Then finally closes Nest, the HTTP server releases 3001, the event loop empties, and the process exits.

Typing the injected worker as Worker is what would have caught the original close() bug. The constructor parameter had no type, so TypeScript had nothing to check it against.

How to test this properly

Don't use Ctrl+C. Get the PID and signal it directly:

npm run start:worker
# in another terminal
kill -TERM $(lsof -t -i :3001)
# wait a few seconds, then
lsof -i :3001
ps -p <pid>
Enter fullscreen mode Exit fullscreen mode

Both should come back empty. If either returns something, your process is ignoring SIGTERM and your deploys are paying for it.

The fix is in temporalio/samples-typescript PR #525.

Top comments (0)