Skip to content
8 min read

The exit code that tells the orchestrator not to restart you

The error path was the careful one: log the failure, then exit rather than limp on. It exits with no argument, the runtime turns that into zero, and a restart policy reading exit codes concludes the work is done. The loudest deaths are the ones nobody hears.

Vyacheslav Pankratov· Fullstack Developer
Cover art for “The exit code that tells the orchestrator not to restart you”
On this page

TL;DR: Two correct decisions compose into a broken one. The error path exits rather than limping on, and it calls process exit with no argument — which the runtime turns into 0. The restart policy restarts on failure, and 0 is not a failure. So an unrecoverable startup error becomes a completed job: the container stops, the platform is satisfied, nothing restarts, and no alert mentions a crash because as far as anything can tell there wasn't one.

Both halves of this were written by people being careful.

The first is an error path that refuses to continue. Something the service needs is missing or malformed, and rather than limp along in an undefined state it logs the problem and stops. That is the right instinct and the one every code review asks for.

The second is a restart policy that distinguishes between a crash and a finish. Restart on failure, not always — because a container that exits cleanly has done its job, and restarting it forever is how you get a boot loop that hides a real problem behind noise.

Neither is wrong. Put them in the same system and the careful error path tells the careful restart policy that everything went fine.

  1. Unrecoverable startup errorsomething required is missing
  2. Logged in fullstack trace, cause, everything
  3. Exit, with no argumentstopping beats limping on
  4. Runtime makes it 0the documented default
  5. Policy reads successa completed job — nothing restarts
Every step is the documented behaviour of its own layer. Nothing here is a mistake, and the composition is still a service that stops and is never restarted.

The argument nobody passed

The mechanism is one omitted parameter and it is documented on both sides.

Node's own reference for process.exit([code]) states the default plainly: "If code is omitted, exit uses either the 'success' code 0 or the value of process.exitCode." So process.exit() on an error path — with nothing set on process.exitCode beforehand — is a request to terminate successfully.

On the other side, the Compose specification describes restart in one line per policy, and the relevant one is unambiguous: on-failure "restarts the container if the exit code indicates an error." Zero does not.

Neither document is hiding anything, and neither is wrong. The gap is that one of them is read by whoever writes the error handler and the other by whoever writes the deploy file, and the sentence that connects them is in neither.

What each layer sees when that error path runs:

Layer What it observes
The log the full reason, a stack trace, everything you wrote
The process's exit status 0
The restart policy a job that completed
Crash-loop alerting nothing to alert on
A view of container state Exited (0)

Only the first row is written for a human, and it is the only row a human is not looking at yet.

A demonstration, not a measurement

There is no number to report here — this is documented behaviour, not a discovery — but it is worth seeing once, because "surely something restarts it" is the reaction everyone has. Two containers, the same image, the same policy, differing only in the argument:

docker run -d --name zero --restart on-failure alpine sh -c 'sleep 1; exit 0'
docker run -d --name one  --restart on-failure alpine sh -c 'sleep 1; exit 1'
# wait a few seconds, then:
docker inspect -f '{{.Name}} restarts={{.RestartCount}} running={{.State.Running}}' zero one
/zero restarts=0 running=false
/one  restarts=6 running=true

The second container is still going round; the first stopped once and was left alone. The restart count on the second line is not a finding — it is however many restarts fit in the time the terminal was open. The finding is the first line, and specifically that nothing about it is an error state. docker ps -a reports Exited (0), which is what a finished job looks like.

The asymmetry, which is the part worth writing down

If this were only "pass an argument to exit", it would be a lint rule and not an article. What makes it worth an afternoon is where it sits.

Four services in one system shared an entrypoint shape — the same startup sequence, copied and adapted as each was carved out. Two of them pass a failure code on the error path. Two pass nothing. That split is not a decision anybody made; it is the residue of which copy happened to be made from which, and it means the answer to "does this service restart when it fails to boot" is different for each service and is written nowhere.

And then the part that makes the fix less obvious than it looks: the two that got it right import the same shared error handler as the other two, and that handler exits without a code. A service can be entirely correct at its own call site and still be terminated cleanly through a helper it did not write. Correctness at the call site does not survive a shared import — which is the same shape as a uniqueness declaration only one of two writers imports and an integer that is load-bearing in a file nobody reads that way: a property that is true in one place and enforced in none.

Why is "die loudly" the wrong instinct here?

It is not wrong — it is incomplete, and the incompleteness is invisible because the loud part works. The log line is written. The stack trace is there. Everything a human would want is present, in a file nobody opens until they are already looking for a problem.

What is missing is the part machines read. An exit status is the only thing the platform above you consumes, and it carries exactly one bit of the thing you were being careful about. Dying loudly without setting it is speaking to the audience that is not listening: a process that exited zero and was never restarted is a dead feed with a green console. No restart, no crash-loop alert, no unhealthy container — just a service that is no longer there, in a system where nothing measures its absence.

A dead feed with a green console

A process that exited zero and was never restarted is exactly that. No restart, no crash-loop alert, no unhealthy container — just a service that is no longer there, in a system where nothing measures its absence.

What to do instead, in order of cost

Set the status on the error path, not just the message. process.exit(1), or set process.exitCode before whatever calls exit. One character, and it converts a silent stop into a restart.

Audit the shared helper before the call sites. The call sites are the visible half and the smaller one. If an error handler that several services import can terminate a process, its exit status belongs to it, and getting it right there fixes every importer at once — including the ones that already looked correct.

Then make the absence measurable, because the exit code is not the whole answer. A restart policy turns a crash into a retry; it does nothing about a service that exits cleanly for a reason nobody anticipated. Something outside the process has to notice that a thing which should be producing output has stopped producing it, and that check is worth having whether or not the exit codes are right.

FAQ

Doesn't restart: always fix this? It papers over it, and it costs you the distinction the policy exists for. With always, a container that genuinely finished is restarted forever, so the boot loop that used to indicate a real problem becomes background noise. The exit code is the right place to carry the information; the policy is just reading it.

Our error path throws instead of calling exit. Are we fine? Probably, and worth checking rather than assuming. An uncaught exception exits non-zero, which is the behaviour you want — but many services install a top-level handler precisely so a crash gets logged nicely, and if that handler ends in a bare exit call you are back here with an extra step in between.

How would we find this in our own system? Grep for exit calls and read each one for an argument, then grep for whatever your top-level error handler is called and read that too. It is a five-minute search and the interesting result is usually not a missing argument — it is discovering that two services you assumed behaved identically do not.

Is this specific to one runtime? No. The default is a language and platform detail — some exit non-zero on an unhandled error, some do not — but the shape is the same anywhere a process's status is the only thing its supervisor reads. The question to ask is not "what does exit do here" but "what does the thing above us do with what exit produced".

The three things worth taking away

An exit status is an API, and it has one consumer that never reads your logs. Everything else the error path does is for people. That one integer is the entire conversation with the platform.

Correctness at the call site does not survive a shared import. Two of the four services had this right and were still killable through a helper they had in common. Audit the thing several places import before the places that import it.

A clean exit is indistinguishable from a completed job, by design. That is the correct behaviour of every layer involved, which is why nothing anywhere reports a problem — and why the only way to notice is to measure the absence of the work rather than the presence of an error.

On a market feed, a process that exited zero is a feed that stopped with every console still green, which is the class of failure market feeds that reconnect without their symbols is about.

If something in your system stopped and the first anyone heard of it was the missing data, that is the kind of thing we come in for.

Was this helpful?

// Build it

Running this in production?

A fan-out breaks where nobody is watching — a replica count that turns out to be a correctness invariant, a handler that deletes its own replacement, a retry hint that decides your loss. Tell us what the consumers are missing; we reply within a day with a concrete next step.