The replica count is a correctness invariant, and nothing says so
One line says how many of a service to run. The line above it says to start the new container before stopping the old. Nobody wrote down that the first line was load-bearing, and the second one is documented to break it.
On this page
TL;DR: A replica count of one is often not a capacity decision. If the service holds state in memory or acts on every event it sees, that integer is a correctness invariant — and it is written in a file with no comment, no assertion and no test. Directly above it sits an update policy chosen to avoid downtime, whose own documentation says the old and new tasks "briefly overlap". We measured that overlap: twelve events handled twice per rollout. Switching the policy takes duplicates to zero and replaces them, one for one, with events handled by nobody.
There are two lines involved and they sit next to each other:
deploy:
replicas: 1
update_config:
order: start-firstThe first is a number somebody typed once. The second is a choice somebody made deliberately —
stop-first is the default,
so this was turned on, presumably during a push to stop deploys from dropping traffic.
Neither line is wrong. Together they say: there is exactly one of this service, and during every deploy there are two.
The integer is doing a job the file does not describe
replicas: 1 reads as a capacity setting, and for most services that is all it is — one is enough,
raise it when one is not. That reading is correct right up until the service starts holding something.
A consumer with an in-memory cursor, a scheduler with a timer, a worker holding a lease, anything with
a Map that accumulates: the moment two of them exist, they both hold their own copy and both act on
everything they see.
At that point the integer has quietly changed job. It is no longer "how much of this do we need", it is "how many of these may exist without the system being wrong" — and the file records the two identically. There is no comment. There is no assertion at boot. There is no test that fails when it becomes 2, because raising it is a configuration change and configuration changes do not run the test suite.
This is the same shape as a cleanup that deletes by key rather than by connection: a correctness property that no code enforces, held up instead by something incidental — there, the timing of a handler; here, an integer in a deploy file. Both are correct until something ordinary moves.
Why is the policy above it annotated as safe?
Because it is safe, for the thing it is annotated about. start-first exists to stop a rollout from
dropping requests: bring the replacement up, let it become ready, then retire the old one. For a
stateless service behind a load balancer that is straightforwardly better, which is why somebody
turned it on and why nobody argued.
And the documentation is not hiding anything. Compose describes the option in one sentence — the new
task is started first, and "the running tasks briefly overlap". The overlap is the feature. It is
stated plainly, in the same page as
replicas, and it reads as an
availability note because that is what it is.
What is missing is anybody connecting the sentence to the line below it. The overlap is harmless when the number is a capacity setting and fatal when it is an invariant, and nothing in the file says which one you have.
What one rollout actually costs
We built a harness for this rather than reason about it: one feed fanned out to every subscriber — a singleton consumer's subscription is a fan-out, not a work queue, so both instances receive everything — a rollout controller that runs either policy, and a shared downstream recording which instance acted on each event. Five trials per arm:
| Update policy | Events handled twice | Events handled by nobody |
|---|---|---|
start-first |
12 | 0 |
stop-first |
0 | 12 |
Twelve is not a magic number: it is the overlap divided by the publish interval, and the overlap is whatever your replacement takes to become ready. Ours becomes ready in 250 ms because it is a goroutine and a dial. A real one pulls an image, passes a health check and warms a cache, so every figure here is a floor — the same arithmetic with a bigger numerator.
The second row is the part worth sitting with, and it contradicts the advice you will find. The
standard operational recommendation for a singleton consumer that double-processes is to switch to
stop-first, and it is offered as a fix. It is not a fix; it is a swap. Duplicates go to zero and
are replaced, one for one, by events that no instance handled at all — because for the length of the
same window there is now nobody rather than everybody. Whether that is an improvement depends
entirely on which failure your system can absorb, and nothing in the compose file knows the answer.
Neither setting is the fix, and that is the finding
There is no third value. The dropdown has two options and they are the two failure modes: process something twice, or miss it. A file that has to choose between them is a file being asked a question it cannot answer, and the instinct to keep tuning it — a grace period here, a health check threshold there — is an attempt to make the window small rather than to make it not matter.
Shrinking the window is worth doing and does not close it. We swept the readiness wait to check that the proportionality is real rather than inferred from one point:
| Replacement ready in | Predicted | Handled twice |
|---|---|---|
| 50 ms | 2.5 | 2 |
| 100 ms | 5 | 5 |
| 250 ms | 12.5 | 12 |
| 500 ms | 25 | 25 |
| 1 s | 50 | 50 |
Five points on the line, so the lever works exactly as arithmetic says — and it reaches zero only
when the overlap does, which start-first is defined not to allow. You also buy each step with a
health check that passes earlier, which is to say with a replacement that is readier on paper than in
memory.
The fix is upstream of the deploy file, and it is one of two decisions, both of which are real engineering rather than configuration:
- Stop requiring the invariant. If the state the consumer holds can live outside the process — a cursor it re-reads, a lease it re-acquires — then two instances are no longer a correctness problem and the integer goes back to being a capacity setting.
- Make the second one harmless. If every effect the consumer produces is idempotent on the event's own identity, a duplicate becomes a no-op and the overlap costs nothing but work. This is the more common answer and it has its own failure mode, which is a piece of its own.
There is no third value
Where the invariant has to live instead
Whichever you choose, the thing to fix today costs an afternoon: make the invariant fail out loud where it is violated. Assert the count at boot from whatever the runtime exposes, and refuse to start if it is not what the code assumes:
deploy:
# CORRECTNESS INVARIANT, not capacity. This service holds its cursor in memory and
# acts on every event it sees, so a second instance double-processes. Boot asserts it.
replicas: 1
update_config:
order: start-first # documented to overlap the old and new tasks — see the assertPut the reason beside the number, in words, because the next person to read it is on call and is deciding whether raising it will help.
That does not make the rollout safe. It makes the assumption checkable, which is the whole difference between an invariant and a habit — and it turns the next incident from "why are we processing everything twice" into a container that refused to start and said why.
FAQ
Can I keep start-first and just make the overlap shorter?
Yes, and it is worth doing, but it is a mitigation rather than a fix. The duplicate count is
proportional to the overlap, so halving readiness halves it — and the count reaches zero only when the
overlap does, which start-first is defined not to allow. Its own documentation is explicit that the
tasks overlap; that is what you turned on.
Isn't stop-first the safe option, then?
Only if dropping events is safer than duplicating them for your system, which is a real question with
a real answer and is not the same answer everywhere. Our run shows the trade is one for one: the
events that stop being handled twice start being handled by nobody. Neither column is free, and
choosing between them without knowing which costs more is choosing at random.
Does this need two hosts or a cluster? No. Everything above is one machine and one process per instance. Scale changes the size of the window, not its existence.
Our consumer is behind a work queue rather than a fan-out. Same problem? Different, and this measurement does not cover it. A broker that routes work to whichever subscribers it believes in behaves differently under an overlap than one that fans everything out to all of them, and we did not test that. It needs its own harness before anyone quotes a number for it.
The three things worth taking away
An operational number can be a correctness invariant, and the file will not tell you which it is. The test is not what the number means; it is what the service holds. Anything with in-process state that acts on what it sees has an invariant written in a file owned by whoever is on call.
The two update policies are the two failure modes, measured one for one. Handle everything twice or miss some of it. There is no setting that does neither, and tuning between them is choosing which kind of wrong you prefer.
Make the assumption fail out loud. An assertion at boot and a sentence beside the number do not make the rollout safe — they make it audible, which is the difference between an invariant somebody wrote down and one everybody is remembering.
On a market feed the overlap is a consumer handling every tick twice for the length of a rollout — taken further under market feeds that reconnect without their symbols.
If nobody on your side can say out loud which of your services must never run twice, that is a cheap thing to have looked at.