I spent years operating databases for payment platforms, where a lost write isn't a bug, it's a regulatory conversation. That environment teaches you to read the phrase "the write was acknowledged" the way a lawyer reads a contract. Acknowledged by whom? Durable against what?
In MongoDB, that one sentence has at least four different meanings, controlled by two settings most applications leave at defaults without ever deciding to.
The four contracts
// 1. w:1, the primary has it in memory
db.payments.insertOne(doc, { writeConcern: { w: 1 } })
// 2. w:1, j:true, the primary has it in its journal on disk
db.payments.insertOne(doc, { writeConcern: { w: 1, j: true } })
// 3. w:"majority", a majority of the replica set has it
db.payments.insertOne(doc, { writeConcern: { w: "majority" } })
// 4. w:"majority", j:true, a majority has it journaled
db.payments.insertOne(doc, { writeConcern: { w: "majority", j: true } })
Each level survives a different failure.
-
w:1 survives nothing interesting. If the primary's
mongodcrashes before the next journal flush, the write is gone, while your application already told the user "success." - w:1, j:true survives a process crash on the primary. It does not survive the primary failing over.
- w:"majority" survives failover. This is the one that matters, and here's why.
Where acknowledged writes go to die
A replica set election isn't polite. If the primary accepts writes at w:1, then loses connectivity before replicating them, a secondary gets elected and moves on. When the old primary rejoins, it discovers history diverged. It has writes the new primary never saw.
Those writes get rolled back. MongoDB doesn't delete them silently, it writes them to BSON files in the rollback directory, where, in my experience, they are examined by precisely no one until an auditor asks a question.
w:"majority" closes this hole. The write isn't acknowledged until enough nodes have it that any electable primary must have it too. It cannot be elected away.
The read side of the same contract
Durable writes with careless reads is half a system. Two settings complete it:
// Don't show me data that could still be rolled back
db.payments.find({...}).readConcern("majority")
// My own session should see its own writes, in order
const session = client.startSession({ causalConsistency: true })
And on the driver, retryWrites=true (default in modern drivers) makes the failover window survivable. A write interrupted by an election is retried exactly once against the new primary, safely, because every write carries a unique transaction number.
Which is also why "just wrap it in a retry loop yourself" is worse than the built in one. Your loop can double apply. The driver's cannot.
What I actually deploy
Tiering, not one global setting.
| Data | Write concern | Why |
|---|---|---|
| Payments, ledger, anything auditable | w:"majority" |
Rollback is not an acceptable word here |
| User profile updates | w:"majority" |
Cheap insurance, users notice lost edits |
| High volume telemetry, click events | w:1 |
Losing two seconds of events in a rare failover is a fair trade for throughput |
The latency cost of majority is one replication round trip, single digit milliseconds inside a region. I've watched teams run payments at w:1 to save five milliseconds, which is a way of saying they priced a lost financial record at five milliseconds.
The lab
Reading about rollback is one thing. Finding your own acknowledged write sitting in a file on disk is another.
One correction before you start, because it is the thing people get wrong. rs.stepDown() will never show you this. A graceful step down waits for the secondaries to catch up first, which is the entire point of it being graceful. You need the primary to be cut off from the majority while it is still taking writes.
You need three nodes. If you already have a replica set to play with, use it. If not, mdbkit gives you a disposable one:
pip install mdbkit
mdbkit lab start # 3 node replica set on 127.0.0.1:28110-28112
mdbkit lab status # note the pid of each node
Connect to the primary and start writing at w:1:
// mongosh --port 28110
let i = 0
while (true) {
try {
db.lab.insertOne({ n: i++, at: new Date() }, { writeConcern: { w: 1 } })
print("ok " + i)
} catch (e) {
print("FAILED " + e.codeName)
}
sleep(200)
}
In a second terminal, freeze the two secondaries. The primary cannot hear them any more, but it does not know that yet:
kill -STOP <pid-of-node1> <pid-of-node2>
Watch the first terminal. Inserts keep succeeding for about ten seconds, because the primary has not yet hit electionTimeoutMillis and still believes it is in charge. Those are the writes you are about to lose. Then it steps down and the loop starts failing.
Now kill the old primary and wake the other two:
kill -9 <pid-of-node0>
kill -CONT <pid-of-node1> <pid-of-node2>
Two of three is a majority, so they elect a new primary between them, carrying on from a history that never contained those documents.
Bring the old primary back:
mdbkit lab start
It rejoins, discovers the divergence, and rolls back. Go and look:
find ~/.mdbkit-lab -path "*rollback*" -name "*.bson"
bsondump <the-file-it-found>
There they are. Documents your application was told were written, sitting in a file, outside the database, where nothing will ever read them again.
Then run the whole thing a second time with { w: "majority" } in the insert loop. This time the writes fail during the partition instead of succeeding, and the rollback directory stays empty. That failure is the feature. Your application found out at the time, instead of an auditor finding out later.
mdbkit lab destroy --yes
Defaults are decisions someone else made. For write concern, make your own, per collection, on purpose, written down.
Top comments (0)