MongoDB Write Concern, Journaling, and the Durability Illusion: What w:1 Actually Guarantees
Most Go services that talk to MongoDB treat w:1 as "write succeeded." That's a category error. w:1 is an acknowledgment that the primary received the write and applied it to its in-memory view. Whether that data survives a sudden primary crash before the next journal flush depends entirely on how you've combined write concern with journaling, and most production codebases get this combination wrong because the driver defaults are designed for throughput, not durability.
This article is about the mechanics underneath the driver call, the failure modes that only surface during ungraceful shutdowns or network partitions, and the design decisions that produce a coherent durability contract.
What the Acknowledgment Actually Means
MongoDB's write path moves through two distinct layers before data is durable on disk:
- In-memory (WiredTiger cache): The write is applied to the storage engine's in-memory B-tree. The primary can acknowledge here.
-
Journal (on-disk WAL): WiredTiger appends to its write-ahead log. Flush frequency is governed by
storage.journal.commitIntervalMs, which defaults to 100ms.
w:1 with no j field defaults to j:false in most driver versions prior to MongoDB 5.0, and to the server's own writeConcernMajorityJournalDefault in 5.0+. The critical distinction: without j:true, MongoDB acknowledges the write after step 1. If the primary crashes within the journal commit window (up to 100ms of writes), those acknowledged writes are gone. They were never in the replica set oplog durably; they exist only in the buffer.
This is not a theoretical risk. Any ungraceful shutdown—OOM kill, EC2 instance termination, kernel panic—creates exactly this scenario.
The Oplog Is Not a Safety Net Here
A common misconception: "the replica set will recover from secondaries." The oplog entry for a write is itself written through WiredTiger. If the primary crashes before flushing the journal, the oplog entry for that write is also lost. Secondaries cannot replicate what the primary never durably recorded. The write is gone from the replica set entirely.
w:majority without j:true has the same exposure at each node—majority of nodes acknowledged in memory. In a three-node replica set, two nodes can lose up to 100ms of writes simultaneously if both crash before their journal flushes.
The Correct Durability Combination
For writes that must survive primary failure:
writeConcern := writeconcern.New(
writeconcern.WMajority(),
writeconcern.J(true),
)
collOpts := options.Collection().SetWriteConcern(writeConcern)
coll := db.Collection("orders", collOpts)
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
result, err := coll.InsertOne(ctx, orderDoc)
WMajority() + J(true) means: a majority of voting members have written the entry to their journals before acknowledging. This is the only combination that provides crash-safe durability across the replica set. The tradeoff is latency: you now pay for two journal flushes on the primary plus network round-trips to secondaries before the call returns.
On NVMe-backed instances (EBS gp3 with provisioned IOPS, or instance store NVMe), this overhead is often 2–8ms per write. On gp2 or network-attached storage with high queue depth, it can exceed 40ms. Your timeout budget must account for this.
Read Concern and the Consistency Surface
Write concern governs what you acknowledge. Read concern governs what you read back. They interact:
-
readConcern: localreturns the latest data from the node you queried, which may include writes not yet replicated. On a secondary, this can return stale data or data that will be rolled back if the primary fails mid-replication. -
readConcern: majorityreturns data that a majority has acknowledged. Combined withwriteConcern: majority + j:true, this gives you linearizable-adjacent semantics for single-document operations without usingreadConcern: linearizable. -
readConcern: linearizableis the only option that guarantees you read your own writes across a failover, but it requires a read quorum check and can block if a majority isn't reachable.
For a payment or inventory system where a Go service writes a reservation and then immediately reads it to pass to a downstream service, using writeConcern: majority + j:true on the write and readConcern: majority on the read eliminates the window where a post-failover read could return a rolled-back state.
Transactions Raise the Stakes
Multi-document transactions in MongoDB use the same write concern and read concern machinery, but the failure surface expands. A transaction that commits with w:1, j:false has all the same crash exposure as a single-document write—but now the crash can leave a partial multi-document update unacknowledged that was never atomically applied. The transaction either committed or it didn't; the issue is that "committed" without journaling means "committed to memory."
In Go, the session-based transaction API accepts write concern at the session level:
sessionOpts := options.Session().SetDefaultWriteConcern(
writeconcern.New(writeconcern.WMajority(), writeconcern.J(true)),
).SetDefaultReadConcern(readconcern.Snapshot())
session, err := client.StartSession(sessionOpts)
if err != nil {
return fmt.Errorf("start session: %w", err)
}
defer session.EndSession(ctx)
_, err = session.WithTransaction(ctx, func(sc mongo.SessionContext) (interface{}, error) {
// multi-document ops here
return nil, nil
})
readConcern: snapshot inside a transaction gives you a consistent view of the data as it existed at the transaction's start timestamp. Without it, reads inside a long-running transaction can observe writes from other transactions that committed after your transaction began.
Operational Consequences of Getting This Wrong
Silent data loss after failover: The most dangerous failure mode. Your service logged a successful write. The primary failed. The write wasn't in the new primary's oplog. Downstream systems never received the event. No error was ever raised. Discovery happens when a user reports missing data days later.
Rollback during initial sync: A new secondary joining the replica set performs an initial sync. If the primary acknowledges a write that hasn't reached a majority, and then rolls back during that sync period, the rollback file captures the rolled-back documents—but your application already returned 200.
Timeout misconfiguration under durable writes: Teams switch to w:majority + j:true and set a 500ms operation timeout that was calibrated against w:1 on gp2 EBS. Journal flush latency under write load now causes spurious timeouts. The fix isn't to revert durability; it's to instrument journal flush p99 via db.serverStatus().wiredTiger.log and set timeouts against measured flush latency.
Choosing Write Concern by Workload Tier
Not every write in a system carries equal loss risk. A practical tiering:
| Write type | Recommended concern | Rationale |
|---|---|---|
| Financial transaction, order record | w:majority, j:true |
Loss is unacceptable; latency budget allows it |
| User-generated content, session state | w:majority, j:false |
Majority replication without full journal overhead |
| Metrics ingestion, event stream buffer | w:1, j:false |
High throughput; loss-tolerant; downstream system is source of truth |
| Cache warmup writes, ephemeral docs |
w:0 (unacknowledged) |
Fire-and-forget; no latency impact on caller |
Apply these at the collection or operation level in Go using options.InsertOne().SetWriteConcern(...), not globally on the client, so each write path carries an explicit durability contract rather than inheriting a blunt default.
Decision Framework
Before choosing a write concern, answer these four questions:
-
What is the cost of losing this write? If it's recoverable from an event stream, queue, or upstream system,
w:1or evenw:0may be appropriate. If it's the authoritative record,w:majority + j:trueis the floor. -
What is your p99 journal flush latency on the actual storage class? Measure this in staging under write load before committing to timeout values.
db.serverStatus().wiredTiger.logexposes flush counts and times. -
Do you read your own writes in the same request context? If yes, use
readConcern: majorityon the read or accept that a failover between write and read can return stale or rolled-back data. -
Are you inside a transaction? Enforce
w:majority + j:trueandreadConcern: snapshotat the session level. Don't let individual operations inside a transaction inherit a weaker concern from a client-level default.
The driver makes it easy to write code that looks correct and fails silently under the one operational condition—ungraceful primary failure—that you're most likely to encounter during a real incident. Durability in MongoDB is not a property of the replica set topology; it's a property of the write concern and journal configuration on every individual operation. Make that contract explicit in code.
Top comments (0)