DEV Community

Saqib Ameen Subhan
Saqib Ameen Subhan

Posted on Edited on

How to delete 2TB from a live MongoDB cluster without anyone noticing

The request was simple. "There is about 2TB of expired data in this collection. Can you delete it tonight?"

My answer was no, and being able to explain why is most of the job. A missing TTL index had let short lived data pile up for months in one of our busiest collections, and the team wanted to run one big deleteMany overnight. Here is what that would have actually done.

Why a big delete hurts, and it is not locking

Most people think a big delete is risky because it locks the collection. In modern MongoDB it does not, because WiredTiger works at the document level, so a huge delete does not freeze the collection. The damage comes from three other places.

1. The oplog fills up. Every deleted document becomes its own entry in the oplog (the log that secondaries copy from). Delete 500 million documents and you have written 500 million oplog entries, and every secondary has to copy and apply each one of them. Replication lag starts climbing. If a secondary falls further behind than the oplog can hold, it can no longer catch up and needs a full resync. At that point your delete has cost you a copy of your data.

2. The cache gets flushed out. To delete a document, WiredTiger first has to load it into memory. A 2TB delete pulls a lot of old data into the cache and pushes out the data your application actually uses. Queries that have nothing to do with the delete start getting slow, and whoever looks into it will probably never connect it back to the delete.

3. Majority writes slow down. If your writes use w:"majority" (and they should, I covered why in an earlier post in this series), every write now waits on the same secondaries that are already struggling to keep up with the delete. The problems add up.

I showed this to the team in a lower environment. I started the big delete, ran rs.printSecondaryReplicationInfo(), and let them watch the lag grow. After that, "tonight" turned into a plan.

The approach: small chunks, a pause, and watching the lag

const CHUNK = 10000;
const SLEEP_MS = 500;
const cutoff = ISODate("2026-01-01");

let total = 0;
while (true) {
  const ids = db.events.find({ created: { $lt: cutoff } },
                              { _id: 1 }).limit(CHUNK).toArray().map(d => d._id);
  if (ids.length === 0) break;

  db.events.deleteMany({ _id: { $in: ids } });
  total += ids.length;

  sleep(SLEEP_MS);   // give the secondaries time to catch up

  if (total % 1000000 === 0) print(`${total} deleted, check the lag`);
}
Enter fullscreen mode Exit fullscreen mode

Two rules we followed while running it.

Only run it in off peak hours. We ran it every night and stopped it before business hours. The full 2TB took around a week and a half, and nobody noticed anything, which was exactly the goal.

Let replication lag decide the speed. If the lag went above the level we were comfortable with, we increased SLEEP_MS. The speed of the loop is decided by how healthy the cluster is, not by how quickly we want it done.

The pause between chunks can look far too slow to developers. That is fine. Once the data has stopped growing, the delete has no deadline, and what stops the data from growing is the real fix below.

The real fix: a TTL index, so this never happens again

db.events.createIndex({ created: 1 }, { expireAfterSeconds: 2592000 })  // 30 days
Enter fullscreen mode Exit fullscreen mode

A few things about TTL indexes that surprise people.

  • The TTL process runs every 60 seconds, so documents are not deleted at the exact second they expire. That is fine for cleaning up old data, but do not rely on it for business logic that needs exact timing.
  • It deletes in batches in the background, which is basically the same chunked loop as above, running forever. On a very large backlog it can fall behind, which is why we cleared the 2TB by hand first and then let TTL handle it from there.
  • Put the TTL index on the field that actually decides when the data is old. A TTL index on the wrong date field will quietly delete data you wanted to keep.

We finished with a runbook covering the script, the lag limits, and the schedule, so the next time someone asks to delete 2TB, the answer is already written down.

Try it yourself

If you want to see this on your laptop, mdbkit lab gives you a 3 node replica set.

pip install mdbkit
mdbkit lab start          # 3 node replica set on 127.0.0.1:28110-28112
Enter fullscreen mode Exit fullscreen mode

Connect to the primary and load a million documents, about half of them older than 60 days.

// mongosh --port 28110
use shop
for (let b = 0; b < 100; b++) {
  const docs = []
  for (let i = 0; i < 10000; i++) {
    docs.push({ created: new Date(Date.now() - Math.random() * 120 * 86400000), payload: "x".repeat(200) })
  }
  db.events.insertMany(docs)
}
Enter fullscreen mode Exit fullscreen mode

Then run the chunked delete. This version prints the time, how many documents have been deleted so far, and how far behind the secondaries are.

function secondaryLag() {
  const s = rs.status()
  const p = s.members.find(m => m.stateStr === "PRIMARY")
  const lags = s.members.filter(m => m.stateStr === "SECONDARY")
                        .map(m => (p.optimeDate - m.optimeDate) / 1000)
  return Math.max(...lags)
}

const cutoff = new Date(Date.now() - 60 * 86400000)
let total = 0
while (true) {
  const ids = db.events.find({ created: { $lt: cutoff } }, { _id: 1 })
                       .limit(10000).toArray().map(d => d._id)
  if (ids.length === 0) break
  db.events.deleteMany({ _id: { $in: ids } })
  total += ids.length
  const t = new Date().toISOString().substr(11, 8)
  print(t + "  deleted " + total + " old documents   secondary lag " + secondaryLag() + "s")
  sleep(500)
}
print("done, deleted " + total + " old documents")
Enter fullscreen mode Exit fullscreen mode

You will see the lag stay close to zero the whole way through. If you want to compare, reload the data and run the same delete as one single deleteMany, then check rs.printSecondaryReplicationInfo() while it runs.

When you are done:

mdbkit lab destroy --yes
Enter fullscreen mode Exit fullscreen mode

When you have to delete a lot of data from a live system, going slowly is the safe way to do it.

Top comments (0)