DEV Community

Cover image for We Version-Control Everything Except Our Own Memory
Sonia Bobrik
Sonia Bobrik

Posted on

We Version-Control Everything Except Our Own Memory

You know the moment. A five-year-old answer describes your exact bug, sends you to the official docs for the details, and the link lands on a glossy product page that has never heard of the API you need. No redirect, no archived copy, no apology. That small dead end is a symptom of something much bigger, and a recent deep dive into how the technology industry is erasing its own history traces the pattern in uncomfortable detail, from acquisitions that swallow engineering blogs to platform migrations that drop a decade of posts on the floor. The part that should sting for developers is who does the erasing. The web rarely forgets because of hackers or hardware failure. It forgets because of ordinary engineering decisions, shipped by people like us, often in a pull request titled "cleanup."

Most Pages Don't Die. Somebody Deletes Them.

In 2024, researchers pulled just under a million pages from Common Crawl snapshots taken between 2013 and 2023 and tried to load every one of them again. According to Pew Research Center's study of disappearing online content, 38% of the pages captured in 2013 were gone a decade later. Freshness offered little protection: 8% of the pages captured in 2023 had already vanished by that October.

The detail engineers should stare at is where the losses happened. In most cases the page disappeared from a website that was still online. The server answered; the page had simply been removed or moved without a forwarding address. Redirects are carrying more weight than most teams realize, too, since about a third of the links on news pages already bounce to a different URL than the one first published. Every one of those hops is a forwarding rule that somebody has to remember to keep alive. Money doesn't solve it either: the most-trafficked news sites were just as likely to contain broken links as the smallest ones. And Pew counted a page as dead only when it returned one of nine error codes that leave no room for doubt, so the real picture is probably worse.

The Link Still Loads. That's the Scary Part.

A 404 at least tells you the truth. The nastier failure is content drift, when the URL resolves but whatever you linked to has changed shape underneath you. Jonathan Zittrain, who teaches law and computer science at Harvard, made this point in his Atlantic essay on why the internet is rotting, drawing on research his team did with a dataset of New York Times articles published between 1996 and mid-2019. Those articles held more than two million outbound links. A quarter of the deep links were completely dead, and age made it brutal: 6% of links from 2018 had rotted, against 72% of links from 1998. When the team checked a sample of 4,500 links that still worked, 13% pointed to content that had drifted significantly from what the reporter originally cited.

Software teams manufacture drift every day without noticing. A GitHub link to a file on main promises that the highlighted lines will stay put, and the next refactor breaks that promise without a sound. A link to the "latest" docs changes meaning with every release. A README that says "see the wiki" outlives the wiki. None of these throw an error. They just start telling the reader something that is no longer true.

Now the Archive Itself Is Getting Locked Out

For most of the web's life, the safety net under all this decay has been the Internet Archive. Its Wayback Machine passed one trillion archived web pages in October 2025, the largest record of the web anyone has ever assembled. Right around the time that counter rolled over, the doors started closing.

In August 2025, Reddit limited the Wayback Machine to little more than its homepage, saying AI companies had been scraping Reddit data through the archive. At the end of 2025, The New York Times added the archive's crawler to its robots.txt and later confirmed it was hard-blocking it. The Guardian excluded itself from the archive's APIs and filtered its article pages out of the Wayback Machine. When Nieman Lab analyzed robots.txt files from a database of 1,167 news sites, it found 241 sites in nine countries disallowing at least one Internet Archive bot, and a follow-up in May 2026 counted more than 340 local US outlets limiting access. No publisher confirmed to Nieman Lab that an AI company had actually scraped its content through the Wayback Machine.

Read that with an engineer's eye and it looks like a textbook incident: a defense aimed at one threat takes down a shared dependency that everyone else relies on. It is the robots.txt version of firewalling your own backups. There's a bitter footnote, too. The newspaper whose outbound links became a landmark dataset for measuring link rot is now keeping its own pages out of the biggest archive built to fight it.

The Quietest Loss: Questions We Stopped Asking in Public

A stranger kind of forgetting is happening inside our own profession. Stack Overflow peaked in early 2014 with more than 200,000 questions a month. In December 2025 it received 3,862, a 78% drop from a year earlier. The slide started years before chatbots, but the cliff came after them. Developers didn't stop hitting bugs; they started solving them in private conversations with AI assistants built into their editors.

Those answers might be excellent. But a private chat has no URL, no public edit history, no accepted answer, no maintainer dropping in to say the behavior changed in v3, and no crawler will ever see it. The fix for that cursed driver bug now lives in one person's chat sidebar. When the next developer hits the same wall a few years from now, there will be nothing to find. We are trading a messy public commons for millions of private, unlinkable notebooks.

Deletion Is a Breaking Change

We have already run the experiment on what happens when shared history vanishes overnight. In March 2016, a developer unpublished left-pad, an 11-line npm package, and builds across the JavaScript world fell over because projects as large as Babel and React pulled it in somewhere down their dependency trees. npm restored the package and tightened its rules so that packages older than 24 hours could no longer be unpublished on a whim.

Nine years later, Google repeated the lesson at web scale. It announced that goo.gl short links would stop resolving on August 25, 2025, after years of steering developers toward Firebase Dynamic Links as the replacement, a service retired on that very same day. Under pressure, Google reversed course for links that still got clicks, but anything with no activity in late 2024 was switched off. Rarely visited old links are precisely what a historical record is made of.

A Preservation Checklist for People Who Ship Code

Tim Berners-Lee wrote in 1998 that cool URIs don't change. Nearly three decades later, it is still the most ignored style guide in our industry. Fixing that doesn't require a grant or a committee, only the habits we already apply to code, pointed at knowledge:

  • Give every moved page a forwarding address. Ship permanent 301s when docs or posts move, keep the redirect map in version control next to the code, and add a CI job that replays your legacy URLs and fails the build on any 404.
  • Link to commits, not branches. On any GitHub file view, pressing y swaps the address for a permalink pinned to the commit SHA, so the lines you highlighted can't wander away from your explanation.
  • Snapshot your sources, not just your code. When a design doc or blog post leans on an outside page, capture it with the Wayback Machine's Save Page Now and keep a local copy in a self-hosted ArchiveBox, since a single archive can be blocked; reach for Perma.cc when the citation is legal or academic.
  • Keep shorteners out of anything meant to last. READMEs, error messages, printed docs, and conference slides deserve full URLs on a domain you control.
  • Deprecate instead of deleting. Version your docs and leave old pages online with a banner pointing to the current release.
  • Write down the why. Commit messages, architecture decision records, and PR descriptions outlive the people who remember the context, so when you squash-merge, carry the PR description into the commit body.
  • Publish the answers you find in private. If an AI session or a 2 a.m. debugging marathon cracks something hard, spend ten minutes turning it into a post, an issue comment, or an FAQ entry.
  • Send your public repos to Software Heritage. Its Save Code Now form archives a repository along with its full history, and the archive already holds more than 22 billion unique source files.

Memory Is a Feature

On October 24, 2026, the Wayback Machine turns 25. It went public in 2001 with more than ten billion pages and now holds over a trillion, yet it only works if the rest of the web keeps its doors open and its URLs stable. The irony of our field is hard to miss. We built Git so that no change is ever truly lost, we argue about semantic versioning with theological intensity, and then we let whole decades of documentation evaporate in a CMS migration nobody thought to review.

Pick one item from that checklist and do it this week. Somewhere in 2031, a developer will land on your page instead of a 404 and never know how close it came.

Top comments (0)