Onpode
Cover art for How the Wayback Machine became the durable record — the permission-less snapshot model

How the Wayback Machine became the durable record — the permission-less snapshot model

August 11, 2026 · 10 min

Eliza Ward & Brian Reed

38% of webpages accessible in 2013 had disappeared by 2023, according to Pew Research Center. The Wayback Machine — launched by Brewster Kahle in 1996 — has captured over one trillion pages using a permission-less crawling model, but copyright losses, DDoS outages, and news sites blocking its crawler now threaten that archive's completeness.

The Internet Archive, founded in 1996 by Brewster Kahle, is a non-profit organization that pioneered systematic web preservation by crawling publicly accessible pages at the HTTP level — fetching and storing whatever web servers return as time-stamped snapshots — without requiring permission from site owners, platform APIs, or user login credentials.

0:009:56
Get the next episode on unrealized & lost content

Follow it free — new episodes land in your feed.

Or make your own — any topic, in minutes
About this episode

The web forgets. That's not a bug anyone fixed — it's the original design. Pew Research found that 38% of webpages from 2013 had simply vanished a decade later. No catastrophe, no hack. Just platforms deleting, pivoting, going bankrupt, taking their content with them. The Internet Archive stepped into that gap in 1996, and the Wayback Machine has now captured more than a trillion pages using a permission-less crawl model: if a page loads publicly, it gets archived. The implied logic — publish openly and you've opted in, robots.txt exists if you want out — held for nearly three decades. This episode examines where that logic is cracking. The Archive's attempt to extend the same reasoning to book lending ended in federal court losses and the removal of over 500,000 titles from Open Library. The founder called it a gutting of the mission. DDoS attacks in 2025 exposed something the legal fights only hinted at: architectural independence from platforms doesn't help when you're still one institution on one budget. And now news publishers are quietly blocking the Archive's crawler — not out of opposition to preservation, but out of fear of AI scrapers they can't distinguish from Heritrix. The historical record thins out as collateral damage. What the episode sits with is the structural question underneath all of it: we may have accidentally made one non-profit's survival synonymous with the survival of significant chunks of the public record. That's worth understanding before we need to find out what happens if the bet fails.

Frequently asked

How does the Wayback Machine archive websites without permission?

The Wayback Machine uses an open-source crawler called Heritrix that fetches any page a browser can load without credentials — no API, no login, no permission request. The Internet Archive treats public publication as implied consent, and site owners can opt out only by adding a robots.txt directive explicitly blocking the crawler.

What percentage of old webpages have disappeared from the internet?

According to Pew Research Center, 38% of webpages that were accessible in 2013 were gone by the time researchers measured a decade later. The loss affects ordinary pages, not just fringe content, because the web was never architecturally designed to preserve what it publishes.

Did the Internet Archive lose its lawsuit over digital book lending?

Yes. A US district court ruled against the Internet Archive in 2023, and the appeals court upheld the ruling in September 2024. The case, brought by Hachette, Penguin Random House, HarperCollins, and Wiley, resulted in more than 500,000 books being removed from Open Library. The Archive declined to seek Supreme Court review.

Why are news websites blocking the Wayback Machine?

News outlets are adding robots.txt blocks that prevent the Wayback Machine's crawler from capturing their pages — not because they oppose archiving, but because they cannot distinguish the Archive's Heritrix crawler from AI model-training scrapers. The practical result is voluntary degradation of the historical web record, with no legal mechanism to reverse it.

Is there a backup for the Internet Archive if it goes down?

The Michigan Digital Preservation Network (MDPN) stores some Archive-It web crawl data, but sourcing describes the arrangement as nascent. When adversarial DDoS traffic took the Internet Archive offline in 2025, the redundancy infrastructure was not yet at a scale capable of serving as a reliable failover for the full public record.

Grounded in 11 sources
Digital memory at stake: News outlets block Wayback Machine - DW · dw.com
Internet Archive’s legal fights are over, but its founder mourns what was lost · arstechnica.com
The Internet Archive Turns 20: A Behind The Scenes Look At Archiving The Web · forbes.com
Link Rot and Digital Decay on Government, News and Other Webpages | Pew Research Center · pewresearch.org
What happens when the internet disappears? | The Verge · theverge.com
The Internet's Most Powerful Archiving Tool Is in Peril | WIRED · wired.com
The Long Now of the Web: Inside the Internet Archive's Fight Against ... · medium.com
​​An Approach to Backing up Internet Archive Web Crawls – bloggERS! · saaers.wordpress.com
Web Archive in 2026: What Has Changed and How It Affects Website Restoration · archivarix.com
The Hidden Costs of Digital Decay: Why It Matters · evolllution.com
The Long Now of the Web: Inside the Internet Archive’s Fight Against Forgetting · hackernoon.com
Read transcript

Brian Reed: Hey — before we get into it, quick question: did you save anything this week? Like, deliberately, intentionally download something because you weren't sure it would be there later?

Eliza Ward: That's a strange opener. But — actually, yes? I screenshot a tweet because I had a feeling the account was going to go private. Which is a totally normal thing I do now that I barely even notice.

Brian Reed: Right, because on some level we all know the web doesn't keep things. We just don't say it out loud. Pew Research Center put a number on it — 38% of webpages that were accessible in 2013 were gone by the time they measured a decade later. Not hacked, not moved. Gone.

Eliza Ward: 38 percent. That's — hold on, that's not edge-case content either. That's the normal web, just evaporating.

Brian Reed: Because the web was never designed to last. That's the uncomfortable baseline. Platforms own their content, they can delete it, they can go bankrupt, they can pivot to video — whatever — and the public record just... goes with them. The landlord demolishes the building and everything inside is gone.

Eliza Ward: Which is what makes 1996 interesting — Brewster Kahle started the Internet Archive that year, and the Wayback Machine has now captured more than one trillion web pages. The scale of that is kind of staggering when you set it against the structural problem.

Brian Reed: One institution, doing this essentially alone, trying to compensate for a design flaw baked into the entire internet.

Eliza Ward: Kahle calls it 'Universal Access to All Knowledge' — and honestly once you hold that 38% number, it stops sounding idealistic and starts sounding like a very specific problem that someone decided to fight. The question is whether the fight is actually working.

Brian Reed: And that question — whether it's working — kind of hinges on how it actually captures things in the first place. Because the mechanism is weirder than I think most people realize.

Eliza Ward: Okay, here's the plain version. Imagine a librarian who walks through a city every day, photocopies every publicly posted flyer on every wall, and files it. No permission asked. The logic is: you put it on a public wall, that's your invitation.

Brian Reed: That's — yeah. That's actually the whole thing, isn't it.

Eliza Ward: The Internet Archive does that at HTTP level. It fetches whatever a web server returns to any browser — time-stamps it, stores it. No platform API, no login, no permission request. The crawler is called Heritrix, it's open-source, and it treats the public web exactly the way your browser does. If the page loads without credentials, it gets captured.

Brian Reed: Sites can opt out, right? Via a robots.txt file. Which is how Brewster Kahle frames this as implied consent rather than... just taking things.

Eliza Ward: Right — and that's where the analogy gets complicated. Because if you never knew robots.txt existed, did you actually consent to anything? Or did you just... not object loudly enough to count as objecting?

Brian Reed: Let me see — so the model is: publish publicly equals opt in, and you have to actively opt out. That's the implied license. But most site owners in, say, 1999 had no idea that lever existed.

Eliza Ward: And the infrastructure built around that assumption — the Archive runs on self-owned hardware called PetaBox, costs way below commercial cloud rates — it was designed from the start to sit outside any platform's control. Which means the whole thing only works if the implied consent model holds.

Brian Reed: And right now, that model is... actually being tested. Like, in ways it wasn't before.

Eliza Ward: Tested is — wait, actually that's understating it. The Archive tried to extend that implied-consent logic to books, through Open Library, and it broke completely.

Brian Reed: Web crawling and book lending are not the same legal territory at all.

Eliza Ward: Right — so Open Library used something called Controlled Digital Lending. The idea being: we own a physical copy, we lend a digital scan, one copy out at a time. Libraries do this. Except four major publishers disagreed. Hachette, Penguin Random House, HarperCollins, Wiley — they sued.

Brian Reed: And won.

Eliza Ward: US district court ruled against the Archive in 2023. Appeals court upheld it — September 2024. More than 500,000 books removed from Open Library. And then the Archive said it wouldn't seek Supreme Court review, which means the legal question of whether Controlled Digital Lending is even legitimate just... stays open. No answer at the highest level.

Brian Reed: Hang on — so the thing that would have settled it, one way or the other, they decided not to pursue. Which means every library trying to do something similar is still operating in the dark.

Eliza Ward: And Kahle — I mean, he publicly mourned this. Called it the gutting of Open Library. That's not legal language, that's someone saying the mission just got hollowed out. Think about what that means practically: a library researcher in 2024, trying to verify a footnote from a book their institution can't afford to license, opens Open Library, and the link is dead. Book pulled under court order. That's a very specific kind of loss.

Brian Reed: The implied-consent model — it held for the web crawling side. But the moment they stepped into books, they were in contested copyright territory and it collapsed fast.

Eliza Ward: And that's before we even get to the structural question underneath all of this — which is that the DDoS attacks in 2025 exposed something the Hachette case only hinted at. That one's going to reframe everything we just said.

Brian Reed: And the DDoS attacks are — I mean, that's where the whole 'we're outside the platforms so we're safe' argument just falls apart. Because it doesn't matter that you're architecturally independent if you're still one institution on one budget.

Eliza Ward: That's the thing. The 2025 outages weren't a legal threat. No lawsuit. Just — adversarial traffic, and the Archive went down. One non-profit, minimal budget, single point of failure.

Brian Reed: So what's the backup?

Eliza Ward: Michigan Digital Preservation Network — MDPN — is storing some Archive-It web crawl data. But the sourcing on this is pretty clear: nascent. Which is... not a word that inspires confidence when we're talking about the primary copy of large chunks of the public record.

Brian Reed: Nascent meaning it exists, but we don't actually know if it works at scale yet.

Eliza Ward: Right — and then layer the robots.txt problem on top. Because now news outlets are blocking the Wayback Machine's crawler. Not because they oppose archiving — wait, this is the part that should be alarming — they're blocking it because they're scared of AI scrapers.

Brian Reed: So they can't tell the difference between the Archive and a model training pipeline, and the historical record just... quietly degrades as a side effect.

Eliza Ward: No court ruling fixes that. No architectural innovation fixes that. It's voluntary opt-out at scale, for a reason that has nothing to do with preservation.

Brian Reed: The institution built to outlast platforms is now facing something its whole design never anticipated — the sites themselves walking away. And there's no lever to pull on that.

Eliza Ward: And that's — I mean, you don't need a court ruling. You don't need a Hachette. You just need enough sites to quietly drop a robots.txt block, one at a time, and the Wayback Machine just... thins out. No verdict. No moment. Just fragmentation.

Brian Reed: The part I'm still sitting with is whether any of this was ever a solvable problem or whether we just... noticed it late. Like, the Internet Archive is one institution. One Brewster Kahle, one budget, one set of servers that went down in 2025 because someone aimed enough traffic at them. And the MDPN backup — nascent, right, that was the word — means the redundancy isn't actually there yet. So the question the whole episode keeps circling isn't really about the Archive surviving. It's whether we built a world where one institution's survival and the survival of the historical record are the same thing. And if they are — that's not durability. That's a very fragile bet.

Eliza Ward: You started us with: did you save anything this week. And I said yes, a screenshot, almost without thinking about it. Which — now that feels like the answer to your question. The infrastructure doesn't hold, so we do it ourselves, quietly, one screenshot at a time. That's not a system.

Brian Reed: No. It really isn't.

Eliza Ward: Thanks for thinking through this one with me. It's a hard place to land.

How the Wayback Machine became the durable record — the permission-less snapshot model · Onpode