Brian Reed: Hey — before we get into it, quick question: did you save anything this week? Like, deliberately, intentionally download something because you weren't sure it would be there later?
Eliza Ward: That's a strange opener. But — actually, yes? I screenshot a tweet because I had a feeling the account was going to go private. Which is a totally normal thing I do now that I barely even notice.
Brian Reed: Right, because on some level we all know the web doesn't keep things. We just don't say it out loud. Pew Research Center put a number on it — 38% of webpages that were accessible in 2013 were gone by the time they measured a decade later. Not hacked, not moved. Gone.
Eliza Ward: 38 percent. That's — hold on, that's not edge-case content either. That's the normal web, just evaporating.
Brian Reed: Because the web was never designed to last. That's the uncomfortable baseline. Platforms own their content, they can delete it, they can go bankrupt, they can pivot to video — whatever — and the public record just... goes with them. The landlord demolishes the building and everything inside is gone.
Eliza Ward: Which is what makes 1996 interesting — Brewster Kahle started the Internet Archive that year, and the Wayback Machine has now captured more than one trillion web pages. The scale of that is kind of staggering when you set it against the structural problem.
Brian Reed: One institution, doing this essentially alone, trying to compensate for a design flaw baked into the entire internet.
Eliza Ward: Kahle calls it 'Universal Access to All Knowledge' — and honestly once you hold that 38% number, it stops sounding idealistic and starts sounding like a very specific problem that someone decided to fight. The question is whether the fight is actually working.
Brian Reed: And that question — whether it's working — kind of hinges on how it actually captures things in the first place. Because the mechanism is weirder than I think most people realize.
Eliza Ward: Okay, here's the plain version. Imagine a librarian who walks through a city every day, photocopies every publicly posted flyer on every wall, and files it. No permission asked. The logic is: you put it on a public wall, that's your invitation.
Brian Reed: That's — yeah. That's actually the whole thing, isn't it.
Eliza Ward: The Internet Archive does that at HTTP level. It fetches whatever a web server returns to any browser — time-stamps it, stores it. No platform API, no login, no permission request. The crawler is called Heritrix, it's open-source, and it treats the public web exactly the way your browser does. If the page loads without credentials, it gets captured.
Brian Reed: Sites can opt out, right? Via a robots.txt file. Which is how Brewster Kahle frames this as implied consent rather than... just taking things.
Eliza Ward: Right — and that's where the analogy gets complicated. Because if you never knew robots.txt existed, did you actually consent to anything? Or did you just... not object loudly enough to count as objecting?
Brian Reed: Let me see — so the model is: publish publicly equals opt in, and you have to actively opt out. That's the implied license. But most site owners in, say, 1999 had no idea that lever existed.
Eliza Ward: And the infrastructure built around that assumption — the Archive runs on self-owned hardware called PetaBox, costs way below commercial cloud rates — it was designed from the start to sit outside any platform's control. Which means the whole thing only works if the implied consent model holds.
Brian Reed: And right now, that model is... actually being tested. Like, in ways it wasn't before.
Eliza Ward: Tested is — wait, actually that's understating it. The Archive tried to extend that implied-consent logic to books, through Open Library, and it broke completely.
Brian Reed: Web crawling and book lending are not the same legal territory at all.
Eliza Ward: Right — so Open Library used something called Controlled Digital Lending. The idea being: we own a physical copy, we lend a digital scan, one copy out at a time. Libraries do this. Except four major publishers disagreed. Hachette, Penguin Random House, HarperCollins, Wiley — they sued.
Eliza Ward: US district court ruled against the Archive in 2023. Appeals court upheld it — September 2024. More than 500,000 books removed from Open Library. And then the Archive said it wouldn't seek Supreme Court review, which means the legal question of whether Controlled Digital Lending is even legitimate just... stays open. No answer at the highest level.
Brian Reed: Hang on — so the thing that would have settled it, one way or the other, they decided not to pursue. Which means every library trying to do something similar is still operating in the dark.
Eliza Ward: And Kahle — I mean, he publicly mourned this. Called it the gutting of Open Library. That's not legal language, that's someone saying the mission just got hollowed out. Think about what that means practically: a library researcher in 2024, trying to verify a footnote from a book their institution can't afford to license, opens Open Library, and the link is dead. Book pulled under court order. That's a very specific kind of loss.
Brian Reed: The implied-consent model — it held for the web crawling side. But the moment they stepped into books, they were in contested copyright territory and it collapsed fast.
Eliza Ward: And that's before we even get to the structural question underneath all of this — which is that the DDoS attacks in 2025 exposed something the Hachette case only hinted at. That one's going to reframe everything we just said.
Brian Reed: And the DDoS attacks are — I mean, that's where the whole 'we're outside the platforms so we're safe' argument just falls apart. Because it doesn't matter that you're architecturally independent if you're still one institution on one budget.
Eliza Ward: That's the thing. The 2025 outages weren't a legal threat. No lawsuit. Just — adversarial traffic, and the Archive went down. One non-profit, minimal budget, single point of failure.
Brian Reed: So what's the backup?
Eliza Ward: Michigan Digital Preservation Network — MDPN — is storing some Archive-It web crawl data. But the sourcing on this is pretty clear: nascent. Which is... not a word that inspires confidence when we're talking about the primary copy of large chunks of the public record.
Brian Reed: Nascent meaning it exists, but we don't actually know if it works at scale yet.
Eliza Ward: Right — and then layer the robots.txt problem on top. Because now news outlets are blocking the Wayback Machine's crawler. Not because they oppose archiving — wait, this is the part that should be alarming — they're blocking it because they're scared of AI scrapers.
Brian Reed: So they can't tell the difference between the Archive and a model training pipeline, and the historical record just... quietly degrades as a side effect.
Eliza Ward: No court ruling fixes that. No architectural innovation fixes that. It's voluntary opt-out at scale, for a reason that has nothing to do with preservation.
Brian Reed: The institution built to outlast platforms is now facing something its whole design never anticipated — the sites themselves walking away. And there's no lever to pull on that.
Eliza Ward: And that's — I mean, you don't need a court ruling. You don't need a Hachette. You just need enough sites to quietly drop a robots.txt block, one at a time, and the Wayback Machine just... thins out. No verdict. No moment. Just fragmentation.
Brian Reed: The part I'm still sitting with is whether any of this was ever a solvable problem or whether we just... noticed it late. Like, the Internet Archive is one institution. One Brewster Kahle, one budget, one set of servers that went down in 2025 because someone aimed enough traffic at them. And the MDPN backup — nascent, right, that was the word — means the redundancy isn't actually there yet. So the question the whole episode keeps circling isn't really about the Archive surviving. It's whether we built a world where one institution's survival and the survival of the historical record are the same thing. And if they are — that's not durability. That's a very fragile bet.
Eliza Ward: You started us with: did you save anything this week. And I said yes, a screenshot, almost without thinking about it. Which — now that feels like the answer to your question. The infrastructure doesn't hold, so we do it ourselves, quietly, one screenshot at a time. That's not a system.
Brian Reed: No. It really isn't.
Eliza Ward: Thanks for thinking through this one with me. It's a hard place to land.