An Update on Wayback Machine Access
144 points by ChrisArchitect 2 hours ago | 73 comments

simonw 2 hours ago
> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've ålso already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

reply
packetslave 2 hours ago
This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
reply
bsimpson 55 minutes ago
It's an open secret that you can often circumvent paywalls by searching Wayback.
reply
gambiting 52 minutes ago
Every single paid article linked on HN has the way back machine link as the very first comment.
reply
ValentineC 51 minutes ago
The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
reply
unkeen 46 minutes ago
*APIs
reply
stronglikedan 19 minutes ago
Yes, that's one acceptable alternative, and another commonly accepted alternative is API's. Although, I'm not sure why you included the asterisk.
reply
pantsforbirds 24 minutes ago
We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice!

Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.

reply
subarctic 19 minutes ago
What if they charged money? Is it something you'd pay for?
reply
bradly 37 minutes ago
Just yesterday from my one of my sessions with Sol:

> Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly

reply
TeMPOraL 27 minutes ago
As it should.

Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

reply
bradly 13 minutes ago
Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.
reply
aaron_m04 5 minutes ago
robots.txt?
reply
bradly 42 seconds ago
Has it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.
reply
RobotToaster 11 minutes ago
Do they offer bulk torrent downloads as an alternative?
reply
echelon 4 minutes ago
I would love to be able to download every page of a given domain as an archive, and I'd pay to do this.
reply
jader201 9 minutes ago
> I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

Appalling, yes. But also expected. I'm surprised they haven't been the target of scrapers for years. But sites putting their content behind login walls and other anti-bot mechanisms has certainly exacerbated this. But again, this isn't at all a surprising progression.

> we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

To be fair, another big motivation was likely users on sites like HN using archive.org (and similar sites) to get around their paywalls. In fact, I'd be surprised if this wasn't a big motivator.

Again, it sucks, but it's not at all surprising to see it progress like this. I wouldn't be surprised to see similar blocks on other archive sites eventually.

reply
luckylion 2 hours ago
What sites would they be targeting? Generic "just give me anything"? Whenever I check regular sites on IA, the coverage is spotty -- they'll have the homepage and a few important pages, but it quickly fizzles out.

Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you want everything.

reply
toomuchtodo 60 minutes ago
It is. They will most likely eventually need to move to a walled model for Wayback due to scraper aggressiveness (like Reddit deprecating anonymous old.reddit.com), or behind Cloudflare for aggressive bot and scraping protection. Hard to defend against abuse of a public resource when its intent is public access with as little restriction as possible.

https://en.wikipedia.org/wiki/Tragedy_of_the_commons

(no affiliation)

reply
ronsor 51 minutes ago
Reddit has no excuses for the anonymous old.reddit.com removal; they're simply greedy.

On the other hand, the Internet Archive is a non-profit offering a free public resource.

reply
toomuchtodo 48 minutes ago
Examples provided as technical examples, strong feelings are out of scope for this thread.
reply
itintheory 36 minutes ago
As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a temporary bandaid.

The cheapest solution is to require a login and rate limit by API key. I also have strong feelings about the tragedy of the commons.

[0] https://people.kernel.org/monsieuricon/creepy-crawlies

reply
msephton 2 minutes ago
I've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.
reply
basilikum 19 minutes ago
Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access.

The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.

If you got some money to spare, consider donating to them. They need it.

reply
superxpro12 8 minutes ago
fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future.

The future is bleak :\

reply
BeetleB 2 hours ago
Wow, but I wonder if there's more to it.

I've not been able to access web.archive.org from my work computer - I always get the 429 error.

But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

reply
flexagoon 59 minutes ago
I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP
reply
novok 12 minutes ago
Your workplace is probably redirecting traffic through a datacenter IP range. Especially if they have their own datacenters like google, microsoft, oracle, amazon, etc.

Try making a vpn via digital ocean for example and you'll see similar patterns.

reply
dotmanish 60 minutes ago
Could be due to some scrapers from either your work ISP block, or the larger block which lends IPs to multiple workplaces.
reply
lousken 50 minutes ago
AI companies should pay billions to wayback machine for access
reply
KPGv2 41 minutes ago
I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.
reply
roblh 34 minutes ago
Shouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?
reply
Joel_Mckay 22 minutes ago
That is essentially what LLM vector search results are, but the misappropriated $9Tn worth of FOSS code "AI" scraped and compacted for isomorphic plagiarism tokens is harder to prove now with watermarking skewed outputs. =3
reply
jMyles 29 minutes ago
It's time for copyright to end anyhow; that's what's gumming up the whole project in the first place.
reply
CqtGLRGcukpy 2 hours ago
> We’re getting better at telling abusive bots apart from the people who depend on the Wayback Machine every day. If you think you were blocked in error, email info@archive.org with your operating system, browser, and IP address, and we’ll look into it.
reply
timpera 2 hours ago
I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them.

Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.

reply
hubraumhugo 3 minutes ago
There is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreement. So how do we go from here? Is adding more difficult Anubis and Cloudflare bot protection really the solution? How many millions of human hours and billions in infra costs are we willing to spend on this arms race?

Some approaches that I think are promising:

- A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).

- Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.

- what else?

[0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada...

[1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...

reply
MattCruikshank 36 minutes ago
There was a feature on Amazon Web Services for a while, and I wish it was still there...

Downloader pays.

I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.

I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.

reply
vlyan 33 minutes ago
unrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?
reply
tech234a 2 hours ago
I wonder if they'll end up behind Anubis at some point. I'm surprised it hasn't happened already.
reply
stickfigure 43 minutes ago
Plenty of threads on HN about this, Anubis does not work.
reply
UltraSane 54 minutes ago
Why not put it in S3 with downloader pays?
reply
Kayvanian 46 minutes ago
As a public resource the hope is for Wayback to be free to access. I imagine putting up a paywall would be their last resort.
reply
Onavo 2 hours ago
Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon.

It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.

I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

reply
imglorp 39 minutes ago
Micropayments would solve so many Internet problems. It's not too late to adopt.

Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage.

The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.

reply
novok 9 minutes ago
Micropayments are blocked by government money laundering regulations increasing the costs significantly to make them untenable.
reply
drdexebtjl 2 hours ago
Sites would just block the Internet Archive crawler as well.
reply
KPGv2 38 minutes ago
> Why not just offer a paid endpoint for the crawlers?

Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.

reply
Ajedi32 18 minutes ago
What if you're not charging for the content, but as compensation for the network bandwidth / server resources consumed by serving that content? The idea isn't to profit from content (the IA is a nonprofit anyway), just to allow the IA to continue to serve its purpose as an archive of public data without being overwhelmed by bots.
reply
Onavo 13 minutes ago
Exactly, it's a question for the lawyers to sort out.
reply
xp84 60 minutes ago
My guess? Because even with a paid endpoint, the type of unscrupulous yahoo that is DDOSing IA today would probably still abuse the free endpoints because they can. The revenue that might come from a paid endpoint could help to scale up, but with how slow IA usually seems, I suspect there is an upper limit to how much traffic they can serve without a LOT more revenue.

This is a major "this is why we can't have nice things" situation in my opinion. IA is one of the most valuable gems of the Internet. The only thing that even comes close to preserving our shared history. The damage being caused (both by the effective DDOSing and by the knock-on impact that abuse has in encouraging publishers to remove their content from the archive) is incredibly serious.

reply
croes 2 hours ago
It’s one thing to archive other companies content, it’s another to sell the access to it
reply
faefox 2 hours ago
Yeah, who does the Internet Archive think it is, (insert literally any AI company here)?
reply
bonoboTP 48 minutes ago
Which AI company is selling access to reliable verbatim copies of websites? I don't mean "it may regurgitate a paragraph", but as a reliable service where you can repeatably get website content snapshots to a reliability level that makes such a use case viable?

Using the information for training purposes is not the same thing. Not legally the same and otherwise.

reply
Onavo 2 hours ago
That's for the lawyers to sort out, they have a lot of flexibility as a US nonprofit. The case law isn't that clear cut for this.
reply
simonw 2 hours ago
Internet Archive was almost destroyed by a copyright lawsuit from book publishers within the last few years. I expect they aren't excited to take on any additional risk of similar lawsuits right now.
reply
celsoazevedo 51 minutes ago
They need access to sites to archive them. It's already hard to do it as it is, imagine if they start selling access to content. They'd be shooting themselves on the foot, independently of what the law says.
reply
xp84 58 minutes ago
major [citation needed] on that. There are very limited exceptions to the massive power of copyright -- and they're mainly granted to libraries in the form of narrow waivers. And just the cost of fighting the most powerful copyright holders can bankrupt you -- especially if you're a relatively modestly-funded nonprofit.
reply
swingandamiss 2 hours ago
[flagged]
reply
kg 2 hours ago
Does xcancel scrape twitter? Isn't it more like a proxy for specific user requests to view tweets?
reply
knowaveragejoe 48 minutes ago
Correct, and nothing wrong with that.
reply
MadameMinty 44 minutes ago
"Kidnapping innocents bad but imprisoning criminals good?? Inconceivable!"
reply
yifanl 57 minutes ago
It's almost as if moral values aren't assigned universally.
reply
righthand 60 minutes ago
No one is upset that the AI companies are scraping the web, they’re upset how poorly implemented the scrapers, but the scraping itself is fine. Lots of people and businesses scrape the web.
reply
akerl_ 33 minutes ago
There are people commenting parallel to you saying they are upset about AI companies scraping the web.
reply
faefox 2 hours ago
Yes, anything that potentially costs Elon Musk money is objectively a good thing. :)
reply
dallen33 2 hours ago
Yeah cuz X is fucking shitty, why would I want to give them any traffic?
reply
xp84 52 minutes ago
Then... don't? If it sucks so much why do you need to read the tweets?

Great take: "This private website is owned by a man I don't like, so I refuse to pay for it - or even give it the possibility to monetize my traffic with ads!"

Still quite mainstream take: "... so I'll use an adblocker on it"

Immature take: "This private website that I hate and boycott is also an important part of our culture, but the posts on it are too important and valuable to ignore, so I'll use a proxy to scrape it"

reply
slig 50 minutes ago
You're giving them attention, thus validating their existence and their numbers.
reply
qwerpy 55 minutes ago
“It’s ok to do bad things to people/things I don’t like”

Feels good when you get to dish it out doesn’t it?

reply
xyst 44 minutes ago
> abusive bots

Are the abusive bots in the room with us?

reply
gooeyblob 11 minutes ago
What reason do you have to doubt the claim?
reply
alex1138 20 minutes ago
I mean there are people who have reported that with their own personal website Facebook's crawlers were essentially DDOSing them
reply