Top
Best
New

Posted by ChrisArchitect 14 hours ago

An update on Wayback Machine access(blog.archive.org)
515 points | 265 comments
simonw 14 hours ago|
> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

Kodiack 8 hours ago||
I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests.

However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a small donation their way. They provide an incredibly valuable service and I love the benefit that I get from them just for personal side projects.

sippingabonedry 7 hours ago|||
How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks?

Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...

Kodiack 6 hours ago|||
They have their own ASN, which I’ve explicitly allowed requests from.

https://www.peeringdb.com/asn/7941

junon 1 hour ago||
TIL they have their own ASN. This is helpful, thanks.
fc417fc802 6 hours ago|||
I thought at least google (and possibly others) provided a way to verify the user agent?
usr1106 4 hours ago||
Sorry, not following. I thought the user agent is a string that the caller can set to anything. There is no immediate, reliable way to tell whether the string is correct.

Yeah, a real browser would produce certain patterns and never certain others. So in some cases one could clearly say it's not a human using a browser. But a scraper could also make efforts to mimic human browsing. Mostly the frequency of requests can tell with high probability tell it's not human. But such algorithms might occasionally give false positives for real users, exactly like it has obviously happened for the archive.

What google service are you referring to? Not sure whether the archove uses any of Google's tracking. I have pretty strong blocking of trackers and ads. But the archive works for me.

grumbelbart2 3 hours ago||
Google's scraper bot at least used to be behind IPs that you could identify via reverse-then-forward DNS. Not sure if that is still up to date, though.

https://developers.google.com/search/blog/2006/09/how-to-ver...

dewey 1 hour ago||
Yep, verifying the IPs is still the way to go. You often see websites that do it wrong when you set your user agent to Google Bot and they give you a different version of the page without validating that.
4thguy 1 hour ago||||
Kudos on that. I don't know what site you're hosting, but I appreciate knowing that it is there
mkatx 7 hours ago|||
This is the way to go! Cut the cat and mouse, win win ish.
packetslave 14 hours ago|||
This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.
bsimpson 13 hours ago||
It's an open secret that you can often circumvent paywalls by searching Wayback.
koolala 4 hours ago|||
One site was doing that which archive in their name but wasn't apart of archive.org
sam_lowry_ 4 hours ago||
archive.is or archive.today?

Why being shy in the era of stealing AI?

LoganDark 3 hours ago||
archive.today uses clients to perform DDoS, I would not recommend using their site.
schnebbau 3 hours ago||
Let's all just believe this baseless assertion shall we.

You have to include more than that, this isn't common knowledge.

jimmydorry 30 minutes ago|||
It's been posted multiple times of the last few months as they were removed from wikipedia and cloudflare. [1] [2] [3]

1. https://news.ycombinator.com/item?id=46843805

2. https://news.ycombinator.com/item?id=47092006

3. https://news.ycombinator.com/item?id=47474255

x______________ 1 hour ago||||
Sure it is! This has been going on for years and global attention was gained at the beginning of this one.[0]

Wikipedia deprecates Archive.today, starts removing archive links (arstechnica.com) 616 points by nobody9999 6 months ago | hide | past | favorite | 368 comments

0 https://news.ycombinator.com/item?id=47092006

LoganDark 2 hours ago|||
<https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidan...> for those without a search engine.
zymhan 8 hours ago||||
Only some of them, it is not universal.
ghostly_s 8 hours ago||
"often"
gambiting 13 hours ago|||
Every single paid article linked on HN has the way back machine link as the very first comment.
ValentineC 13 hours ago||
The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).
eek2121 10 hours ago|||
Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.
normie3000 9 hours ago|||
> Why folks continue to use them confuses me.

I use them. I haven't ever heard mention that the content is edited. Do you have a source?

uqers 8 hours ago|||
See

https://arstechnica.com/tech-policy/2026/02/wikipedia-might-...

https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidan...?

Besides tampering with content, the site was also using visitors to DDOS a blog that mentioned the owner of archive.today.

sam_lowry_ 4 hours ago||
[dead]
Mogzol 8 hours ago|||
See the "Background" section of the Wikipedia RFC on banning archive.today links: https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment...

They bulk replaced one string (a name) with another one across many archived pages, and added malicious code to all archive pages that would rapidly send requests to gyrovague.com in an attempt to DDOS them.

sam345 7 hours ago||
Just out of curiosity, how do we know that what is in the Wikipedia comments is accurate? I have no skin in the game. I was just wondering. Anybody can post anything on Wikipedia comments. I find it odd that Ars Technica would use that as a source. Maybe it's fine for gossip and speculation but it shouldn't be in Ars Technica then.
Mogzol 3 hours ago|||
There is a megalodon.jp archive of an archive.today page that shows the edits: https://megalodon.jp/2026-0219-1654-07/https://archive.ph:44...

There's also a bunch of previous hackernews discussions about it:

- https://news.ycombinator.com/item?id=47474255

- https://news.ycombinator.com/item?id=46624740

- https://news.ycombinator.com/item?id=47092006

- https://news.ycombinator.com/item?id=46843805

And you obviously have no reason to believe me, but I was following this when it was happening at the start of this year and can confirm that the DDoS script and archive text replacements really did happen.

fc417fc802 6 hours ago|||
Because a lot of us watched the drama unfold in real time.
fn-mote 7 hours ago||||
> Why folks continue to use them confuses me.

Seems like the “confused” is disingenuous when not paying for content you read is a clear motivation.

opello 4 hours ago||
Convenience as a higher order motivator than disgust at the bad behavior of archive.{today,ph,...} mentioned elsewhere, I think is the point of the comment to which you replied.
DaSHacka 9 hours ago|||
Ironically, your framing of the situation is infinitely more disingenuous to push a personal agenda versus anything the archive.today guy did.
petcat 12 hours ago|||
ehh it's a distinction without a difference. The point is that alternative links are available to circumvent paid access for anyone that wants them.
sandcat_ 12 hours ago|||
That isn’t the point being discussed. The point being discussed is that it’s bad form to abuse a service (archive.org) that is provided for free, for the public good in order to run commercial scraping operations.
petcat 12 hours ago||
It's bad form to scrape the scrapers?
sandcat_ 12 hours ago||
Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)
petcat 12 hours ago||
You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason.

The end result is exactly the same.

fc417fc802 6 hours ago|||
It is different, precisely because the end result is not the same - one broadly benefits the public while the other doesn't.

Substitute almost any disruptive public service to see the issue with your line of reasoning. For example - you seem to think that [ bulldozing private property ] to "construct an emergency fire break" is somehow different than [ bulldozing private property ] for any other reason.

Never mind that the sort of scraping being objected to is actually harmful to service health while what the wayback machine does is almost entirely unnoticeable.

DaSHacka 9 hours ago||||
The minuscule traffic generated by the wayback machine, which serves to preserve the content for years to come, is completely incomparable to the scrapers that hammer every single href linked on a website.
sippingabonedry 8 hours ago|||
You missed the XCancel flamewar yesterday. The consensus is we should be allowed to scrape data and bypass login walls if we dislike the site owner, or feel we are owed free access by arbitrary criteria, it's sort of an unwritten rule. Unless of course it's Google or Meta properties we're talking about because that might impact RSUs.
organsnyder 12 hours ago|||
They're different sites, with different goals, run by different people.
petcat 12 hours ago||
That provide the same functional service....

Hence, distinction without a difference.

celsoazevedo 12 hours ago|||
They are 2 different services, run by different people, one goes out of their way to bypass paywalls while the other doesn't, one is banned by Wikipedia and the other isn't, etc.

I think it's a distinction worth making.

Not to mention that the Wayback Machine itself isn't exactly a good tool to bypass paywalls as most paid sites don't let them archive paywalled content anyway.

fluffybucktsnek 12 hours ago||||
Given that the root of the discussion is about Internet Archive being hit with huge traffic and not the functionalities provided by Wayback Machine, it very much is a distinction with a difference.
petcat 12 hours ago||
Bot traffic or human traffic doesn't matter. The goal is to read websites without having your own access.

So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.

publlus_enigma 7 hours ago|||
I suspect you may be conflating two different things.

Archive.org exists to preserve historical snapshots of the public parts of websites, and not to bypass subscriptions or pay walls.

fluffybucktsnek 8 hours ago||||
Internet Archive's traffic may not matter to you, but that's the main topic of this discussion, regardless of what you care or use website archival tools for.
HDBaseT 8 hours ago|||
I think you have the wrong impression of the Internet Archive.

The internet archive is not designed to circumvent anything. It is not designed to "grant access without having your own access".

rpdillon 11 hours ago|||
Yeah, you're mistaken. One archives web pages, the other maintains a list of paid-access accounts and fetches information from behind paywalls as a service.
DaSHacka 9 hours ago||
Exactly this

archive.org is the more straight-laced archive that doesn't circumvent sites that try to block it, and removes content they deem 'problematic' even if not illegal or requested by the site owner.

Meanwhile archive.today/ph/is/etc is the guerrilla alternative run by a die-hard datahoarder that seeks to archive the information itself, bypassing whatever blockers/login pages/whathaveyou to achieve the result.

It's nice to have both options. When I archive a site, I usually use both for added resiliency.

pantsforbirds 13 hours ago|||
We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice!

Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.

subarctic 13 hours ago||
What if they charged money? Is it something you'd pay for?
bonestamp2 12 hours ago|||
I was thinking the same thing... paid access for high volume users or scrapers could actually help fund the non-profit. Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it. If news and other sites were getting paid, maybe they could go back to optimizing for good content instead of clicks.
progval 1 hour ago|||
They already provide this service at https://archive-it.org/archive-it/ though for some reason they don't seem to publicize it. Some info at https://help.archive.org/help/archive-it-information/ as well.
DaSHacka 9 hours ago|||
I think that would get into murky water really quickly with the rights holders (/content creators) not exactly being thrilled the Wayback Machine is essentially monetizing their IP behind their back.
usr1106 4 hours ago|||
Interesting, in all the years I have never noticed that IP has 2 meanings (well probably more...) Yeah, I am an engineer and usually try to avoid the legal BS. Although I hate that AI has made stealing legal if you are big enough.
g-b-r 9 hours ago|||
If the money was guaranteed to only be used to pay the costs, there probably wouldn't be any problems
usr1106 4 hours ago||
No. What happened to their remote library scheme? They did not make if for profit, but still...
msephton 12 hours ago||||
I'd pay for it, but only if they implemented the changes the community of users have been requesting for years.
carlosjobim 11 hours ago||
No matter what they did, you'd have a new excuse for why you won't pay.
msephton 11 hours ago||
Ah, the old ad hominem attack. How refreshing.

But anyway, no, I wouldn't keep finding reasons. I donate to them every year already. Somebody asked if I would be willing to pay and my answer was "yes, but".

It would need to be improved because certain aspects of it suck right now, not only the error this post is about. They only need go as far as their forums and github repos to see the community feedback.

jakderrida 10 hours ago|||
Is it really an ad hom if he doesn't know the hom?

Their reply is 100% based on the content of your post.

Dylan16807 9 hours ago||
If you make up a person to insult then yeah it's still ad hominem.
carlosjobim 8 hours ago|||
If you're donating, then you are evidently willing to pay without any "buts". So aren't you arguing against your own actions?
msephton 8 hours ago||
Not at all. I donate to Internet Archive, but we're talking here about paying for unobstructed access to but one part of their service: Wayback Machine. Two different things.
pantsforbirds 3 hours ago||||
I mean we already paid for the article from the source itself. I guess I'd expect a better "diff" source from them, but if they dont even update the article itself, i guess i wouldn't expect a paid service to have those updates either?
pantsforbirds 3 hours ago||
ah, i think i misunderstood your original post. if you mean the wayback-machine/arkive, then I suspect it'd be hard to justify? You are essentially paying a third-party source to validate that diffs didn't go through on the source material.

with llms, at some point it probably becomes easier to use your paid api connection to manage your own cached version yourself?

bee_rider 12 hours ago|||
I wonder if there would be concern on their part about appearing to be a company that was basically offering paywall circumvention as a product.
cloakley 11 hours ago||
It wouldnt be a paywall, more like an option for companies to not pay scrappers. At least the payment deviates to the source.
bradly 13 hours ago|||
Just yesterday from my one of my sessions with Sol:

> Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly

TeMPOraL 13 hours ago||
As it should.

Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).

bradly 13 hours ago||
Do you think there should be a way for a site to tell an agent it isn't allowed access? I'm not sure where I land on this exactly tbh, so no judgment cast.
aaron_m04 12 hours ago|||
robots.txt?
bradly 12 hours ago||
Has it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.
dhx 1 hour ago|||
robots.txt was only intended to help search index crawlers not get stuck in endless crawl loops for badly designed websites.

What you suggest is explicitly not a purpose of robots.txt per RFC9309[1]:

"These rules are not a form of access authorization."

HTTP 429 and HTTP 403 are what servers are meant to return to clients to slow them down or tell them to stop doing something without having first gained authorisation.

[1] https://datatracker.ietf.org/doc/html/rfc9309#section-1

bityard 10 hours ago||||
robots.txt applies (or should, in my opinion) to anything that automatically follows a link. Basically any software that is not a human-controlled web browser or single-shot curl command. Everything else: robot.
xena 11 hours ago|||
AI bros think they should be exempt from robots.txt. Administrators of big services beg to differ. No solid consensus has arisen. I bet it's gonna take a lawsuit or two to see how it shakes out.
ghaff 10 hours ago|||
From the start, robots.txt has always been an indicator of a site's preference with no actual legal significance.
Dylan16807 9 hours ago||||
wget ignores robots.txt outside of recursive mode. I think it's correct to do so, and I think an AI loading a handful of pages in response to a command should be about the same.
recursive 10 hours ago|||
If a new directive was introduced that allows for an explicit setting in robots.txt, do you think the bros would follow it anyway? Something like `ALLOW AGENTS` or `DISALLOW AGENTS`
Analemma_ 12 hours ago|||
I want agents to be able to act on my behalf, that’s the entire point. An agent should be able to do anything I can do sitting at my browser.
compiler-guy 11 hours ago|||
I suspect most people would be ok with this if they could only do it at the rate and frequency you yourself can do it. The problem is largely one of scale.
daveoc64 10 hours ago|||
Is scale what we're discussing though?

e.g. a prompt of "fetch <article URL> and summarise it for me" is very close to what a human would be doing with a web browser, and doesn't seem to involve any kind of scaling issue.

kelnos 2 hours ago|||
Sure, but all the time I'll ask Claude a question, and then I'll see it fetch 5-10 different URLs to come up with answer. I certainly would not be fetching those URLs at that rate if I were doing it myself. I would probably be visiting those pages, one by one, over the span of 10-20 minutes.

That's the scale argument.

compiler-guy 9 hours ago|||
It’s easy to write instructions that have the agent check once every fifteen minutes, or even once an hour, in perpetuity, which never sleeps. And people do write such instructions. A human can’t do that by hand for very long.

The problem is that it is hard to distinguish your one off (which seems perfectly fine) from the tidal wave of bad actors.

cruffle_duffle 11 hours ago|||
Then make agent friendly content. Take the text and make a markdown version.
fineIllregister 10 hours ago||
People doing this say it makes things worse because then the bots download both.
compiler-guy 10 hours ago||
Not to mention that it solves none of the rate issues. If the scrapers are hitting your site 10,000 times a day, adding markdown isn’t going to change that at all.
ryandrake 10 hours ago||||
Technically, even your browser is an agent. It says it in the HTTP: User-Agent. So is cURL. Every application the user runs is acting on the user's behalf.
cruffle_duffle 11 hours ago|||
Dunno why the downvotes. I feel that is reasonable as well. Owners that block that stuff are doing so only to their detriment.
matt_heimer 9 hours ago|||
I wonder if the entire internet is going to slowly move behind logins and allow lists for specific trusted crawlers at some point.

Open access doesn't seem sustainable.

But I might just grumpy about spending another hour this week adjusting rules to prevent bots.

intrasight 8 hours ago||
Rather than logins or regional filters, how about they just be a content provider to local libraries and perhaps use an app like Libby.
autoexec 10 hours ago|||
I've personally been using the Wayback Machine more often because I increasingly find myself being blocked from websites who are trying to keep out scrapers even though I'm just a regular person with JS disabled (along with a bunch of other stuff)
eek2121 10 hours ago||
Sites are getting too overzealous with blocking IMO. I got blocked for several hours by huggingface simply because my download didn't complete and I had to retry. It gave me error 429, suggested I login, and the login page wouldn't load because error 429.

A popular tech news site blocked my phone because of Apple Private Relay. That didn't last long because their traffic fell off a cliff when that happened.

Many sites are throwing more captchas at the problem, without understanding that captchas don't actually help with LLMs, they just hinder normal users and primitive scripts. LLMs solve captchas just fine.

Some big sites have put up improved paywalls. I'm fine with subscribing to a quality site, however, WSJ and all the other big media sites routinely spit out regurgitated garbage that can be had for free elsewhere (and due to political spin, their garbage is less valuable than the free versions of said content).

Some folks are declaring the internet dead. I wouldn't go that far, however, I will say that a reckoning is going to happen, especially when advertisers figure out that most ads served on basically every website are no longer viewed by humans.

simonjgreen 2 hours ago|||
Nearly every time a link is posted to HN to a site behind some form of wall, a high voted comment on the post will be a link to an archive site bypassing the owners wall. Bot owners are not the only ones routinely circumventing the choices of content owners.
RobotToaster 13 hours ago|||
Do they offer bulk torrent downloads as an alternative?
QuantumNomad_ 12 hours ago|||
Once upon a time some people explored backing up the Internet Archive.

However, that experiment ended. They mention there were some learnings and they then say:

> The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to a number of projects that are still in use.

https://wiki.archiveteam.org/index.php/INTERNETARCHIVE.BAK

I would really like to know if any sort of thing like that is still ongoing and if it’s accessible to people in general. Would be nice to mirror some data from IA to my local drives, for example via BitTorrent or IPFS, to have it for offline exploration and personal archive.

I know that individual items have torrents. And I’ve downloaded a few that way but always it ends up only using the “web seed” (i.e. the BitTorrent client is retrieving the files from IA via HTTP) because there are no one seeding some random single item I found. Plus, those torrents are unreliable sometimes because they include meta data files that were since updated but the torrent was not updated and so the web seed is giving the updated files that don’t match what the torrent says their hashes should be. So then you have to jump through some extra hoops to fix that and then resume the download, and all the while the HTTP connections to IA servers time out because their servers are overloaded. So when I say I wonder about possibilities of using BitTorrent I mean to retrieve whole collections of many items instead of individual ones, and with actual other peers instead of just having it put load on IA HTTP servers.

alightsoul 9 hours ago|||
The internet archive's decentralization project is paused as far as I can tell. They have too many things to do and too little funding to do it all. Their current strategy seems to be establishing new legal entities outside the us like in Canada and Switzerland, but they don't accept web traffic even though they hold full copies of the internet archive. There used to be a full copy in Egypt at the library of Alexandria and another in the Netherlands. Not sure if they're still in use, but they did accept web traffic. They hold a decentralized web camp every year in the middle of a forest
giantrobot 10 hours ago|||
The Internet Archive's torrents are a sick joke. I've yet to find one that actually manages to complete. They always get stuck at 90-something percent but that final blocks always fail verification and get retried, fail, and the process repeats forever. Because they're web seeds they're hitting IA infrastructure and not offloading to a real swarm. So their broken torrents are just screwing themselves.
echelon 12 hours ago|||
I would love to be able to download every page of a given domain as an archive, and I'd pay to do this.
msephton 12 hours ago|||
They provide a free cli tool to do this.
petcat 12 hours ago||||
isn't that what wget -m does? what is there to pay for?
carlosjobim 11 hours ago|||
You'd pay the domain owner for it? How much?
ezekiel68 8 hours ago|||
You might be right but -- why would they have watied until these recent weeks?
luckylion 13 hours ago|||
What sites would they be targeting? Generic "just give me anything"? Whenever I check regular sites on IA, the coverage is spotty -- they'll have the homepage and a few important pages, but it quickly fizzles out.

Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you want everything.

hedora 6 hours ago|||
My use of wayback has skyrocketed recently due to anti-bot measures.

I often cannot get past captchas, and archive.org is one of the fallbacks I try.

However, archive.is, etc are more reliable.

I wish the internet archive acted more like a library system, where multiple organizations could mirror the content.

They are a big single point of failure, and I’m shocked Trump/SCOTUS haven’t intentionally burnt the archives down yet.

jader201 13 hours ago|||
> I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

Appalling, yes. But also expected. I'm surprised they haven't been the target of scrapers for years. But sites putting their content behind login walls and other anti-bot mechanisms has certainly exacerbated this. But again, this isn't at all a surprising progression.

> we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

To be fair, another big motivation was likely users on sites like HN using archive.org (and similar sites) to get around their paywalls. In fact, I'd be surprised if this wasn't a big motivator.

Again, it sucks, but it's not at all surprising to see it progress like this. I wouldn't be surprised to see similar blocks on other archive sites eventually.

throwawayk7h 5 hours ago|||
Perhaps it would be sensible for the wayback machine to not show paywalled articles for the first, say, 3 months.
toomuchtodo 13 hours ago|||
It is. They will most likely eventually need to move to a walled model for Wayback due to scraper aggressiveness (like Reddit deprecating anonymous old.reddit.com), or behind Cloudflare for aggressive bot and scraping protection. Hard to defend against abuse of a public resource when its intent is public access with as little restriction as possible.

https://en.wikipedia.org/wiki/Tragedy_of_the_commons

(no affiliation)

ronsor 13 hours ago||
Reddit has no excuses for the anonymous old.reddit.com removal; they're simply greedy.

On the other hand, the Internet Archive is a non-profit offering a free public resource.

toomuchtodo 13 hours ago||
Examples provided as technical examples, strong feelings are out of scope for this thread.
itintheory 13 hours ago||
As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers. We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation. Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a temporary bandaid.

The cheapest solution is to require a login and rate limit by API key. I also have strong feelings about the tragedy of the commons.

[0] https://people.kernel.org/monsieuricon/creepy-crawlies

userbinator 3 hours ago|||
I say put your data up in torrents, host a few KB of plain HTML linking to them, and let decentralisation do the rest.
toomuchtodo 11 hours ago|||
No strong feelings here is what I meant. Certainly, that energy is best directed into aggressive countermeasures and defense in depth of public goods.

https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...

TZubiri 9 hours ago|||
Nah, there's actual value in hitting historical versions and with agents the gap between "how long has this product been offered by this company" and "I should go to wayback machine and do a binary search to find the earliest snapshot that contains this product offering " has closed.
e40 6 hours ago||
I say name and shame!
basilikum 13 hours ago||
Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access.

The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.

If you got some money to spare, consider donating to them. They need it.

ternaryoperator 11 hours ago||
I donate to them every year b/c I fully agree they’re doing a thankless critical job very well.
j79 10 hours ago||
Thank you for the inspiration! I just made my first donation.
niuzeta 8 hours ago|||
I've been donating $5 to them monthly for I don't know how long. I've only recently bumped it up to $25. They're the heroes of the internet age
superxpro12 12 hours ago||
fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future.

The future is bleak :\

mrguyorama 11 hours ago||
I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side.

It's going to suddenly be extremely valuable that wikipedia didn't settle for having a small rainy day fund and instead ceaselessly grabbed every fucking donation they could for two decades so they can fight such a legal battle.

jasonfarnon 9 hours ago|||
I don't think there's much grounds for a serious legal attack at this point, and also it would look so horrible from a PR standpoint no non-desperate would dare it. Anyway, google has been steering enough traffic away, now that its AI jumps in with an answer and the wikipedia link which 99% of the response is based on is always hidden with a bunch of other overlaid icons. Google has surely managed to kill off a bunch of reddit and stackoverflow traffic with this trick.
pvab3 10 hours ago||||
on what basis?
nephihaha 10 hours ago|||
Wikipedia is not a challenge to the system. If it were, it would be excluded from search results, much like blogs are now.
userbinator 3 hours ago||
The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic

Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that.

I knew something was up when a few alternative YouTube front-ends I use suddenly put up the 'nubis and complained about the high volumes of traffic they were getting flooded with; of course someone actually going after that data would be aiming their "AI bots" at YouTube directly instead of trying to suck it through a tiny little-known proxy-site, so it really strained the credibility of the argument.

robotmay 9 hours ago||
Unrelated, but this week I've been on a memory binge with the Wayback Machine, trying to find old content of mine from the early 2000s. Took me a while but I've finally put together a good bit of info about myself at the time that I'd completely forgotten, and it's all thanks to the Internet Archive storing my little gaming review website from when I was 16. I could barely remember any of the other stuff, it's been genuinely surprising figuring out what I'd forgotten. I couldn't even remember most domains I owned aside from one, which I used as the starting point.

Still can't remember what my Tripod site address was, but that might be lost to time.

Thank you, Archive.org.

BeetleB 13 hours ago||
Wow, but I wonder if there's more to it.

I've not been able to access web.archive.org from my work computer - I always get the 429 error.

But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

flexagoon 13 hours ago||
I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP
hedora 6 hours ago|||
Are there any decent/reputable residential proxy companies? I’m pretty sure I’ll end up needing one occasionally, for those days when my residential IP has a poor reputation score.
jcrawfordor 11 hours ago||||
There are definitely factors beyond IP being used. A week ago I found that all requests from Chrome-like browsers got a 429 across more than a half dozen networks and several machines, while Firefox reliably worked. I assume this was an overzealous policy on UAs.
BeetleB 11 hours ago||
In my work, it's failing on both Firefox and Chrome.
BeetleB 12 hours ago|||
No - my phone is not connected to work's WiFi.

Wonder who the bad actors in my company are...

flexagoon 11 hours ago|||
Doesn't have to be someone at your work, it could be a block on a whole ISP network or at least an IP subnetwork that is shared between many clients

You can try emailing the address mentioned in their post so they adjust their filters to match just the bot networks more precisely

iamacyborg 10 hours ago||||
It might just be your corporate VPN and whatever ASN it’s being routed through.
ButlerianJihad 11 hours ago|||
You should file a support ticket with your manager and the IT security or support desk. Show them the evidence of 429s that are blocking your assigned tasks during working hours. Also include the screenshots and files that you downloaded on your personal device in order to access your work-related materials. Be sure and thank them for adequately configuring the MDM on your personal mobile device so that you could do these work-related tasks. You should definitely also file an expense report to request reimbursement of your personal mobile bill, any data charges incurred, and the hours of networking or collaborating with external colleagues, while you were working on these work-related projects with your personal device.
BeetleB 7 hours ago||
Not sure if you're posting as a joke or sarcasm but accessing archive.org is not relevant to my work.
novok 13 hours ago|||
Your workplace is probably redirecting traffic through a datacenter IP range. Especially if they have their own datacenters like google, microsoft, oracle, amazon, etc.

Try making a vpn via digital ocean for example and you'll see similar patterns.

GetSMS 3 hours ago|||
My buddy said he could not access it even from a residential IP, it was blacklisted for some reason.
dotmanish 13 hours ago|||
Could be due to some scrapers from either your work ISP block, or the larger block which lends IPs to multiple workplaces.
giantrobot 9 hours ago||
They seem to be aggressively blocking IPv6 source IPs. I ran into this problem over the past month traveling. I got nothing but 429 errors until I switched on my VPN (which is IPv4 only) and magically the Wayback machine worked again. The lack of transparency on the part of IA is very frustrating.
timpera 13 hours ago||
I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them.

Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.

emaro 12 hours ago||
It's shame that the AI arms race causes such collateral damage. Free resources were always exploited, but the stakes ($T) and capabilities around AI allow unprecedented abuse. I wish we could go back... :/

I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.

zdragnar 12 hours ago|
I'm a little more skeptical that this is "AI is big so it is worse" issue. Yes, AI is big in scale, but this has been the case for almost every popular free service. They either start:

- charging (news / journalist services)

- gate-keeping (X forcing log-ins)

- enshittifying (lots of ads and degraded service)

The fact that the way back machine is incredibly useful but most people didn't know about it or use it very much doesn't change the fact that it has basically become very popular... only with LLM agents rather than humans. Ads alone aren't enough to support human traffic for many sites with human traffic.

roughly 1 hour ago||
Bonus points for anyone who’d like to guess how the tragedy of the commons was resolved in the times before the enclosure movement.
RobotToaster 32 minutes ago|
Torches and pitchforks?
CqtGLRGcukpy 14 hours ago||
> We’re getting better at telling abusive bots apart from the people who depend on the Wayback Machine every day. If you think you were blocked in error, email info@archive.org with your operating system, browser, and IP address, and we’ll look into it.
delis-thumbs-7e 3 hours ago|
I recently remembered a wonderful comic blog from 2010’s that is not online anymore. It was a sonderful Finnish LGTG-thened comic blog that I use to read, then forgot completely until few weeks ago. WM had it stored of course, so I could read through this amazing piece of internet art again.

I really so through some money their way, they do wonderful work.

More comments...