Top
Best
New

Posted by ChrisArchitect 16 hours ago

An update on Wayback Machine access(blog.archive.org)
528 points | 269 commentspage 3
tech234a 15 hours ago|
I wonder if they'll end up behind Anubis at some point. I'm surprised it hasn't happened already.
stickfigure 15 hours ago|
Plenty of threads on HN about this, Anubis does not work.
danbolt 11 hours ago|||
I’ve read a few different experiences with hosts having had success with Anubis to cut down on excessive scraping. One that comes to mind is the Dolphin project.[1]

I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit?

[1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...

stickfigure 9 hours ago|||
From a few weeks ago: https://news.ycombinator.com/item?id=49500040

Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated can easily afford to solve workloads higher than your users will tolerate.

You might stop casual scrapers, but you're not going to stop someone who cares. AI scraping companies care.

murderfs 7 hours ago|||
That's not even the fundamental problem. Even if the payload runs optimally in the browser, the cost of CPU is so small that it's basically irrelevant.

If you waste your user's time with something that would take a full minute to run on a datacenter core, you're costing the scraper something like $0.000005: 360 W TDP on a 128-core EPYC 9754 * $0.10/kWh. In reality, it'll be substantially less than that, because CPUs don't use 0W at idle.

The only way this would make any sense is if there were many more scrapers than users and scrapers cared more about latency than real users, but that's the exact opposite of reality. The entire endeavor is so fundamentally misguided that it almost seems like a psyop.

hedora 7 hours ago|||
[dead]
nubinetwork 2 hours ago|||
> Is there a chance you could clarify the “not working” bit?

I see anubis, 90% of the time I close the tab before it finishes.

ShadowOfThePit 1 hour ago||
But why?
nubinetwork 57 minutes ago||
Impatience and annoyance
autoexec 11 hours ago||||
It always seems to keep me, a normal human, locked out of any site that uses it.
phendrenad2 11 hours ago|||
Plenty of threads saying it works, too.
vlyan 14 hours ago||
unrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?
msephton 14 hours ago||
They get marked as inaccessible, but still exist in IA data
vlyan 13 hours ago||
is it possible to access somehow? it seems the site got excluded because of robots.txt set by some domain squatter, not manually.
alightsoul 10 hours ago|||
They have all the WARC file cataloged on their site outside the way back machine
msephton 12 hours ago|||
Get a job at IA?
tech234a 14 hours ago||
See also: https://wiki.archiveteam.org/index.php/List_of_websites_excl...

Note that Archive Team is separate from the Internet Archive.

int32_64 13 hours ago||
Are any AI companies using residential proxies to scrape?
xena 13 hours ago||
Yes. It's impossible to tell which because the split is residential proxies, dataset curators, and AI companies all being separate actors. However I fucking guarantee you it's out there and people are too cowardly to be honest about it so they don't get sued out of existence.
oasisbob 12 hours ago||
Oh yeah, absolutely.
tgtweak 12 hours ago||
Can't wayback machine just offer direct access to the archive for a premium and in doing so, pay for the service?
edelbitter 11 hours ago|
Not while the new dukes of the internet wielding massive armies of hijacked smart TVs have a better time browsing the web than I have; as a mere peasant with just a few IP addresses. There would be no reason to sign up and pay up for bulk access, unless open access is shut down.
MattCruikshank 15 hours ago||
There was a feature on Amazon Web Services for a while, and I wish it was still there...

Downloader pays.

I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.

I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.

brador 13 hours ago||
The only solution is to make visitors do compute. Compressing files for the archive to access other files would be perfect for this.

Cross verify hashes to prevent cheating.

Ez.

hubraumhugo 14 hours ago||
There is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreement. So how do we go from here? Is adding more difficult Anubis and Cloudflare bot protection really the solution? How many millions of human hours and billions in infra costs are we willing to spend on this arms race?

Some approaches that I think are promising:

- A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).

- Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.

- what else?

[0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada...

[1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...

edelbitter 11 hours ago||
- Find some new way for Cloudflare to acquire paying customers. If their business did not depend on the status quo, they would be exceptionally well positioned to roll out the technical & organizational frameworks that that make massive botnets a thing of the past.
maxrev17 14 hours ago||
Yeah it’s kinda crazy to me that what was once a back alley python script is now accepted as ‘fine, free for all’. The new era of bros really are smth else.
ignoramous 15 hours ago||
https://archive.vn/WWENg
UltraSane 15 hours ago||
Why not put it in S3 with downloader pays?
charcircuit 14 hours ago||
S3 price gouges on bandwidth.
Kayvanian 15 hours ago||
As a public resource the hope is for Wayback to be free to access. I imagine putting up a paywall would be their last resort.
lousken 15 hours ago|
AI companies should pay billions to wayback machine for access
KPGv2 15 hours ago|
I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.
roblh 14 hours ago|||
Shouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?
Joel_Mckay 14 hours ago|||
[flagged]
autoexec 11 hours ago||||
They wouldn't be paying for the content, just the bandwidth. Like buying a linux OS on a CD ROM was about the cost of media not profiting off of the software.
0xDEAFBEAD 9 hours ago||||
Isn't that already a big part of reddit's business model?
jMyles 14 hours ago|||
It's time for copyright to end anyhow; that's what's gumming up the whole project in the first place.
autoexec 11 hours ago||
I'd have a lot less of a problem with AI if everything that went into their training was public domain and made easily available to anyone for any use. It'd feel less like AI companies were just stealing the work of others and charging for it.
jMyles 10 hours ago||
Seems like a reasonable norm:

* If you train AI on it, you have to afford public access to it.

* Nobody can exact violence against anybody else in response to that person providing public access to any data anymore (ie, all bytestrings are public domain).

That's the world I'd like to try in the coming years.

More comments...