Top
Best
New

Posted by ChrisArchitect 15 hours ago

An update on Wayback Machine access(blog.archive.org)
515 points | 265 commentspage 2
1vuio0pswjnm7 5 hours ago|
"Here's what's going on."

Thank you

https://news.ycombinator.com/item?id=49571448

I had a feeling it was due to "AI" companies and developers using "agents"

Not surprised

pelican0 13 hours ago||
Is it established that the scraping scourge of late is primarily driven by AI companies? Anyone aware of any relevant studies?

Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.

userbinator 3 hours ago|
It's not. There is no evidence, just propaganda.
sicktriple 8 hours ago||
Anyone else feel like making a new internet and starting over
righthand 1 hour ago|
People will just bring their bots over and you'll be back to square one. Bots scraping existed before LLM companies decided to go nutso on the internet.
thimabi 13 hours ago||
I wonder why doesn’t the Internet Archive require logging-in prior to accessing the Wayback Machine. It would probably help them distinguish humans from bots, at a very little cost to humans.
extralongdivisi 12 hours ago|
Gatekeeping information is not the solution
hamandcheese 11 hours ago||
Why not? If its the difference between the information being available at all, then I choose login any day of the week.
extralongdivisi 10 hours ago|||
> If its the difference between the information being available at all

That's the point. The solution should avoid information not being available. Requiring login will incentivize bots to create spam accounts and move the battle to a new frontier, hurting real people in the process.

TechSquidTV 5 hours ago|||
This really only inconveniences people, not bots.
xacky 12 hours ago||
The anti virus industry needs to crack down on crawler and proxy malware, plus ISPs FINALLY need to replace CGNATs with iov6 to stop crawlers banning everyone behind a NAT.
sicktriple 6 hours ago||
the root of the root of all evil: NAT
HDBaseT 9 hours ago||
How exactly does IPv6 "stop crawlers".

If anything, it will make it harder to block due to the vastness of the IPv6 address space.

fulafel 2 hours ago||
CGNAT makes all ISP users appear to come from one v4 address, so blocking by v4 address becomes unworkable.
petterroea 7 hours ago||
I'd be happy to pay a 5$/month donation to get a higher rate limit/more lenient filter put on me
throwaway456754 2 hours ago|
If you get something, it isn't a donation.
ilamont 12 hours ago||
Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website?

My blogs are getting slammed and there are issues with cloudflare or captchas.

iamacyborg 11 hours ago|
> Shouldn't the solution be to gate bulk access for automated services for a price?

Fine in theory but determined scrapers will use residential proxies in bulk.

LastTrain 7 hours ago||
These should be illegal unless users sign off on every fucking byte.
potato-peeler 5 hours ago||
Wayback can’t be accessed through vpn, atleast on proton. Heck, most sites simply block you for using vpn.
xbar 4 hours ago||
Thank you for the Wayback Machine. It is immensely powerful for good.
msephton 8 hours ago|
Why can't they capture OS, Browser, and IP address at the time of error?

All that information is available at the point of failure, the user should not need to email it in.

ericpauley 8 hours ago|
Presumedly they collect that, but the vast majority of blocks are correct and not errors. This info allows them to look up the user’s request to label it as legitimate.
More comments...