Top
Best
New

Posted by ChrisArchitect 17 hours ago

An update on Wayback Machine access(blog.archive.org)
537 points | 272 commentspage 4
Onavo 16 hours ago|
Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon.

It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.

I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

oasisbob 13 hours ago||
> It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs

The problem with this perspective is that it ignores the victimization which is happening to all sorts of sites right now.

On one hand, you have content owners/suppliers which are trying to place restrictions on how much free bulk use is allowed.

When scrapers go to exotic lengths to evade the blocks, eg by using thousands of ephemeral IP addresses to collect an entire corpus, saying stuff like that makes it sound like it's all a wash.

"Oh, what a silly situation... How did we ever end up like this? It's not good for anyone ..."

No, there is a victim trying to defend themselves from rampant theft of resources, and a corporate asshole which doesn't care about the effects of their actions.

drdexebtjl 16 hours ago|||
Sites would just block the Internet Archive crawler as well.
imglorp 15 hours ago|||
Micropayments would solve so many Internet problems. It's not too late to adopt.

Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage.

The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.

novok 15 hours ago|||
Micropayments are blocked by government money laundering regulations increasing the costs significantly to make them untenable.
Analemma_ 15 hours ago|||
Micropayments would solve all the problems except for the problem that people absolutely loathe micropayments. Like, vein-popping furiously hate them.

Whenever the topic of micropayments for internet content comes up, a bunch of people start talking about payment processors and their floor on prices, and so on. That's not wrong, but it can be designed around and I think it's a scapegoat to avoid confronting the fact that users despise micropayments and we'd rather blame credit card companies for the lack of adoption.

mindcandy 14 hours ago|||
Micropayments would solve so many problems for the internet. And, cryptocurrencies would solve so many problems for micropayments. But, it's a non-starter because any proposal gets flooded with people popping veins about how crypto can't solve anything.
imglorp 13 hours ago|||
It doesn't need to be crypto, or payment processor based.

My ideal experience would be I load $20 into the browser somewhere like a wallet in one block (that could be a payment processor step). If I visit a participating page, it decrements my wallet $.01 or whatever.

The downside is the possibility of abuse and tracking by governments, which would have to be handled at the source, not the symptom.

landgenoot 3 hours ago||
I worked on this before. You can solve the privacy issue using "statistical payments", by lack of a better word.

You visit website A,A,A,B,C,D,A,A

At the end of the month, you send your entire 20$ randomly to one of the websites you visited.

This will level out everyone's contribution and reward websites with lots of traffic. It eliminates the need for micropayments.

xp84 16 hours ago|||
My guess? Because even with a paid endpoint, the type of unscrupulous yahoo that is DDOSing IA today would probably still abuse the free endpoints because they can. The revenue that might come from a paid endpoint could help to scale up, but with how slow IA usually seems, I suspect there is an upper limit to how much traffic they can serve without a LOT more revenue.

This is a major "this is why we can't have nice things" situation in my opinion. IA is one of the most valuable gems of the Internet. The only thing that even comes close to preserving our shared history. The damage being caused (both by the effective DDOSing and by the knock-on impact that abuse has in encouraging publishers to remove their content from the archive) is incredibly serious.

katatue 5 hours ago||
IA might be large enough to earn consideration, but generally scrapers just don't care about being good citizens. I work in the GLAM space and we offer OAI-PMH interfaces for the harvesting of our collections data - which doesn't stop companies from preferring to scrape our website for worse (less complete, less structured, less standardized) data instead.
KPGv2 15 hours ago|||
> Why not just offer a paid endpoint for the crawlers?

Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.

Ajedi32 15 hours ago||
What if you're not charging for the content, but as compensation for the network bandwidth / server resources consumed by serving that content? The idea isn't to profit from content (the IA is a nonprofit anyway), just to allow the IA to continue to serve its purpose as an archive of public data without being overwhelmed by bots.
Onavo 15 hours ago||
Exactly, it's a question for the lawyers to sort out.
croes 16 hours ago||
It’s one thing to archive other companies content, it’s another to sell the access to it
faefox 16 hours ago|||
Yeah, who does the Internet Archive think it is, (insert literally any AI company here)?
bonoboTP 15 hours ago||
Which AI company is selling access to reliable verbatim copies of websites? I don't mean "it may regurgitate a paragraph", but as a reliable service where you can repeatably get website content snapshots to a reliability level that makes such a use case viable?

Using the information for training purposes is not the same thing. Not legally the same and otherwise.

Onavo 16 hours ago|||
That's for the lawyers to sort out, they have a lot of flexibility as a US nonprofit. The case law isn't that clear cut for this.
simonw 16 hours ago|||
Internet Archive was almost destroyed by a copyright lawsuit from book publishers within the last few years. I expect they aren't excited to take on any additional risk of similar lawsuits right now.
quotemstr 13 hours ago||
They brought it on themselves by marketing a read-for-free product
celsoazevedo 15 hours ago||||
They need access to sites to archive them. It's already hard to do it as it is, imagine if they start selling access to content. They'd be shooting themselves on the foot, independently of what the law says.
xp84 16 hours ago|||
major [citation needed] on that. There are very limited exceptions to the massive power of copyright -- and they're mainly granted to libraries in the form of narrow waivers. And just the cost of fighting the most powerful copyright holders can bankrupt you -- especially if you're a relatively modestly-funded nonprofit.
halfblood_walks 11 hours ago||
[flagged]
sehw 6 hours ago||
[dead]
josefritzishere 13 hours ago||
[dead]
unkeen 15 hours ago||
[flagged]
tomhow 13 hours ago||
We detached this subthread from https://news.ycombinator.com/item?id=49716735 and marked it off topic.
stronglikedan 15 hours ago||
Yes, that's one acceptable alternative, and another commonly accepted alternative is API's. Although, I'm not sure why you included the asterisk.
maxrev17 15 hours ago||
Unkeen on the apostrophe that’s why! Gotta keep HN proper and correct guize
xyst 15 hours ago||
[flagged]
plorkyeran 14 hours ago||
If you're in a room with a TV then literally yes, there's a good chance there's an abusive bot in the room.
gooeyblob 15 hours ago|||
What reason do you have to doubt the claim?
alex1138 15 hours ago||
I mean there are people who have reported that with their own personal website Facebook's crawlers were essentially DDOSing them
swingandamiss 16 hours ago||
[flagged]
kg 16 hours ago||
Does xcancel scrape twitter? Isn't it more like a proxy for specific user requests to view tweets?
yifanl 16 hours ago|||
It's almost as if moral values aren't assigned universally.
knowaveragejoe 15 hours ago|||
Correct, and nothing wrong with that.
MadameMinty 15 hours ago||
"Kidnapping innocents bad but imprisoning criminals good?? Inconceivable!"
righthand 16 hours ago|||
No one is upset that the AI companies are scraping the web, they’re upset how poorly implemented the scrapers, but the scraping itself is fine. Lots of people and businesses scrape the web.
akerl_ 15 hours ago||
There are people commenting parallel to you saying they are upset about AI companies scraping the web.
righthand 6 hours ago||
Yeah I dont think they know why they think that.
faefox 16 hours ago|||
Yes, anything that potentially costs Elon Musk money is objectively a good thing. :)
dallen33 16 hours ago||
Yeah cuz X is fucking shitty, why would I want to give them any traffic?
xp84 16 hours ago|||
Then... don't? If it sucks so much why do you need to read the tweets?

Great take: "This private website is owned by a man I don't like, so I refuse to pay for it - or even give it the possibility to monetize my traffic with ads!"

Still quite mainstream take: "... so I'll use an adblocker on it"

Immature take: "This private website that I hate and boycott is also an important part of our culture, but the posts on it are too important and valuable to ignore, so I'll use a proxy to scrape it"

ImPostingOnHN 13 hours ago||
You're confusing the site for the content on it.

Some of the content is good, the site sucks and is run by a guy who seig-heils crowds.

Even if the content sucked, your post has big "you want to improve `X`, yet you participate in `X`"[0] energy.

0 – https://kitzy.com/content/assets/images/we-should-improve-so...

xp84 9 hours ago||
Twitter is not society. It's a privately-owned web site, nothing more. Always has been. Participating in it is giving your bogeyman power.
slig 15 hours ago||||
You're giving them attention, thus validating their existence and their numbers.
qwerpy 16 hours ago|||
“It’s ok to do bad things to people/things I don’t like”

Feels good when you get to dish it out doesn’t it?

msephton 15 hours ago|
I've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.
jolmg 14 hours ago||
> Asking users to email them with details of their OS, browser, IP address is just crazy.

It's surely to serve as data to help tell humans apart from bots.

> Changes made by IA shouldn't become my responsibility.

They're a free service. It's ultimately not their responsibility to service you either.

msephton 13 hours ago||
Imagine if Apple or Microsoft introduced a bug and said, ah yes we know about it we did that on purpose and we know it affects a huge number of people, if each of you could email us these details that'd be great. It's just such an insane request.

IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.

kjs3 12 hours ago|||
What is 'insane' here is the shear level of entitlement displayed here, including lumping a niche, free, volunteer supported service in with billion dollar, for profit corporations and demanding they pander to your inflated expectations.

Wild.

msephton 10 hours ago||
It doesn't matter who or what the service is, how much they have, or whatever else. They created a problem and now users have to pay for the inconvenience by emailing(!) specific details that could be captured automatically through web logs: OS, browser, IP address. It's ridiculous.
jolmg 8 hours ago||
> details that could be captured automatically through web logs

You can't be serious. Are you ok? The entire point is that they're trying to tell bots and humans apart. They're trusting email (and how you write your email) as a good signal that you're human. What are you talking about getting it from the log? The point is to correlate. How do you expect them to know who you are in the log unless you give them that info?

> They created a problem

No, they're dealing with a problem, and compromised that some human users may unfortunately get blocked.

> and now users have to pay for the inconvenience

You don't have to anything. You can just not use them. They don't owe you their service.

Somebody is handing out free apple lollipops, they ran out, compromised on giving grape ones, and now you're complaining you're being forced to eat a grape one and you don't like grape. Don't eat it.

msephton 7 hours ago||
Lollipops?
jolmg 12 hours ago||||
No, it's more like you're requesting something from them and they're telling you they may need some technical, non-personally-identifiable info from you to fulfill your request.
HDBaseT 10 hours ago|||
Who do you think the Internet Archive is? They are not Google, they have very low funding, very high expenses and are constantly under legal pressure.

The fact you can even access the Internet Archive for free is a result of tens thousands of human hours striving for one goal. Digital Preservation. If you rely so much on IA, you should consider donating.

msephton 10 hours ago||
I do donate. But that doesn't mean I have to thank them before every meal or think that the service is perfect.
HDBaseT 9 hours ago||
What do you expect them to do though? You have to be a reasonable person.
msephton 7 hours ago||
Data analysis would be a good start, better blocking heuristics, an off-the-shelf solution used by other organisations that don't have this problem, etc.
doctor_radium 9 hours ago||
OTOH I do appreciate their openness. It beats those times when I try visiting a site, only to get a cryptic 403 error or similar and no suggestion the site would like to hear from me.