Top
Best
New

Posted by zdw 1 day ago

Creepy Crawlies(people.kernel.org)
1239 points | 621 commentspage 11
bilater 19 hours ago|
Instead of trying to block why not monetize? So the proof of work can be directed at something you can be paid for (bitcoin mining)?
desterothx 38 minutes ago||
If everyone had devices that could do some compute worth paying for, people would be doing the work all of the time. The problem is actually useful POW is a lot more expensive compute-wise, making it non viable for normal users
inigyou 17 hours ago||
Monero would work better.
DrJThomasHusk 14 hours ago||
When a hapless user visits my site well

muahahahahah

Sorry, just the thought of it

But when they do… boy do I have a trap waiting for them.

My wife calls me The Genius. I’m the guy she calls when her battery dies or when her instagram breaks like when it shows that random guy in her DMs, stupid bugs LOL

I digress. Alas, when a user lands on my page. My page wants to know exactly 2 things:

1. Why are you here and who are you

And 2. Can you produce a working solution to Pharoah’s Fortune

…those of you aren’t familiar Pharoah’s Fortune is an old chestnut little poem, a riddle if you will I like to ask candidates and so far nobody’s solved it

And the reason nobody has solved it is Pharoah’s Fortune is a very tricky problem. It’s not something you can “solve” per se it’s more like you arrive there.

So far no one has solved it. They all fall for the same trick! It is of course what separates those who write elegant C versus those write poor quality JavaScript.

So I always say to my students to keep an open mind because you never know who - or should I say where you’re talking to.

I’m bookish.

phyzome 10 hours ago|
I don't see any relevant reference online to "Pharoah’s Fortune".
hnisjafx40 19 hours ago||
Learned this the expensive way
0xdeadbeefbabe 9 hours ago||
> permanently tying up a chunk of capacity spent on producing output that is only useful for a single purpose — feeding a learning model.

The horror.

pbronez 19 hours ago||
“Expect to lose some functionality, at least when accessing our resources anonymously.”

This seems fine to me. It would be a better world if we could have anonymous bulk data access. But if aggressive scrapers are bloating host costs, I’m fine with logging in.

Now, the flip side is that ONCE logged in, I want my bulk access. The worst of all worlds with when you demand authentication and then STILL block bulk access.

Case in point, I want to automatically download my Amazon and Target order records. This is easy to automate with playwright or whatever, but authentication stays annoying. My sessions expire quickly and I have to re-auth all the time. There should be an API to pull this data down.

api 20 hours ago||
The AI companies should have their AI fix their crappy inefficient crawler code.
bluedino 21 hours ago||
> But no, let's in fact choose the stupidest possible way of doing it — by rendering everything as HTML commit by commit and then parsing it.

I feel like I'm at work.

We had some web crawler using Selenium to make queries and scrape the data instead of just downloading the whole file.

Every day it seems like we have some people that know just enough to be dangerous creating things like that. And then of course it's our fault that things are slow, or we won't give them infinite system resources, etc

iririririr 20 hours ago||
anyone knows how Jwz solution is working?

dont click next link because he will show a nutsack image if the referrer contains hackernews. love the guy.

www.jwz.org/blog/2025/01/exterminate-all-rational-ai-scrapers/

basically, instead of blocking, he just poison it. and if a human sees it, it takes less effort to ignore the nonsense than it takes your pocket computer to deal with proof of work.

UltraSane 12 hours ago||
$1 dollar a year subscriptions would help.
oowa 17 hours ago|
have a hackathon to solve for this. OP says it's not a problem for him right now but if we extrapolate what he's talking about it's definitely a problem aaaaaaand It's totally solvable, Even with all of the crazy combinations he's talking about it's still solvable. And it's already been solved using patterns we see in streaming services. This is completely hackathonable. but why do we even need to bother with this? The slurpers are the cause of this, and they can cause this problem because of Murphy's Law. well you can only account for Murphy's Law with good architecture or something like that or whatever. Ha ha hackathon.
More comments...