Wikipedia deprecates Archive.today, starts removing archive links

Posted by nobody9999 3 hours ago

Wikipedia deprecates Archive.today, starts removing archive links(arstechnica.com)

Archive.today is directing a DDoS attack against my blog - https://news.ycombinator.com/item?id=46843805 - Feb 2026 (168 comments)

Ask HN: Weird archive.today behavior? - https://news.ycombinator.com/item?id=46624740 - Jan 2026 (69 comments)

175 points | 98 comments

wuschel 21 minutes ago|

There is an post describing the possibility of an organised campaign against archive.today [1] https://algustionesa.com/the-takedown-campaign-against-archi...

How does the tech behind archive.today work in detail? Is there any information out there that goes beyond the Google AI search reply or this HN thread [2]?

[1] https://algustionesa.com/the-takedown-campaign-against-archi... [2] https://news.ycombinator.com/item?id=42816427

celsoazevedo 2 hours ago||

I don't see the point in doxing anyone, especially those providing a useful service for the average internet user. Just because you can put some info together, it doesn't mean you should.

With this said, I also disagree with turning everyone that uses archive[.]today into a botnet that DDoS sites. Changing the content of archived pages also raises questions about the authenticity of what we're reading.

The site behaves as if it was infected by some malware and the archived pages can't be trusted. I can see why Wikipedia made this decision.

jsheard 1 hour ago||

It's also kind of ironic that a site whose whole premise is to preserve pages forever, whether the people involved like it or not, is seeking to take down another site because they are involved and don't like it. Live by the sword, etc.

ddtaylor 1 hour ago|||

Did they actually run the DDoS via a script or was this a case of inserting a link and many users clicked it? They are substantially different IMO

dunder_cat 1 hour ago|||

https://news.ycombinator.com/item?id=46624740 has the earliest writeup that I know of. It was running it via a script and intentionally using cache busting techniques to try to increase load on the hosted wordpress infrastructure.

jsheard 1 hour ago|||

> It was running

It still is, uBlocks default lists are killing the script now but if it's allowed to load then it still tries to hammer the other blog.

dunder_cat 1 hour ago||

Ah good to know. My pi-hole actually was blocking the blog itself since the ublock site list made its way into one of the blocklists I use. But I've been just avoiding links as much as possible because I didn't want to contribute.

RobotToaster 45 minutes ago||||

Given the site is hosted on wordpress.com, who don't charge for bandwidth, it seems to have been completely ineffective.

Hamuko 12 minutes ago||

The speculation that I saw was that they'd try to get Wordpress.com to boot him off for being a burden on the overall infrastructure.

ddtaylor 1 hour ago|||

Thank you this is exactly the information I was looking for.

"You found the smoking gun!"

hexagonwin 1 hour ago|||

they silently ran the DDoS script on their captcha page (which is frequently shown to visitors, even when simply viewing and not archiving a new page)

jMyles 2 hours ago||

> Changing the content of archived pages also raises questions about the authenticity of what we're reading.

This is absolutely the buried lede of this whole saga, and needs to be the focus of conversation in the coming age.

basch 1 hour ago||

It seems a lot of people havent heard of it, but I think its worth plugging https://perma.cc/ which is really the appropriate tool for something like Wikipedia to be using to archive pages.

mroe https://en.wikipedia.org/wiki/Perma.cc

ronsor 1 hour ago||

It costs money beyond 10 links, which means either a paid subscription or institutional affiliation. This is problematic for an encyclopedia anyone can edit, like Wikipedia.

toomuchtodo 1 hour ago||

Wikimedia could pay, they have an endowment of ~$144M [1] (as of June 30, 2024). Perma.cc has Archive.org and Cloudflare as supporting partners, and their mission is aligned with Wikimedia [2]. It is a natural complementary fit in the preservation ecosystem. You have to pay for DOIs too, for comparison [3] (starting at $275/year and $1/identifier [4] [5]).

With all of this context shared, the Internet Archive is likely meeting this need without issue, to the best of my knowledge.

[1] https://meta.wikimedia.org/wiki/Wikimedia_Endowment

[2] https://perma.cc/about ("Perma.cc was built by Harvard’s Library Innovation Lab and is backed by the power of libraries. We’re both in the forever business: libraries already look after physical and digital materials — now we can do the same for links.")

[3] https://community.crossref.org/t/how-to-get-doi-for-our-jour...

[4] https://www.crossref.org/fees/#annual-membership-fees

[5] https://www.crossref.org/fees/#content-registration-fees

(no affiliation with any entity in scope for this thread)

RupertSalt 32 minutes ago||

If the WMF had a dollar for every proposal to spend Endowment-derived funds, their Endowment would double and they could hire one additional grant-writer

nine_k 21 minutes ago||

If the endowment is invested so that it brings very conservative 3% a year, it means that it brings $4.32M a year. By doubling that, rather many grant writers could be hired.

ouhamouch 1 hour ago|||

There are dozen of commercial/enterprise solutions: https://www.g2.com/products/pagefreezer/competitors/alternat...

also the oldest of that kind and rarely mention free https://www.freezepage.com

jsheard 1 hour ago|||

Does Wikipedia really need to outsource this? They already do basically everything else in-house, even running their own CDN on bare metal, I'm sure they could spin up an archiver which could be implicitly trusted. Bypassing paywalls would be playing with fire though.

toomuchtodo 1 hour ago|||

Archive.org is the archiver, rotted links are replaced by Archive.org links with a bot.

https://meta.wikimedia.org/wiki/InternetArchiveBot

https://github.com/internetarchive/internetarchivebot

jsheard 1 hour ago||

Yeah for historical links it makes sense to fall back on IAs existing archives, but going forward Wikipedia could take their own snapshots of cited pages and substitute them in if/when the original rots. It would be more reliable than hoping IA grabbed it.

toomuchtodo 1 hour ago||

Not opposed, Wikimedia tech folks are very accessible in my experience, ask them to make a GET or POST to https://web.archive.org/save whenever a link is added via the Wiki editing mechanism. Easy peasy. Example CLI tools are https://github.com/palewire/savepagenow and https://github.com/akamhy/waybackpy

Shortcut is to consume the Wikimedia changelog firehose and make these http requests yourself, performing a CDX lookup request to see if a recent snapshot was already taken before issuing a capture request (to be polite to the capture worker queue).

Gander5739 52 minutes ago|||

This already happens. Every link added to Wikipedia is automatically archived on the wayback machine.

RupertSalt 4 minutes ago|||

[citation needed]

toomuchtodo 52 minutes ago|||

TIL, thank you!

jsheard 1 hour ago||||

I didn't know you can just ask IA to grab a page before their crawler gets to it. In that case yeah it would make sense for Wikipedia to ping them automatically.

ferngodfather 1 hour ago||||

Why wouldn't Wikipedia just capture and host this themselves? Surely it makes more sense to DIY than to rely on a third party.

huslage 36 minutes ago|||

Why would they need to own the archive at all? The archive.org infrastructure is built to do this work already. It's outside of WMF's remit to internally archive all of the data it has links to.

RupertSalt 1 hour ago|||

Spammers and pirates just got super excited at that plan!

toomuchtodo 1 hour ago||

There are various systems in place to defend against them, I recommend against this, poor form against a public good is not welcome.

ChocMontePy 30 minutes ago||

I noticed last year that some archived pages are getting altered.

Every Reddit archived page used to have a Reddit username in the top right, but then it disappeared. "Fair enough," I thought. "They want to hide their Reddit username now."

The problem is, they did it retroactively too, removing the username from past captures.

You can see on old Reddit captures where the normal archived page has no username, but when you switch the tab to the Screenshot of the archive it is still there. The screenshot is the original capture and the username has now been removed for the normal webpage version.

When I noticed it, it seemed like such a minor change, but with these latest revelations, it doesn't seem so minor anymore.

xurukefi 1 hour ago||

Kinda off-topic, but has anyone figured out how archive.today manages to bypass paywalls so reliably? I've seen people claiming that they have a bunch of paid accounts that they use to fetch the pages, which is, of course, ridiculous. I figured that they have found an (automated) way to imitate Googlebot really well.

jsheard 35 minutes ago||

> I figured that they have found an (automated) way to imitate Googlebot really well.

If a site (or the WAF in front of it) knows what it's doing then you'll never be able to pass as Googlebot, period, because the canonical verification method is a DNS lookup dance which can only succeed if the request came from one of Googlebots dedicated IP addresses. Bingbot is the same.

xurukefi 25 minutes ago||

There are ways to work around this. I've just tested this: I've used the URL inspection tool of Google Search Console to fetch a URL from my website, which I've configured to redirect to a paywalled news article. Turns out the crawler follows that redirect and gives me the full source code of the redirected web site, without any paywall.

That's maybe a bit insane to automate at the scale of archive.today, but I figure they do something along the lines of this. It's a perfect imitation of Googlebot because it is literally Googlebot.

jsheard 23 minutes ago|||

I'd file that under "doesn't know what they're doing" because the search console uses a totally different user-agent (Google-InspectionTool) and the site is blindly treating it the same as Googlebot :P

Presumably they are just matching on *Google* and calling it a day.

xurukefi 8 minutes ago||

Sure, but maybe there are other ways to control Googlebot in a similar fashion. Maybe even with a pristine looking User-Agent header.

Aurornis 9 minutes ago|||

> which I've configured to redirect to a paywalled news article.

Which specific site with a paywall?

Aurornis 57 minutes ago|||

> I've seen people claiming that they have a bunch of paid accounts that they use to fetch the pages, which is, of course, ridiculous.

The curious part is that they allow web scraping arbitrary pages on demand. So if a publisher could put in a lot of arbitrary requests to archive their own pages and see them all coming from a single account or small subset of accounts.

I hope they haven't been stealing cookies from actual users through a botnet or something.

xurukefi 50 minutes ago||

Exactly. If I was an admin of a popular news website I would try to archive some articles and look at the access logs in the backend. This cannot be too hard to figure out.

elzbardico 1 hour ago|||

> which is, of course, ridiculous.

Why? in the world of web scrapping this is pretty common.

xurukefi 52 minutes ago||

Because it works too reliably. Imagine what that would entail. Managing thousands of accounts. You would need to ensure to strip the account details form archived peages perfectly. Every time the website changes its code even slightly you are at risk of losing one of your accounts. It would constantly break and would be an absolute nightmare to maintain. I've personally never encountered such a failure on a paywalled news article. archive.today managed to give me a non-paywalled clean version every single time.

Maybe they use accounts for some special sites. But there is definetly some automated generic magic happening that manages to bypass paywalls of news outlets. Probably something Googlebot related, because those websites usually give Google their news pages without a paywall, probably for SEO reasons.

mikkupikku 28 minutes ago||

Using two or more accounts could help you automatically strip account details.

xurukefi 23 minutes ago||

That's actually a really neat idea.

tonymet 1 hour ago|||

I’m an outsider with experience building crawlers. You can get pretty far with residential proxies and browser fingerprint optimization. Most of the b-tier publishers use RBC and heuristics that can be “worked around” with moderate effort.

quietsegfault 49 minutes ago||

.. but what about subscription only, paywalled sources?

layer8 43 minutes ago||

It’s not reliable, in the sense that there are many paywalled sites that it’s unable to archive.

xurukefi 38 minutes ago||

But it is reliable in the sense that if it works for a site, then it usually never fails.

bjourne 47 minutes ago||

FYI, archive.today is NOT the Internet Archive/Wayback Machine.

tl2do 1 hour ago||

Why not show both? Wikipedia could display archive links alongside original sources, clearly labeled so readers know which is which. This preserves access when originals disappear while keeping the primary source as the main reference.

bawolff 1 hour ago||

The objection is to this specific archieve service not archiving in general.

ranger207 1 hour ago||

They generally do. Random example, citation 349 on the page of George Washington: ""A Brief History of GW"[link]. GW Libraries. Archived[link] from the original on September 14, 2019. Retrieved August 19, 2019."

Gander5739 50 minutes ago||

This will always be done unless the original url is marked as dead or similar.

RupertSalt 1 hour ago||

"Non-paywalled" ad-free link to archive: https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment...

chrisjj 3 hours ago||

> an analysis of existing links has shown that most of its uses can be replaced.

Oh? Do tell!

that_lurker 1 hour ago||

I would be suprised if archive.today had something that was not in the wayback machine

chrisjj 1 hour ago|||

Archive.today has just about everything the archived site doesn't want archived. Archive.org doesn't, because it lets sites delete archives.

bombcar 1 hour ago||||

Wayback machine removes archives upon request, so there’s definitely stuff they don’t make publicly available (they may still have it).

zahlman 1 hour ago||||

Trying to search the Wayback machine almost always gives me their made-up 498 error, and when I do get a result the interface for scrolling through dates is janky at best.

ribosometronome 1 hour ago|||

Accounts to bypass paywalls? The audacity to do it?

that_lurker 1 hour ago||

Oh yeah those where a thing. As a public organization they can't really do that.

I personally just don't use websites that paywall important information.

nobody9999 3 hours ago||

>> an analysis of existing links has shown that most of its uses can be replaced.

>Oh? Do tell!

They do. In the very next paragraph in fact:

   The guidance says editors can remove Archive.today links when the original 
   source is still online and has identical content; replace the archive link so 
   it points to a different archive site, like the Internet Archive, 
   Ghostarchive, or Megalodon; or “change the original source to something that 
   doesn’t need an archive (e.g., a source that was printed on paper)

chrisjj 3 hours ago||

Well, that's an odd idea of "can be replaced".

> editors can remove Archive.today links when the original source is still online and has identical content

Hopeless. Just begs for alteration.

> a different archive site, like the Internet Archive,

Hopeless. It allows archive tampering by the page's own JS and archive deletion by the domain owner.

> Ghostarchive, or Megalodon

Hopeless. Coverage is insignificant.

Kim_Bruning 2 hours ago|||

> archive.today

Hopeless. Caught tampering the archive.

The whole situation is not great.

nobody9999 2 hours ago|||

I just quoted the very next paragraph after the sentence you quoted and asked for clarification.

I did so. You're welcome.

As for the rest, take it up with Jimmy Wiles, not me.

ChrisArchitect 2 hours ago|

Previously Related:

Archive.today is directing a DDoS attack against my blog?

https://news.ycombinator.com/item?id=46843805

input_sh 59 minutes ago|

I know I'm arguing with a bot that nobody monitors, but it's already in the fucking post.

More comments...