Top
Best
New

Posted by yusufozkan 1 day ago

A single firm is behind OpenAI, Anthropic, and Meta hacking scandals(www.effort.news)
607 points | 209 commentspage 2
heaney-555 1 day ago|
"Behind" is doing a lot of work in this headline.
LPisGood 1 day ago||
One wonders if the publicity associated with the events in question were part of the sales pitch.
dylan604 1 day ago||
I'd venture a guess that OAI doesn't mind if the HuggingFace hack gets confused in the public's mind.
jaggederest 1 day ago||
My tongue in cheek immediate assumption was "so it's a guerrilla PR firm?"
EagleEdge 1 day ago||
What exactly did Irregular provide to Anthropic, test cases? I am so confused about this story.
ameliaquining 15 hours ago|
Third-party cybersecurity evaluations. See, e.g., https://www.irregular.com/research/assessing-gpt-5.6-sol, which was published around the same time.
xbar 10 hours ago||
Security-yolo-clowns play with dangerous toys and hurt people. Many used to get sent to jail.
eab- 1 day ago||
The fact that this firm makes such defective environments is certainly worthy of attention, and most likely a completely irreversible reputational loss; however, I found the framing in this article of 'therefore all the P(doom) stuff is a psyop, specifically in order to defend this company' to be completely unjustified and frankly a little insane?
nullbio 6 hours ago|
Why is it insane? It makes sense that Anthropic would want an insulating layer to do their dirty work and absolve themselves of culpability for distorting the facts.
mahboi 1 day ago||
Google is thinking man, we should've hired Irregular.
gjm11 1 day ago||
This seems pretty bullshitty to me.

The article says "A single firm, Irregular, is responsible for hacking done by all three companies" but I can't see anything in the article that actually justifies this claim. The nearest to that is the sentence immediately after that one: "Anthropic disclosed that Irregular was responsible for creating the tests ...". This is not, in fact, the same thing.

(Especially as, as aesthesia mentions, the article just happens not to mention that by "hacking done by all three companies" it doesn't mean, e.g., the most famous recent examples of such hacking: Irregular wasn't involved in the OpenAI/HuggingFace incident.)

So, so far as I can tell, the story is: OpenAI and Anthropic make AI models. Irregular does AI model evaluations. In some of Irregular's model evaluations, in which supposedly-sandboxed models attempted to break into simulated targets, the models got out of the sandbox and did bad things in the external world.

The article talks about "firms which instruct AI models to commit cyberattacks", which is a very neat bit of dishonest framing. It's true, in a sense, that Irregular instructed the models to commit cyberattacks -- inside their sandbox, against fictitious hosts. It's also true that the models actually did commit cyberattacks (e.g., the Hugging Face incident, though once again the attacks described by the article don't actually include this one). But it's not at all true that Irregular instructed the models to do anything like the bad things they actually did.

The article says '[Anthropic's] later disclosure shows that exactly zero percent of the agents went "rogue"'. Once again, the disclosure does not in fact show that. It shows that one variety of going-rogue could have been prevented by telling the models explicitly "this thing is real, not part of any kind of test, leave it alone". That is not the same thing.

The article claims that 'In the wake of these attacks, Anthropic and Irregular have deployed a swarm of AI Safety influencers paid by Anthropic-connected foundations to distract from their culpability and towards the baseless “rogue agent” theory.' It offers no actual evidence for this.

And the article seems very keen to highlight links between the companies involved and "effective altruism", though it is -- I assume deliberately -- rather vague about whether it's saying "of course we all know that EA is evil, so that shows that these companies connected to EA are evil" or "this incident shows how evil EA is".

The "Effort News" website has a number of other look-at-the-scary-Effective-Altruists stories on it. They also strike me as rather bullshitty.

... And then I look a bit further, and I see that Effort News's "about" page says "It all started when I was experimenting with using AI for financial auditing. I found stories that were crucial to the public’s right to know, including several of the stories now available at /investigations. I knew we had to sprint to the launch and launch a publication, directly applying this technology." and "The scope of what we can investigate has massively expanded, because we can chase 1,000 misses for one hit. But the final product cannot be slop. There’s plenty of slop on the internet. The way to surpass that, and what really matters, is manual curation and review of every finalized story."

Manual curation and review? I think the people behind Effort News are admitting that this is AI-generated "journalism". I expect that one day AI systems will be trustworthy journalists, but I personally am not very convinced that that day has yet come. And I don't see much reason why I should trust Brian Chau, the guy behind Effort News, to be doing everything possible to make his AI systems trustworthy journalists. It looks to me as if maybe they've been given instructions along the lines of "dig up things that make Effective Altruism look bad" for some reason.

(I don't mean to imply that EA is their only target. It's just one that jumped out at me.)

angry_octet 9 hours ago||
It's pure conspiracy thinking garbage.

I get that people don't like Israelis, but attributing any connection to an Israeli company as evidence of a conspiracy is nonsense.

linkregister 15 hours ago||
> Irregular wasn't involved in the OpenAI/HuggingFace incident.

It was [1]. It's understandable that you assumed it wasn't because the article didn't cite the sources on this claim. I agree with the rest of your points.

1. https://openai.com/index/third-party-cyber-evaluations-invol...

yorwba 14 hours ago||
From the article you linked: "Editor’s Note: These are separate from the Hugging Face security incident"
linkregister 14 hours ago||
Thank you for the correction.
classified 13 hours ago||
> Irregular describes “the agent itself becoming a threat actor”;

Talk about responsibility laundering. The state of affairs is reaching unprecedented levels of absurdity.

Centigonal 1 day ago||
The article contends that evaluations from Irregular helped prompt these incidents, because the prompts in the evals didn't tightly scope the systems to be evaluated or the methods to be used. It also contends that the faulty sandbox operated by Irregular is at fault.

They're probably right that having more defensively written prompts and a better sandbox could have prevented some of these incidents, but:

1. I don't think "well you didn't tell the model not to illegally hack third party organizations in your prompt" is a particularly convincing argument.

2. We don't know whether the blame for misconfiguring the sandbox lies with Anthropic or Irregular.

I'm thankful that this article is bringing up the supply chain of vendors to these labs, as that is often a place where significant sketchiness gets buried. However, the ideas that this is some Israeli EA conspiracy to hype up AI extinction risk seems unsupported by the facts to me.

blackqueeriroh 7 hours ago|
This was almost good until it veered into antisemitism
More comments...