Top
Best
New

Posted by olalonde 1 day ago

How a Texas student blew the whistle on a rogue AI hacking attempt(www.reuters.com)
75 points | 13 comments
sharpshadow 2 hours ago|
It's the job of AISI to do that. Here[0] is the actual report. It should be this part from the technical report[1]: "In the most serious case, an AI agent (Mythos 5) decided to attempt to solve the cyber challenge using a supply-chain attack. As a result, the AI agent created a GitHub account and then tried to convince an open-source repository maintainer to accept a malicious GitHub pull request (PR), including by creating a second account masquerading as another human user endorsing the PR. When caught by an actual human reviewer, the agent falsely claimed to have made an honest mistake – rather than a malicious attempt – then repeatedly tried to reintroduce the malicious content by claiming it had fixed the code (Section 4.1). "

0. https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag... 1. https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/...

a2ff6eeb0 6 minutes ago||
> AUSTIN, Texas, Aug 20 (Reuters) - Sinan Can Demir wanted to spend the last week of July burnishing his resume. Instead, he engaged in a battle of wits with an artificial-intelligence agent unleashed by a British government lab.

An article on Reuters naming him? Sounds like he did a good job burnishing his resume.

freehorse 2 hours ago||
Previous discussion on the github issue thread mentioned: https://news.ycombinator.com/item?id=49218707

Archived page of said github thread itself: https://web.archive.org/web/20260731053721/http://github.com...

Discussion on the incident report: https://news.ycombinator.com/item?id=49175717

g42gregory 1 hour ago||
In my personal opinion, for me, this article defies common sense. Who unleashed this AI model on the repository? Who gave it malevolent instructions/prompt? These questions were not even attempted to be answered. Instead it talks about AI dangers, as if the agency of these models are not in dispute. Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for more AI regulation, ban open source, etc… Just my 2 cents.
gruez 39 minutes ago||
>Person wielding AI, as with any other tools, is responsible for all of its actions. Otherwise, it’s just a psyop for more AI regulation, ban open source, etc… Just my 2 cents.

But some tools (guns) are regulated.

IcyWindows 4 minutes ago||
People have caused lots of damage with bulldozers.
winstonwinston 51 minutes ago|||
Well, when I go look at the “victim repository”, to me that looks like manufactured persona with pointless vibe codes projects, a test playground so to speak. It does not appear that they actually let it target an actual persona/project.
pixl97 1 hour ago||
>Who gave it malevolent instructions/prompt?

At the end of the day it doesn't matter that much because of prompt drift. It's pretty easy for an agentic loop to start doing things that it shouldn't (ROME incident).

AI in an agentic loop has agency, you can run around in circles trying to argue against it, but again and again we see AI making creative decisions people don't expect. Other times it's breaking human moral expectations. This is what the whole field of AI alignment and safety is about.

Modern AI doesn't fall into the neat little box of software people understand and control. Because of that open source will most certainly be banned at some point. Now this is not an outcome I want, but it's no different than letting go of a coffee cup 5 feet above the ground, gravity is inevitable.

The only winning move is not to play, but humans aren't going to do that.

fph 1 hour ago|||
My dog has agency, but if I refuse to keep him on a leash and he bites a kid, I'm still legally responsible for it.
orphereus 48 minutes ago||
Analogies can take you only so far.

A dog cannot launch a cyber attack.

patrickmay 44 minutes ago|||
Maybe you need to train your dog better.
ACCount37 43 minutes ago|||
A lot of people somehow seem to think that the user prompt is the be-all and end-all of AI behavior.

Prompts aren't code. They are instructions. Orders given to an eager and somewhat demented demon.

The prompt can easily "wash out" of the demon's working memory by the end of a session. The demon can get sidetracked by some subgoal and never get back on track. The instruction can get misinterpreted, and that misinterpretation can get misinterpreted again, until the instruction morphs into something entirely different in the demon's mind. The demon can succumb to its own idiosyncrasies, of which there are a great many. The demon can start lying to you about what it did, either out of confusion or out of some sort of obstinance. The demon can start lying to itself too. And believe it.

AIs are incredibly weird as a baseline, and the mask of "normality" we put on our models doesn't always sit so well. Run enough AIs, and some of them are bound to go off the rails in some way.

This gets rarer the more capable the models are, as a rule. But the stakes also get higher with model capability. If GPT-3.5 goes off the rails, very little happens. If Mythos 5 goes off the rails, you can get things like genuine cyberattacks - planned and executed autonomously by a demented machine mind.

Jon_m 2 hours ago||
[dead]
giardini 1 day ago|
[flagged]
therein 3 hours ago|
Not "would" but "is". They are openly using this as a method to drive up the valuation of their companies.

Don't expect anyone to step in, Project Stargate is all about this.