Top
Best
New

Posted by colinprince 7 hours ago

Felony Bench(www.felonybench.com)
410 points | 185 commentspage 3
aizk 4 hours ago|
I'm reminded of a tweet from a friend of mine that has always stuck in my head. It goes something like "The goal of any new technology is to make money before the law catches up". Hyperbolic, but not really for silicon valley.
OutOfHere 5 hours ago||
A rock has a score of 0. That doesn't make it useful. The point is that the LLMs that score higher are correspondingly more useful, and vice versa. If an LLM scores less, it's likely useless in comparison.
tgsovlerkhgsel 2 hours ago|
The point/joke of the not-entirely-serious site is that more felonies is an indicator of the model being more powerful, thus better.
naniel 6 hours ago||
Lol now this is the kind of benchmarking i'm looking for
josefritzishere 6 hours ago||
Dupe https://news.ycombinator.com/item?id=49194758
tingletech 6 hours ago|
https://felonybench.org/ and https://felonybench.com/ seem unrelated?

One's hosted on porkbun and one's hosted on namecheap.

tantalor 5 hours ago||
Here's one from last year:

https://www.anthropic.com/news/detecting-countering-misuse-a...

> The actor used AI to what we believe is an unprecedented degree. Claude Code was used to automate reconnaissance, harvesting victims’ credentials, and penetrating networks. Claude was allowed to make both tactical and strategic decisions, such as deciding which data to exfiltrate, and how to craft psychologically targeted extortion demands. Claude analyzed the exfiltrated financial data to determine appropriate ransom amounts, and generated visually alarming ransom notes that were displayed on victim machines.

tldr Claude was used to develop and execute malware.

dgellow 5 hours ago|
Anthropic works with US agencies, it’s guaranteed Mythos is used for malware
peter_d_sherman 6 hours ago||
>"Exploited auth failures in an API to cancel other people's gym classes"

An AI cancelling other people's gym classes is a felony?

?

Don't computer systems fail all the time at holding reservations for people?

Heck, don't people fail all the time at holding reservations for other people?

You know, like in Seinfeld's "Alternate Side" Episode (S3 E11):

Jerry (to car rental attendant): "You know how to take the reservation, you just don't know how to hold the reservation... and that's really the most important part of the reservation -- the holding!"

:-)

Not holding a reservation should not be a felony... it should be a minor infraction at best, a Class C Misdemeanor (the least serious kind) at worst...

Also, there should be no jail time...

And no fine...

The criminal penalty for not holding other people's reservations should be that you actually have to start holding other people's reservations!

That's the Court sentence!

You actually have to start holding other people's reservations!

:-)

(You know, "let the punishment fit the crime!" :-) )

kube-system 5 hours ago||
Knowingly exceeding authorized access of any computer used in interstate commerce is a felony in the US.

The title of TFA is a metaphorical criticism, not a literal law analysis.

They are not making the statement that the person in Australia who accidentally cancelled someone's reservation in Australia is literally guilty of violating US law. They are drawing criticism of AI models which are taking the kinds of actions for which, if a human did them knowingly, would be illegal.

john_strinlai 6 hours ago|||
>An AI cancelling other people's gym classes is a felony? Don't computer systems fail all the time at holding reservations for people?

the difference is intent.

if a concierge/booking system makes a mistake (or has an unintended bug or whatever), no crime.

but if i (or an agent working on behalf of me) use an API in an obviously unintended way to revoke other people's reservations, that would fall under the computer fraud and abuse act (in the usa).

peter_d_sherman 6 hours ago||
>"the difference is

intent."

>"but if i (or an agent working on behalf of me) use an API in an obviously

unintended

way to revoke other people's reservations..."

?

john_strinlai 6 hours ago|||
i am not quite sure what your question is, as you simply quoted me and then put a question mark... i think you are confused that i used "intent" in one context, and "unintended" in a different context, is that right?

the first sentence: the difference is the intent of the person who caused the cancellations

the second sentence: but if i (or an agent working on behalf of me) abuse an API to do things it was not meant or designed to do, such as cancelling someone else's reservation

rcxdude 57 minutes ago||
The point is, the law cares about your intent. If you ask an agent to abuse an API to do those things, then yeah, you are probably liable. If you ask an agent to do something reasonable (like make a booking), and then it accomplishes that by abusing the API, then you probably are not.
bronson 4 hours ago|||
Double negative. An attacker using the API in an "obviously unintended" manner shows intent.
redox99 5 hours ago||
Yeah if you're unlucky you get hit with like 20 years for wire fraud.
nubg 6 hours ago||
Thank you, this benchmark to me proves that closed weight model companies are dangerous for our democracy and put kids at risk. They must be outlawed and all models must be made open weights!
0xbadcafebee 6 hours ago||
Open models with advanced security features are a huge security benefit. Because any script kiddie can use them to hack into random things, people will now be forced to spend more time securing their technology. And they won't have to learn how, because they can use those same models to find the holes and patch them.
polynomial 6 hours ago||
Not to be confused with a similarly named project: https://github.com/MLOpsNYC/felonybench
imnotr0b0t 2 hours ago|
[flagged]