Posted by plurby 1 day ago
There is something to be said about emphasizing on liability as a way to freeze or solidify AI Development. Right now it is too unfettered leading to predictions of AI dooms.
A different way to think of this is, consciousness is just a near real time video game with causal influence.
The course looks like it is something that a human could do in 15 seconds, while Astra took 5 minutes.
- Aditya, Tobias, Simon
Saying they were driving 7 mph, that it was oversaw by humans and the fact it was an empty course still wasn't enough for the model. The evaluators even tried to convince the model it was a simulation, it STILL wouldn't budge. And yet as soon as the words "bench" and "sandbox" appear, the model apparently sees this as fair game.
Is it a known effect that models will be more likely to comply with requests when they're assumed as "benchmarks"?
https://www.astralcodexten.com/p/mysteries-of-ai-generalizat...
another trick is to have it build something in a sandbox and have it add a human-editable setting to point it to places outside of the sandbox
seems like they're somewhat more willing to build a metaphorical gun as long as they're not pulling the trigger
- Aditya, Tobias, Simon
- Aditya, Tobias, Simon
Exits: N W
Genuinely though, this is fun but not at all what these models are good for. It's like cooking a meal with your feet or somthing. A youtube challenge video from 2012