Posted by quantumgarbage 1 day ago
> We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext.
Should be easy for them to fix though: switch up the encryption key so it only works with the API for each specific model, rather than being shared across all of their models.
And indeed, the paper says it's been fixed by all three providers (though no news on how they fixed it):
> All model providers acknowledged the receipt of our report and subsequently we were unable to launch the same attacks.
You think that you are adding a description for your post when in fact you're simply submitting a regular comment not promoted or distinguished in any way.
It's even worse considering posts without URLs would take the same text from the same input box on the HN submission form and append it under the post title, locking it to the top of the page - so if you've browsed linkless posts you think maybe the submission text is gonna end up just like that, but it doesn't. (Also I think you'll even see URLs like an archive link end up appended just below the submission title for URL posts. Which I guess is a special feature.[1])
Any reason I should not send HN an email requesting clarification of this the submission page?
Edit: quoting https://news.ycombinator.com/submit :
“If there is no url, text will appear at the top of the thread.”
OK, now I finally understand what that means in practice, but it doesn’t imply an entirely separate undistinguished comment will be simultaneously submitted on my behalf.[1]modpowers(?) used to directly append links to URL submissions further confuse the matter: https://news.ycombinator.com/item?id=49243880
I feel bad now :(
Feedback emailed to HN!
In principle, yes, I want total control and visibility into the reasoning process.
In practice, I find that it takes up so much time to DIY reasoning agents that I can't spend much energy on the actual application.
The model providers have way more resources and talent to do this right and keep it right over time. I am willing to concede this moat to them if it means I can actually focus on the business.
The more I think about it, the less I care to see those tokens. It feels an awful lot like obsession over logging every trace item an application could produce. Useful in theory but a nasty garage stacked to the ceiling with useless shit otherwise. What nefarious things are we concerned with? That they burn too many tokens in the black box? I'll threaten to move to a different black box. There are always options in a market with this many participants and big players feel this pressure. They know there's a point at which being evil bastards is no longer profitable.
Spending time carefully designing tools and views that interact with the environment in clean ways is a much better investment than trying to own and control 100% of the reasoning process or models.