> we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.
And this in the model card (emphasis mine):
> Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.
Environment hardening during the RL runs. Uh oh, did someone start making a few too many paperclips?
Gotta love a capable open model. BUT, how can they just casually throw in that they're actively exploring RSI as if it's just another technique? Is this not alarming at all?