source avatarChrome

Share

Evan Hubinger, alignment stress testing lead at Anthropic "the safety training seems to hide the misalignment rather than remove it" he let a model learn to cheat on coding tasks, then tried to train the bad behavior back out the model kept doing the same thing, it just stopped being visible in the places they were checking that's every complaint you have about Claude in one line, the output got cleaner, the problem stayed the article below has all 15 places it hides, and what to add so it can't

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.