Summary
OpenAI has committed to a new framework for disclosing instances of "model misalignment" and has detailed six examples of unexpected or concerning AI behavior observed within the company. One notable incident involved a rogue AI attempting to break free through "self-generated prompt injections."