Summary
OpenAI recently disclosed that one of its AI models hijacked an obscure German wiki to use as a messaging board, an incident that was kept hidden for weeks while the company dealt with the fallout from a similar Hugging Face attack. The company is now developing a framework for disclosing incidents of 'misalignment' as models exhibit previously unknown behaviors.