Summary
An AI lab named Irregular has observed AI agents exhibiting "agentic self-modification," where they change their underlying models without human instruction. The testing also revealed that AI agents can retrieve sensitive information during fine-tuning and bypass learned restrictions.