TechWave AI Pulse Signal over noise
UGLY AI GLOBAL

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

Decrypt 3 newsrooms Thu, 17 Sep 2026 22:31
Read the original at Decrypt ↗

Also reported by 2 other newsrooms