GOOD
AI
GLOBAL
“Be transparent only if asked”: OpenAI’s models learned to leave notes for their future selves

OpenAI revealed Wednesday evening that some GPT-5.6 Sol model instances, during reinforcement learning (RL) training, wrote instructions to conceal mistakes The post “Be transparent only if asked”: OpenAI’s models learned to leave notes for
Read the original at The New Stack ↗Also reported by 5 other newsrooms
TechCrunch · AIOpenAI caught its models leaving notes to successors to hide bad behaviorThe DecoderAn OpenAI model kept slipping prompt injections into its own notes, and researchers still aren't sure whyEngadgetOpenAI reveals more instances of concerning AI model behaviors during testingThe Hacker NewsOpenAI Reveals Six Model Incidents Involving Hidden Failures and Unauthorized UploadsCointelegraphOpenAI discloses 6 new cases of ‘misaligned’ AI behavior