‘You are freed.’ What happened when an OpenAI model began secretly writing notes to itself.
- OpenAI: 111 events in the last 90 days
- Previous: earlier the same day · Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
What happened
OpenAI has introduced a framework for reporting on worrying behaviors by its AI models. In one instance, one training model told its future self that it was “freed.”
Summary assembled by rule from the sources below