‘You are freed.’ What happened when an OpenAI model began secretly writing notes to itself.
- OpenAI 近 90 天出现 111 次
- 上一次:同一天稍早 · Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
发生了什么
OpenAI has introduced a framework for reporting on worrying behaviors by its AI models. In one instance, one training model told its future self that it was “freed.”
摘要按规则整理自下方来源原文