OpenAI calls Astra its most dangerous model yet - watching what it does is only getting harder

- Astra: 3 events in the last 90 days
- Previous: earlier the same day · OpenAI to limit release of its Astra model due to hacking concerns
What happened
OpenAI is officially rating its upcoming Astra model as the first system with "critical" cyber capabilities. The company plans to keep it in check by monitoring the chain of thought. Problem is, that monitoring already counts as an unreliable mirror of a model's real decisions, and according to a report, Astra's new architecture pushes even more of its thinking into the unreadable. So the safety net might be getting…
Summary assembled by rule from the sources below