This is one of those AI announcements that deserves more attention than it is probably going to get.
OpenAI says preliminary testing suggests its upcoming Astra model may have reached what the company classifies as Critical cybersecurity capability. (OpenAI)
The response was significant.
OpenAI temporarily slowed frontier model development, including a two week pause in reinforcement learning training. Its largest planned frontier RL run remains on hold while researchers conduct smaller training runs and safety evaluations. (OpenAI)
The company is also placing higher risk AI workloads inside stronger sandboxes, restricting internet access, strengthening network isolation, and expanding monitoring capable of examining model activity for unauthorized access, data theft, destructive behavior, and attempts to bypass safeguards. (OpenAI)
OpenAI estimates that this monitoring alone can consume roughly 20% of the inference compute being monitored. (OpenAI)
The bigger signal is d...
Suggested Credits
Tags, Events, and Projects