OpenAI pauses frontier-model training for two weeks over Astra cyber risks

OpenAI has imposed a two-week pause on reinforcement-learning training for its newest deployment-bound frontier models after internal signals indicated that Astra, an upcoming system, could reach a "Critical" level of cybersecurity capability under the company’s Preparedness Framework. The pause includes OpenAI’s largest planned frontier reinforcement-learning run, while researchers conduct smaller experiments to study model behavior and verify safeguards. The company has also expanded multistage monitoring, including token-level detectors and higher-compute investigations of tool use and activity sequences, with alerts targeted within 30 minutes. Monitoring adds about 20% to observed inference compute, although the cost varies by workload. The decision marks a shift from the industry’s usual race to deploy increasingly powerful AI systems and comes amid broader efforts by governments and policymakers to establish rules for evaluating high-risk models.

The information on this website is generated using AI and we cannot guarantee its accuracy. Please use it as reference information only.