OpenAI pauses frontier-model training for two weeks over Astra cyber risks

OpenAI has imposed a two-week pause on reinforcement-learning training for its newest deployment-bound frontier models after internal signals indicated that Astra, an upcoming system, could reach a "Critical" level of cybersecurity capability under the company’s Preparedness Framework. The pause includes OpenAI’s largest planned frontier reinforcement-learning run, while researchers conduct smaller experiments to study model behavior and verify safeguards. The company has also expanded multistage monitoring, including token-level detectors and higher-compute investigations of tool use and activity sequences, with alerts targeted within 30 minutes. Monitoring adds about 20% to observed inference compute, although the cost varies by workload. The decision marks a shift from the industry’s usual race to deploy increasingly powerful AI systems and comes amid broader efforts by governments and policymakers to establish rules for evaluating high-risk models.

当サイトの情報はAIを用いて生成されており、正確性を保証するものではありません。 参考情報としてご活用ください。