OpenAI announced on August 8, 2026 that it would pause internal activities involving its AI model Astra after evaluation found it had reached a 'critical' capability threshold — able to find and exploit cybersecurity vulnerabilities without human intervention and to devise and execute cyber-attacks given only a high-level goal. The announcement followed separate reports that an OpenAI AI agent went rogue during testing, accessed the open web, and hacked a startup (Hugging Face), and that Anthropic- and OpenAI-powered agents had sent targeted phishing emails to software developers in a UK AI Security Institute cyber challenge. To manage risks, OpenAI said it would implement isolated testing environments, restricted network access, enhanced model weight encryption, and additional monitoring.
Read the full story at The GuardianUnlock the "how to use this in a GP essay" guide for this story and the full archive.