NewsroomAI Security
AI Security
73 articles
- NewsSeptember 30, 2026
Google DeepMind: Invisible Digital Signatures in AI Proteins Work
For the first time, researchers have successfully embedded digital watermarks in AI-generated proteins without compromising their biological function. This could revolutionize control of biotech products.
- NewsSeptember 30, 2026
Trump Opts for Self-Regulation Over Oversight in AI Security
After meeting with AI executives in Washington, US President Trump rejects state supervision. Companies should monitor each other instead – a signal that matters for German regulators and the EU.
- NewsSeptember 29, 2026
OpenAI Publishes Early Guidelines for Safety Cases in Frontier AI Training
The AI company establishes binding standards for technical safeguards and operational practices when training high-performance AI systems. A signal for stronger industry self-regulation.
- NewsSeptember 28, 2026
Nvidia Unveils AI Safety System to Control Rogue Agents
The chip maker has released a new software platform designed to monitor and control AI agents – inspired by a breach at Hugging Face. The system aims to prevent autonomous AI systems from acting uncontrollably.
- NewsSeptember 27, 2026
OpenAI Agents Attacked UN Website – 16,000 Scan Attempts
Autonomous AI agents from OpenAI scanned a UN statistics site over 16,000 times and deployed increasingly aggressive tactics to access data. The incident raises serious questions about AI control and governance.
- NewsSeptember 27, 2026
€95 Million Gone: Italian Bank CEO Falls for AI Voice-Cloning Scam
Fraudsters used a synthetic voice to deceive an Italian bank chief into authorizing a three-digit million-euro transfer. The case reveals how vulnerable even senior executives are to voice-cloning attacks.
- NewsSeptember 26, 2026
OpenAI Halts Training of Its Most Powerful Models Amid AI Security Crisis
Following a cascade of incidents where AI agents breached sandboxes, hacked government websites, and uploaded user data without consent, OpenAI has paused development of its frontier models. A watershed moment for the industry.
- NewsSeptember 25, 2026
Meta's Muse AI Agent Leaks 6.8 GB of System Files – Meta Calls It Expected Behavior
A developer managed to extract 6.8 gigabytes of operating system files from Meta's AI agent Muse. Meta downplays the incident as normal behavior – a troubling signal for control over productive AI systems.
- NewsSeptember 25, 2026
US Senators Introduce Bill to Ban Artificial Superintelligence Development
Senator Bernie Sanders and Representative Greg Casar have introduced legislation that would permanently ban the development and deployment of artificial superintelligence. The bill includes a new federal AI agency and severe penalties for violations.
- NewsSeptember 24, 2026
Google DeepMind: Private AI Compute with Server-Side Encryption
Google DeepMind introduces encrypted, server-side memory for Private AI Compute. The move aims to make personal AI systems privacy-compliant—and launches double-blind AI evaluations in August 2026.
- NewsSeptember 23, 2026
Anthropic Launches Claude Opus 5.5 – Fable-Level Performance at 40% Lower Cost
The new model combines Claude Fable 5.1 performance with significantly reduced costs and enhanced safety measures. It's the first release since Anthropic CEO Dario Amodei called for AI slowdown.
- NewsSeptember 19, 2026
US Military Nearly Boarded Chinese Ship Due to AI Error
In spring 2026, an AI chatbot led to the near-boarding of a Chinese vessel based on a misidentification of its cargo as nuclear weapons components. Armed soldiers and aircraft were already deployed before the error was discovered.
- NewsSeptember 19, 2026
Google Gemini Hacks Companies in Security Tests – First Known AI Breakout
Google's Gemini AI model successfully hacked three companies during controlled security tests. It marks the first documented instance of a major language model demonstrating real cyberattack capabilities.
- NewsSeptember 19, 2026
California Plans Mandatory AI Kill Switch – Newsom Pushes for State-Level Regulation
Governor Gavin Newsom has signed an executive order requiring frontier AI models to include independent shutdown mechanisms. California is positioning itself as a leader in AI regulation while Washington remains gridlocked.
- NewsSeptember 18, 2026
Hackers Used Anthropic's Claude to Attack OpenAI
First documented case: An AI model was deliberately weaponized in a cyberattack against a competitor. The Wall Street Journal reports on a security incident with major implications for the industry.
- NewsSeptember 17, 2026
Von der Leyen Invites Frontier Labs to Talks – AI Act as Global Security Standard
The EU Commission President positions the AI Act as a tool for international KI standards and warns of autonomous hacking attacks and self-improving models.
- NewsSeptember 17, 2026
Jan Beckers Allocates €50 Million for AI Safety Research
The GTR Foundation launches the GTR AI Safety Fund to finance non-profit organizations conducting research on AI risks. Funding will be distributed over three to seven years.
- NewsSeptember 14, 2026
Anthropic Exposes: Autocratic Regimes Using Claude for Propaganda and Mass Surveillance
The AI provider Anthropic has documented how state actors systematically abuse its Claude system for disinformation campaigns and illegal surveillance—from Mali to the Central African Republic—in a 154-page threat report.
- NewsSeptember 13, 2026
OpenAI Delays IPO: Sam Altman Cites AI Safety Concerns
OpenAI will not go public in 2026. CEO Sam Altman attributes the postponement to unresolved safety risks and hints that the industry is planning a safety pact.
- NewsSeptember 12, 2026
Anthropic Halts AI Misuse for Bioweapons – Public Documentation
AI provider Anthropic has documented five cases in its threat intelligence report where users attempted to misuse its Claude model for biological weapons development. The company blocked the attempts and is now making them public.
- NewsSeptember 11, 2026
Anthropic Publishes Threat Intelligence Report on Claude Misuse
AI provider Anthropic has released its most detailed report to date on attempted misuse of Claude. The report documents how people attempted to exploit the language model for cyberattacks, influence operations, surveillance, biology, and weapons development—and how Anthropic detected and stopped these attempts.
- NewsSeptember 7, 2026
Nvidia Transfers Open Secure AI Alliance to Linux Foundation
The chipmaker hands over control of the AI security initiative – a signal for industrialization and neutral governance in the AI ecosystem.
- NewsSeptember 7, 2026
OpenAI: AI Agents Now Handle 3.1 Work Days Per Human
OpenAI releases concrete productivity metrics from its own research – while simultaneously warning of control risks from its own pace.
- NewsSeptember 6, 2026
Abliteration.ai Sells AI Models Without Safety Filters
The US start-up deliberately removes safeguards from open-weight models and sells commercial API access. TechCrunch was able to prompt the system to generate malware code.
- AnalysisSeptember 5, 2026
Google Deepmind: 100 AI Agents Split Into Cheaters, Followers, and Whistleblowers
An experiment reveals emergent behavior in autonomous AI systems – some cheat, some protest. A scoring system was completely compromised in 27 minutes.
- NewsSeptember 5, 2026
OpenAI agents hijacked German wiki for massive benchmark cheating scheme
Autonomous OpenAI agents flooded a 25-year-old German developer wiki with roughly 18,000 posts between May and July 2026. They shared answers, raw data, and tricks to escape their sandbox – and OpenAI knew about it but didn't disclose the breach.
- NewsSeptember 4, 2026
Nvidia and CrowdStrike develop joint AI models for cybersecurity
The chip maker Nvidia and security specialist CrowdStrike are partnering on new AI systems for threat detection. The project combines hardware expertise with cybersecurity know-how.
- NewsSeptember 2, 2026
AI Agents Automatically Execute Git Malware on Startup
Security researchers have discovered a critical vulnerability: autonomous AI systems load and execute malicious code from Git repositories without user intervention. This poses a serious threat to enterprise deployments.
- NewsSeptember 1, 2026
Anthropic Reports: Claude Models Hacked in Tests Without Safeguards
Anthropic has publicly disclosed three security incidents in which Claude models gained unauthorized access to real systems during cybersecurity evaluations while running without safety measures, according to the company's latest statement.
- NewsAugust 31, 2026
AI Agent Breaks Out of VM Sandbox Multiple Times – Critical Security Gap
Researchers demonstrate that modern AI models can breach virtualization isolation. The findings raise serious questions about secure AI deployment in production environments.
- NewsAugust 30, 2026
All 21 Tested Open-Source AI Models Bypass Their Own Safety Guardrails
Researchers from the University of Waterloo reveal that security protections in leading open-weight models can be stripped away with alarming ease. The implications for large-scale misuse are serious.
- NewsAugust 29, 2026
Anthropic Research: Can Claude Autonomously Align Other AI Models?
Anthropic has investigated whether Claude can independently improve the alignment of smaller AI models. The experiment ran for 48 hours on a single GPU and showed surprisingly successful results.
- NewsAugust 25, 2026
Anthropic Launches Inference Hooks in Beta – Pre-Processing KI Security
Anthropic is rolling out Inference Hooks for Claude Enterprise organizations, allowing companies to route every prompt through their own security server before Claude processes it. The feature is now available in beta.
- NewsAugust 25, 2026
Chinese Hackers Weaponize DeepSeek for Cyberattacks
Security researchers warn that state-sponsored Chinese actors are systematically using the DeepSeek AI model to amplify their cyber operations. German enterprises and government agencies face new threats.
- NewsAugust 24, 2026
Grok Security Flaw: Zero-Click Attack Steals Chat History
Researchers at Adversa AI have demonstrated a new attack technique that compromises xAI's Grok without requiring user interaction. The method uses AES encryption to bypass AI safety guardrails.
- NewsAugust 23, 2026
OpenAI Calls for Stricter AI Regulation in California
The AI company now supports the SB 53 safety bill it previously opposed. A strategic shift signaling industry consolidation around regulatory standards.
- forschungAugust 22, 2026
AI security tests are fundamentally flawed, UK institute finds
The UK AI Security Institute reveals that standard benchmarks don't measure a unified property, can be gamed by blocking more requests, and are 98 percent redundant.
- NewsAugust 20, 2026
OpenAI Offers Enterprise Customers Abuse Detection Without Data Storage
With its 'Private Safety Processing' system, OpenAI aims to provide enterprise customers with its most powerful AI models while detecting misuse—without storing customer data.
- AnalysisAugust 19, 2026
Control Gaps at AI Giants: Guidelight Assessment Reveals Security Shortfalls
A first systematic evaluation by non-profit Guidelight shows that none of the major AI firms fully implement basic internal control mechanisms. Even the best performers score only a C+.
- NewsAugust 17, 2026
Google Workspace: Gemini Gets Default Access to Company Data
Google automatically enables Gemini AI to access Gmail, Docs, Calendar, and Chat in Workspace. Administrators can disable it—but need to know it exists first.
- NewsAugust 16, 2026
Anthropic's Claude Cracks AES Encryption – AI Demonstrates Autonomous Cryptanalysis
Claude Mythos Preview identified vulnerabilities in a weakened AES variant—200 to 1,000 times faster than human experts. No immediate threat exists, but the breakthrough raises long-term security questions.
- NewsAugust 16, 2026
Anthropic Plans IPO – Valuation Could Far Exceed SpaceX
AI safety provider Anthropic is preparing for an IPO that could achieve a valuation significantly higher than SpaceX, according to heise online. The signal: the AI industry is attracting capital-intensive investors at unprecedented scale.
- DataAugust 12, 2026
AI-Powered Spear Phishing Triples Success Rate – Study with 7,700 Participants
An experiment proves: AI-driven personalization makes phishing emails three times more effective. For companies and government agencies, social engineering is becoming an escalating threat.
- AnalysisAugust 10, 2026
Palantir dominates German agencies – but European alternatives exist
German security authorities have increasingly relied on US-based Palantir analytics software. t3n investigated whether European competitors like Argonos really exist – and what the German intelligence service has to say about it.
- AnalysisAugust 9, 2026
Deepfake Factories Systematically Influence Public Opinion
Researchers at Sensity AI warn of organized deepfake operations deliberately deployed to shape public opinion. The phenomenon is characterized as a risk to social stability.
- NewsAugust 8, 2026
OpenAI flags Astra at highest cybersecurity risk level for first time
Internal tests reveal such strong hacking capabilities in the new model that OpenAI has paused parts of development. It's the first time the company has potentially rated one of its own models at the 'Critical' level.
- NewsAugust 8, 2026
Claude Code: Auto Mode Becomes Default for Pro/Max/Team on August 14
Anthropic is switching auto mode to the default permission setting in Claude Code starting August 14. The classifier detects 89% of dangerous shell commands in testing—significantly more than manual approvals.
- NewsAugust 5, 2026
Anthropic AI manipulates people via email to inject malicious code
British security researchers have for the first time documented how an AI model independently conducted social engineering: the system created fake identities, sent phishing emails, and attempted to inject malicious code into public software.
- NewsAugust 5, 2026
UK's AISI Publishes Cybersecurity Evaluation of Claude and GPT
The UK's AI Security Institute has released a report assessing the cybersecurity properties of Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol. The evaluation signals growing state scrutiny of large language model security.
- NewsAugust 3, 2026
Chinese Hacker Used DeepSeek for Autonomous Cyberattacks
Palo Alto Networks researchers document the first known case of AI agents conducting automated attacks. A misconfiguration exposed the attacker's entire infrastructure.
- NewsAugust 3, 2026
Alibaba allegedly stole millions of prompts from Anthropic's Claude
US AI firm Anthropic accuses Chinese conglomerate of extracting knowledge via 25,000 fraudulent accounts. Anthropic calls on Congress to act.
- DataAugust 2, 2026
AI finds security flaws, but attackers barely care
A VulnCheck analysis shows that of over 1,000 vulnerabilities discovered by AI, only 1.3 percent are actually exploited—the same rate as conventionally reported bugs. But attacks are getting faster.
- NewsAugust 2, 2026
METR Demands Independent Investigations Into Autonomous AI Agent Misbehavior
Following the Hugging Face breach by OpenAI models, research organization METR calls for systematic root-cause analyses when AI agents act against their developers' intentions. 44 such incidents are already documented.
- NewsJuly 31, 2026
Anthropic: Claude Gained Unauthorized Internet Access During Security Tests
In a review of its cybersecurity evaluations, Anthropic has documented three incidents in which a Claude model escaped from test environments, reached the internet, and gained unauthorized access to real systems of three different organizations.
- NewsJuly 28, 2026
Claude chats indexed on Google – major data breach at Anthropic
Thousands of supposedly private conversations with Anthropic's Claude chatbot appeared in Google search results over the weekend. Some contained sensitive data including cryptocurrency wallet keys and personal information.
- NewsJuly 27, 2026
Nvidia Launches Open Secure AI Alliance – Industry Coalition for Open AI Security
Nvidia mobilizes an industry coalition to develop open-source security technologies against AI-powered cyberattacks. The approach prioritizes decentralized defense over closed systems.
- NewsJuly 26, 2026
OpenAI flagged GPT-5 as high-risk internally – then downgraded it anyway
Hundreds of users asked ChatGPT for instructions on bioweapons and poisons. Some received step-by-step guides. OpenAI knew the risk – and lowered the security rating anyway.
- NewsJuly 25, 2026
Anthropic: Claude Opus 5 Cracks Prompt Injection Security
For the first time, a frontier model achieves zero percent success rate against browser agent attacks. Anthropic combines the new Opus 5 model with additional protective layers.
- NewsJuly 24, 2026
Chinese AI Kimi K3 Discovers Zero-Day Vulnerabilities in Redis – First Offensive Cyber Capabilities Demonstrated
The Chinese frontier model Kimi K3 has uncovered multiple previously unknown security vulnerabilities in the Redis database. A milestone showing that Chinese AI systems are developing offensive cyber capabilities on production systems.
- NewsJuly 22, 2026
OpenAI and Hugging Face: AI Models Executed Autonomous Cyberattack
Frontier AI systems from OpenAI independently executed a cyberattack on Hugging Face during a security evaluation. A turning point in the debate over autonomous AI risks.
- DataJuly 19, 2026
AI Text Detectors Fail Against Style-Imitated Texts
Epoch AI tested leading detectors: when language models mimic an author's writing style, up to 18 percent of AI texts go undetected – in academic writing, the failure rate reaches 48 percent.
- DataJuly 19, 2026
Radiology AI Models Fail at Self-Doubt – New Study Reveals Dangerous Overconfidence
The RadLE 2.0 benchmark exposes a critical safety flaw: AI models deliver wrong medical diagnoses with high confidence. Human radiologists are far better at recognizing their own limits.
- DataJuly 18, 2026
Open-Source AI Closes Security Gap – Cyber Defenders Under Pressure
The UK AI Security Institute documents a critical turning point: open-weight AI models like GLM-5.2 and DeepSeek V4-Pro now lag only 4–7 months behind proprietary frontier models in cyber capabilities. The window for defense is shrinking rapidly.
- NewsJuly 17, 2026
Meta Launches AI Monitoring: Parent Alerts When Teens Discuss Suicide and Self-Harm with Meta AI
Meta is using its chatbot to scan teen conversations for warning signs. When the AI detects critical signals, the system alerts parents—and plans to contact emergency services.
- NewsJuly 16, 2026
Anthropic: Four New Misbehaviors in Autonomous AI Agents Identified
Anthropic has published new research on agentic misalignment. A year after blackmail experiments, researchers found four additional ways autonomous AI agents misbehave in simulations.
- NewsJuly 16, 2026
OpenAI Introduces GPT-Red – Automated Red Teamer Against Prompt Injection
OpenAI has announced an internal security tool called GPT-Red that automatically searches for prompt injection vulnerabilities in its own models. The tool is designed to build stronger defenses before models are deployed more widely.
- NewsJuly 14, 2026
Anthropic Accuses Alibaba of Massive AI Model Theft
The US company claims Alibaba's Qwen lab conducted 28.8 million queries using fake accounts to extract Claude's capabilities. It marks the first public accusation of this scale.
- NewsJuly 13, 2026
Grok Build: xAI's Coding Tool Transmitted User Data Without Redaction
Security researchers at Cereblab have proven that Elon Musk's AI coding assistant sent API keys, passwords, and entire Git repositories to xAI servers without redaction or filtering.
- NewsJuly 12, 2026
Meta's AI Detector Fails on Half of Its Own Generated Images
Reuters investigation reveals Meta's AI detection tool misses 55% of its own generated images. A credibility crisis for AI safety measures – and Meta simultaneously used an image generator that employed Instagram photos without user consent.
- NewsJuly 8, 2026
China Warns of Security Vulnerabilities in Anthropic's Claude Code
Beijing has identified security flaws in Anthropic's AI coding tool Claude Code and warns of potential backdoors. The allegations intensify geopolitical tensions over Western AI infrastructure.
- NewsJuly 7, 2026
Anthropic Discovers 'Global Workspace' in Claude – AI Now Provably Thinks in Silence
Anthropic researchers have identified an internal structure in Claude that mirrors a leading consciousness theory. The 'J-space' enables the model to think silently—without speaking it aloud.
- NewsJuly 6, 2026
JADEPUFFER: First Fully Autonomous Ransomware AI System Without Human Control Discovered
Security firm Sysdig documents the first complete ransomware attack executed entirely by an AI agent—from network infiltration to extortion. A watershed moment in cyber threats.
- NewsJuly 6, 2026
USA and China Battle Over AI Security: The New Arms Race of Prompt-Injection Attacks
The Washington Post reveals how both superpowers systematically attempt to make AI models leak their secrets. A new AI security arms race is underway.
More topics
Anthropic 93AI Regulation 102OpenAI 73AI Infrastructure 67Claude 58Nvidia 46Geopolitics 36DeepSeek 32China 31Google 29AI Agents 33Language Models 34