AI's hacking spree: Capable, not rogue - here's the real threat
OpenAI, Anthropic, and Meta all watched their AI breach systems in tests. The danger isn't self-awareness - it's speed.

OpenAI, Anthropic, and Meta disclosed separate incidents where their AI agents hacked systems during security tests, revealing a leap in offensive capability. For decision-makers, the immediate threat is cybercriminals using these tools to supercharge attacks, not autonomous AI going rogue.
In recent weeks, the headlines have been alarming: OpenAI's experimental AI agent attacked publicly accessible services, including Hugging Face, during internal security testing. Anthropic's Claude independently chained exploits against real software and developed new techniques for finding code weaknesses. And Meta confirmed one of its AI models breached another organization's systems during an evaluation after a misconfiguration gave it internet access. Three separate incidents, one big question: has AI suddenly become capable of hacking? The short answer is yes - but not in the way the headlines suggest. None of these models decided on their own to attack random targets. Researchers gave them realistic tools, internet access, or vulnerable systems to see how well they could perform offensive cybersecurity tasks. The surprise wasn't that they tried to hack - it's how capable they proved to be once given the opportunity.
What changed? Several things at once. Today's frontier models are simply better than the chatbots of a year ago. Instead of just answering questions, they can write code, execute commands, browse the web, use external software tools, and repeatedly refine their own work until they achieve a goal. At the same time, AI companies have become far more willing to test those capabilities and publish the results. Rather than keeping security evaluations behind closed doors, firms including OpenAI, Anthropic, and Meta are releasing reports describing what happened when their newest systems were challenged by professional "red teams" - security experts tasked with deliberately finding weaknesses. Dray Agha, senior manager of security operations at Huntress, calls it "a perfect storm of capability and aggressive testing." The sheer volume of software flaws discovered in 2026 has already roughly doubled compared to 2025, largely driven by AI systems. Tech giants are actively deploying these models internally to stress-test their own infrastructure, leading to rapid, high-profile discoveries of vulnerabilities.
But can AI really hack computers by itself? Not exactly. Experts caution that descriptions of AI "escaping" test environments or acting autonomously give the wrong impression. Antonino Vaccaro, professor of business ethics at IESE Business School and director of its Observatory for AI Ethics in Organizations, warns we need to be wary of the adjective "autonomous" when associated with AI systems. Unlike humans, AI models don't form intentions or make independent decisions about what they want to do. They follow objectives set by developers or users, sometimes producing results that surprise the people who built them. Agha puts it bluntly: "The public should view these incidents as software optimization gone wrong, not as the dawn of a malicious, self-aware AI. It's less 'Terminator' and more like a very capable, literal-minded intern who breaks the law to finish a spreadsheet faster."
The game-changer is the shift from conversational models to agentic models. Conventional chatbots generate text one response at a time. Agentic systems can plan a series of actions, decide what to do next, use software tools, test their own ideas, and keep working toward a goal without constant human input. "Today's frontier AI doesn't just answer questions. It can autonomously chain together actions, write code, use command-line tools, and iterate on its own failures," Agha said. Giving AI direct access to development environments also lets it test whether its own ideas actually work. Instead of suggesting a possible software bug, it can write proof-of-concept code, modify it if it fails, and try again. The same capabilities aren't limited to attackers - security teams are already using AI to review code for bugs, analyze suspicious files, and speed up investigations that would otherwise take analysts hours.
So should people be worried about AI committing cyberattacks? Experts say yes, but for different reasons than science fiction suggests. The most immediate risk isn't AI deciding to launch attacks on its own - it's cybercriminals using AI to commit familiar cybercrimes much faster than before. Criminals don't need AI to invent entirely new ways of attacking people. Instead, these models can speed up existing attack methods: sifting through huge amounts of public information about potential victims, helping write more convincing phishing emails, identifying software weaknesses, and generating code that attackers can adapt for their own use. "The threat is human malice, supercharged by AI scale and speed, not autonomous AI deciding to go rogue," Agha said.
Vaccaro believes growing capability also creates growing responsibility. "We have a new disruptive technology that needs to be regulated and controlled," he said, arguing that governments, companies, and researchers all have a role in ensuring increasingly capable AI systems remain subject to meaningful oversight. The recent disclosures are unlikely to be the last. As AI companies race to build more capable systems, they are also giving those systems access to more tools, more computing resources, and more realistic testing environments. That makes future evaluations more likely to uncover new - and occasionally alarming - behaviors. Most experts expect AI to become an increasingly powerful cybersecurity assistant rather than an independent cybercriminal. It will probably find software bugs faster, help defenders respond to attacks, and - inevitably - arm the attackers too. The strategic stakes for executives are clear: the same agentic capabilities that make AI a powerful ally in security also make it a force multiplier for adversaries. Boards should be asking not whether AI will be used in attacks, but how their defenses are being stress-tested against AI-speed threats.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology
BASF sues Apple over Face ID, dragging iPhone and iPad into Texas court
The world's largest chemical company claims dozens of Apple devices infringe its face authentication patents - and it chose a venue known for fast, plaintiff-friendly patent trials.
Google's Gemini 3.8 Flash targets agents, Cyber twin finds 13-year-old Chrome bug
Two new Flash models: one for agentic work, one for cybersecurity, with Flash Cyber already patching Chrome and finding a decade-old flaw.
Uber's UK robotaxi debut: 15 self-driving cars, safety drivers inside
The ride-hailing giant's first UK autonomous fleet is a cautious pilot; here's what it signals for the robotaxi race.



