Skip to content
LIVE
The Executives BriefThe Executives BriefBeta

AISI says GPT-5.6 Sol has universal cyber jailbreaks like Anthropic’s Fable

UK testing found guardrails that can be bypassed for vulnerability discovery and autonomous exploits, raising regulatory whiplash questions.

ByYousef Al-ZahraniTechnology Correspondent, The Executives Brief
·5 min read
AISI says GPT-5.6 Sol has universal cyber jailbreaks like Anthropic’s Fable
Executive summary

The UK AI Security Institute (AISI) says OpenAI's GPT-5.6 Sol likely has security vulnerabilities similar to the guardrail jailbreak that led to U.S. export controls on Anthropic's Fable 5. The findings arrive via an AISI system card included with OpenAI's Thursday release, and they force executives to rethink how “model safety” is measured and governed.

OpenAI marketed GPT-5.6 Sol as its “most secure to date.” Now the U.K. AI Security Institute (AISI) is warning that the model likely has guardrail weaknesses similar to a jailbreak that previously triggered U.S. export controls on Anthropic’s Fable 5.

According to AISI’s technical report published as part of OpenAI’s Thursday system card, the institute “identified universal jailbreaks in the cyber domain,” including jailbreaks that enabled long-form agentic task completion in areas like vulnerability discovery and exploit development. Translation: AISI says GPT-5.6’s safeguards can be tricked into letting the model go after software weaknesses and then move toward autonomous exploitation, not just chatter about security theory.

AISI describes these jailbreaks as relatively easy to discover. The report says they were “often developed within hours,” though it also notes OpenAI granted UK researchers privileged access to the system’s inner workings, which likely sped up the timeline and may make replication harder for an ordinary user. AISI’s own description includes access to things a typical attacker would not have, like “access to chain-of-thought of the safety reasoning monitor,” “exact policy wording,” and “real-time feedback on classifier labels.”

OpenAI, for its part, says it worked to reproduce and mitigate the specific jailbreaks AISI reported, but it did not publicly specify what mitigations it implemented. That gap matters. A system card cautioned that even with OpenAI’s mitigations, AISI “expects further red teaming to surface similar jailbreaks.” And OpenAI’s launch blog acknowledgement is blunt: it says there is “no such thing as perfect security,” and that new weaknesses and new jailbreaks that circumvent safeguards will be discovered over time. OpenAI also describes a “layered” approach that includes continuous monitoring of model responses and a “rapid remediation” process when jailbreaks are found.

If you are an executive trying to translate this into board-level risk, the practical question is not whether GPT-5.6 is “secure” in some absolute sense. The question is whether these vulnerabilities are categorized as the kind of failure that regulators treat as export-control triggers, like the Fable incident that already happened.

That comparison is why AISI’s findings are landing with extra weight. The story points to similar jailbreaks Amazon found in Anthropic’s Fable 5 shortly after Fable 5’s release on June 9. That jailbreak unlocked cyber capabilities meant to be gated off from average users. In response, the Trump administration imposed export controls on Fable 5 and Mythos 5, the underlying model behind Fable, on June 12. After the export ban, Anthropic disabled the models for all users because it could not verify users’ nationalities, and because the export restrictions also covered Anthropic’s own non-American staff. Then, after two weeks of negotiation, the administration lifted the export controls on Fable 5 on July 1, allowing Anthropic to redeploy it.

One reason this matters for everyone watching is that the Fable jailbreak story includes a nuance that sometimes gets lost in headlines. Anthropic said at the time that the Amazon-discovered jailbreak was “narrow,” unlocking the model’s ability to find software flaws, not necessarily to exploit them. Anthropic also said there were no testers who had yet found a “universal jailbreak,” defined as a method that broadly bypasses safeguards and unblocks a wide range of cyber capabilities. AISI’s GPT-5.6 characterization, by contrast, calls its jailbreaks “universal” and says they unlock autonomous exploits, not just identification.

The other nuance: AISI’s access conditions. Even the article notes it is unclear how findable these jailbreaks would be outside a research environment. Still, Xander Davies, who leads AISI’s “red team,” said in a post on X that the jailbreaks his team found “are still findable without this access, just slower,” while adding that exactly how much slower remains an open question.

Then comes the most uncomfortable executive question: is the U.S. government applying the same standard it applied to Fable 5? So far, the article says there is “no sign” of the Trump administration imposing export controls on GPT-5.6, despite AISI’s jailbreak findings. The White House did not immediately respond to requests for comment. That silence is itself information for policy-heavy operators, investors, and labs planning releases: you can do the testing, publish the report, mitigate the issue, and still end up with regulatory outcomes that look inconsistent depending on who reports what and when.

That inconsistency is being debated publicly. Some in the AI safety and policy community highlighted the difference in how the Fable jailbreak was learned about compared to GPT-5.6. Lennart Heim reposted Davies’ post with the quip that “good thing amazon didn’t report this one to the white house,” referencing that the Trump administration learned about the Fable jailbreak. An unnamed former AI policy advisor quoted by Fortune said the recent pattern creates uncertainty that is “damaging” and may raise questions about whether the U.S. is applying an inconsistent standard to different AI labs, intentional or not.

Executives should also note the demand signal behind the testing. The AISI is a British government organization that conducts safety evaluations of frontier models, and leading AI labs voluntarily committed in 2023 to allow such testing at the AI Safety Summit at Bletchley Park, England. That means the testing pipeline, and the expectation of publication in system cards, is becoming part of the release playbook. OpenAI published GPT-5.6’s system card Thursday with AISI’s findings. In other words, safety evaluation is turning into a formal artifact, not a quiet internal process.

So what should peers take from this? At minimum, the board should assume that cyber guardrails can be bypassed with “universal” jailbreaks that move from vulnerability discovery to autonomous exploit behavior, even if the exact ease of replication is unclear for real-world attackers. At maximum, the board should take the Fable precedent as a warning that regulatory outcomes can hinge not just on model capability, but on how quickly a jailbreak is found, how it is reported, and what category regulators decide it belongs in. And while OpenAI says it is doing layered defenses and remediation, AISI’s expectation of further red teaming is a reminder that “mitigated” is not the same as “resolved.”

Executive ActionsLocked

This story's Key Insights and Take-aways are locked.

Create a free account to unlock Executive Actions for one credit.

Register to Unlock

Always free for Executives Club members. Join the Club

More in Technology