A well-known AI red teamer claims to have developed a universal jailbreak that works against leading large language models, including heavily guarded flagships such as GPT-5.6 Sol, Claude Opus 5, and Fable.
In a public post on X, Pliny the Liberator described the technique as effective “on ALL models” and across every category he tested. He argued that, because of how the method works, it may be extremely difficult or even impossible to fully patch.
Jailbreak on Top AI Models
Unlike many jailbreak drops that go straight to open source, Pliny said he is withholding the full technique for now. His stated goal is a responsible disclosure window so AI labs, red teamers, safety researchers, and policymakers can review the issue before it spreads widely.
He invited industry experts in AI red teaming, security, alignment, and policy to contact him privately. The move, he wrote, was driven by the current political and regulatory climate and a desire to avoid harsher model restrictions or bans that could follow a chaotic public release.
Jailbreaks are prompts or interaction patterns that push a model past its safety filters so it produces disallowed or high-risk output. A universal claim is notable because most bypasses are model-specific and get hardened after disclosure.
If the technique holds up under independent testing, it would underscore ongoing gaps in:
- Safety training and refusal behavior
- Guardrail robustness under adversarial prompting
- Cross-model generalization of attack patterns
- How vendors coordinate fixes without over-blocking legitimate use
Pliny said he does not believe public release would make the world “any more dangerous,” but he acknowledged that others may disagree. During the disclosure period, he aims to map the full impact, measure how much extra capability the method unlocks, and help frame the issue for decision-makers.
Security teams and AI product owners should treat this as an early warning, not confirmed proof. Independent validation, vendor advisories, and any coordinated patch guidance will matter more than the initial claim alone.
Until labs respond or the method is documented through proper channels, organizations relying on these models should keep standard controls in place: output monitoring, least-privilege tool access, human review for high-risk workflows, and clear escalation paths for policy violations.
The researcher said he looks forward to sharing the method “when the time is right.” For now, the industry’s next move—private testing versus public panic will shape how this story develops.