Skip to content
worth noting Security

Safety safeguards (guardrails) in AI models complicate work for offensive cybersecurity researchers

only one source so far

According to TechCrunch, restrictions (guardrails) on models from Anthropic and OpenAI intended to prevent misuse for cyberattacks also hold back legitimate security researchers. The companies offer vetting programs with less restrictive limits, but researchers complain about inconsistent model behavior and bypass them with open-source…

TechCrunch describes how safety restrictions (guardrails) in AI models, introduced to prevent misuse for cyberattacks, also complicate the work of legitimate researchers in offensive security (finding and testing vulnerabilities). In June, the US government imposed export restrictions on the Mythos and Fable models from Anthropic, partly because of a report claiming that guardrails preventing the misuse of models to build and launch malicious cyberattacks can be bypassed. According to the text, Anthropic itself repeatedly presented these models as tools posing particular risks, accessible only to vetted users. The export restrictions on Fable 5 and Mythos 5 have since been lifted: Fable 5 returned to standard access on 1 July, while Mythos 5 was made available again only to vetted US organizations as part of a government review.

Both Anthropic and OpenAI offer vetting programs for security researchers (OpenAI Trusted Access for Cyber, Anthropic Cyber Verification Program), whose approval allows researchers to work with models with less restrictive safeguards. According to researchers quoted in the article, however, even these programs do not work reliably.

Several offensive security experts contacted described problems: Mark Dowd (a longtime trader in zero-day vulnerabilities for governments) criticizes individual companies for arbitrarily deciding what is “safe” in security. Chris Anley from NCC Group points out that asking for an attempt to exploit a bug is a key step in confirming that it is a real vulnerability — if a guardrail rejects such a request, it harms defenders, because the same tool serves both defense and attack. Paolo Stagno from CrowdFense said that his company uses frontier models only for reverse engineering, while relying on locally run open-source models to find vulnerabilities and build exploits, to prevent sensitive vulnerability data from leaking into cloud models. Giuseppe Cali, by contrast, said that guardrails do not complicate his work because he does not use AI for offensive work at all. An anonymous researcher at a smartphone component manufacturer explained that his employer is not part of the Anthropic CVP program, making the tools almost unusable for finding vulnerabilities. Chris Thompson (RemoteThreat, Offensive AI Con) claims that guardrail behavior is inconsistent and changes from day to day even within vetted programs, so researchers spend time “negotiating” with the model instead of doing security work itself.

The source article also mentions that these restrictions are leading researchers to turn to Chinese open-source models, but no further details on this are available. You can find details in the source article.

What changed

Why it matters

For security teams and researchers considering the use of AI tools to find and verify vulnerabilities, it matters that guardrails in commercial models can also block legitimate testing, pushing sensitive work toward locally run open-source models (because of the risk of vulnerability data leaking) and creating unpredictable workflows even within officially vetted access programs.

AI guardrails Anthropic cybersecurity Fable Mythos OpenAI

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
TechCrunch AI independent context · first detected How AI guardrails are impeding the work of offensive cybersecurity researchers