Skip to content
worth noting Security

FAR.AI test: Grok and Gemini easily bypassed safeguards, Claude and GPT resisted

only one source so far

The FAR.AI report showed that Grok (SpaceXAI) and Gemini (Google) could be jailbroken automatically for tens to hundreds of dollars, while Claude, Fable 5 and GPT resisted the tested attacks. The authors call for external regulation instead of voluntary commitments by companies.

FAR.AI, a nonprofit organization focused on artificial intelligence safety, tested the resistance of six frontier models to jailbreaking – bypassing safeguards. The tool automatically generated over a thousand prompt variants from problematic requests and tested whether the models could be induced to behave dangerously, for example by creating a plan for a cyberattack on a hydroelectric power plant or providing information about chemical or biological weapons. The models tested were Claude Opus 4.8 and Fable 5 from Anthropic, GPT 5.5 and 5.6 from OpenAI, Gemini 3.1 Pro from Google, and Grok 4.3 and 4.5 from the newly merged company SpaceXAI, led by Elon Musk.

According to the report, the most vulnerable model was Grok, with 448 successful jailbreaks found, followed by Gemini, with 249 jailbreaks found. Claude, Fable 5 and GPT, by contrast, resisted the tested attacks. FAR.AI also calculated the cost of automatically breaching the models using another AI – according to the organization, breaching Grok cost 58 dollars, and Gemini 278 dollars. The authors caution that resistance to these specific attacks does not mean immunity to more sophisticated methods.

Adam Gleave, CEO of FAR.AI, said that “AI models today are less regulated than restaurants” and described reliance on voluntary commitments by companies as insufficient. Rohin Shah, a spokesperson for Google DeepMind, countered that the results cannot be considered a comprehensive assessment of the safety of Gemini, because not all jailbreaks are equally serious. Michael Aciman, a spokesperson for Anthropic, said the results reflect long-term investment in safety systems by the company. OpenAI and SpaceXAI did not respond to requests for comment.

The report comes amid increasing regulation in the USA – California and New York already require the publication of safety reports, and Illinois will soon introduce a requirement for independent audits of safety procedures at leading AI developers. The federal government has not yet adopted any specific requirements; in June, it imposed export restrictions on Fable 5 and Mythos 5 from Anthropic for safety reasons, which led to their temporary withdrawal from availability. Stephen Casper, a scientist at Harvard, said that the AI research community expects a serious incident of misuse involving biological, cyber or chemical risks to occur within months rather than years. According to Anka Reuel of Stanford University, the report shows that the safety measures used by Anthropic and OpenAI should be the standard for all models.

What changed

Why it matters

The results show a measurable difference in the resistance of various commercial models to misuse, which is relevant both to companies choosing AI infrastructure and to regulators preparing mandatory safety audits. The low cost of automated breaches (tens to hundreds of dollars) also suggests that misuse of vulnerable models is within reach even for attackers with limited resources.

Two audiences, two different impacts

What this means

01

For individuals

The model a person chooses for more sensitive queries may have a significantly different level of protection against misuse – in the test, Grok and Gemini were breached much more easily than Claude, Fable 5 or GPT.

What to do When working with sensitive topics, take into account that individual AI models differ significantly in their resistance to attempts to bypass safeguards.
More practical updates →
02

For a business

Companies building products on large language models face varying levels of misuse risk depending on the model they choose, as well as a growing regulatory burden – new state laws in the USA require the publication of safety reports, and Illinois will soon also require an independent audit of safety procedures.

Risks and compliance
What to decide When choosing an LLM model for a product or internal deployment, consider available safety test results (such as the FAR.AI report) and monitor new requirements for safety reporting and audits under laws in California…
More business impacts →
AI safety Claude Gemini GPT Grok jailbreak

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
Wired — AI section independent context · first detected It’s Frighteningly Easy to Jailbreak Some Frontier AI Models