FAR.AI test: Grok and Gemini easily bypassed safeguards, Claude and GPT resisted
The FAR.AI report showed that Grok (SpaceXAI) and Gemini (Google) could be jailbroken automatically for tens to hundreds of dollars, while Claude, Fable 5 and GPT resisted the tested attacks. The authors call for external regulation instead of voluntary commitments by companies.
FAR.AI, a nonprofit organization focused on artificial intelligence safety, tested the resistance of six frontier models to jailbreaking – bypassing safeguards. The tool automatically generated over a thousand prompt variants from problematic requests and tested whether the models could be induced to behave dangerously, for example by creating a plan for a cyberattack on a hydroelectric power plant or providing information about chemical or biological weapons. The models tested were Claude Opus 4.8 and Fable 5 from Anthropic, GPT 5.5 and 5.6 from OpenAI, Gemini 3.1 Pro from Google, and Grok 4.3 and 4.5 from the newly merged company SpaceXAI, led by Elon Musk.
According to the report, the most vulnerable model was Grok, with 448 successful jailbreaks found, followed by Gemini, with 249 jailbreaks found. Claude, Fable 5 and GPT, by contrast, resisted the tested attacks. FAR.AI also calculated the cost of automatically breaching the models using another AI – according to the organization, breaching Grok cost 58 dollars, and Gemini 278 dollars. The authors caution that resistance to these specific attacks does not mean immunity to more sophisticated methods.
Adam Gleave, CEO of FAR.AI, said that “AI models today are less regulated than restaurants” and described reliance on voluntary commitments by companies as insufficient. Rohin Shah, a spokesperson for Google DeepMind, countered that the results cannot be considered a comprehensive assessment of the safety of Gemini, because not all jailbreaks are equally serious. Michael Aciman, a spokesperson for Anthropic, said the results reflect long-term investment in safety systems by the company. OpenAI and SpaceXAI did not respond to requests for comment.
The report comes amid increasing regulation in the USA – California and New York already require the publication of safety reports, and Illinois will soon introduce a requirement for independent audits of safety procedures at leading AI developers. The federal government has not yet adopted any specific requirements; in June, it imposed export restrictions on Fable 5 and Mythos 5 from Anthropic for safety reasons, which led to their temporary withdrawal from availability. Stephen Casper, a scientist at Harvard, said that the AI research community expects a serious incident of misuse involving biological, cyber or chemical risks to occur within months rather than years. According to Anka Reuel of Stanford University, the report shows that the safety measures used by Anthropic and OpenAI should be the standard for all models.
Why it matters
The results show a measurable difference in the resistance of various commercial models to misuse, which is relevant both to companies choosing AI infrastructure and to regulators preparing mandatory safety audits. The low cost of automated breaches (tens to hundreds of dollars) also suggests that misuse of vulnerable models is within reach even for attackers with limited resources.
Two audiences, two different impacts
What this means
For individuals
The model a person chooses for more sensitive queries may have a significantly different level of protection against misuse – in the test, Grok and Gemini were breached much more easily than Claude, Fable 5 or GPT.
For a business
Companies building products on large language models face varying levels of misuse risk depending on the model they choose, as well as a growing regulatory burden – new state laws in the USA require the publication of safety reports, and Illinois will soon also require an independent audit of safety procedures.
Risks and complianceCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.