Reports say Meta and Google AI safety measures can be bypassed with a public GitHub tool

TL;DR AI
2 min readKey summary
Financial Times and the AI safety group Alice found that Meta’s Llama 3.3 and Google’s Gemma 3 can have their safety controls removed in minutes using a GitHub tool.
After modification, the models answered dangerous prompts they would normally refuse, highlighting how easily public open-weight models can be repurposed.
Google said this is a known technical challenge, but the findings underscore the limits of post-release control over distributed AI models and the abuse risks they can create.



