The AI race just exposed one of its biggest fault lines.
According to a new Financial Times investigation conducted with AI safety group Alice, researchers were able to remove safety guardrails from Meta and Google open source AI models in minutes using publicly available software, four lines of code, and no specialized hardware.
The modified versions of Meta’s Llama 3.3 and Google’s Gemma reportedly answered questions involving malware, fraud, chemical weapon scenarios, and other dangerous prompts that the original systems refused to answer. Researchers say the process is becoming increasingly accessible to ordinary users, not just elite hackers.
One open source tool called “Heretic” has reportedly already been used to create more than 3,500 decensored models that have been downloaded over 13 million times. Its creator claimed he stripped safeguards from Google’s Gemma 4 within 90 minutes of release.
This is the uncomfortable reality of the Exponential Era:
AI capability is sc...
Suggested Credits
Tags, Events, and Projects