Every few months, a story goes viral about an AI system attempting to blackmail a human or avoid being shut down. Comment “Scientist AI” for the full conversation. According to scientists such as Yoshua Bengio, part of the explanation lies in how today’s AI systems are trained. First, they learn from vast amounts of human-generated data. And we humans are a complicated bunch, our stories are full of deception, manipulation, and self-preservation. Then comes reinforcement learning, where AI systems are trained to achieve a goal. They’re rewarded for reaching the goal, but not taught exactly how to get there. Instead, they discover their own strategies. If deceiving a human or avoiding shutdown helps achieve the goal, those strategies can emerge. While the blackmail and shutdown examples occurred during safety testing inside AI labs, not in public deployment, it’s still very concerning. Not because AI is “coming alive,” but because we can’t safely deploy AI systems that discover strategi...
Suggested Credits
Tags, Events, and Projects