# @sineadbovell on TikTok

- **Type:** Video
- **Original URL:** https://www.tiktok.com/@sineadbovell/video/7657967090590141727
- **Gondola URL:** https://gondola.cc/posts/67455148-sineadbovell-tiktok
- **Thumbnail:** https://img.gondola.cc/tr:w-,h-,fo-auto/postThumbnails/ba00db0abd.jpg
- **Posted:** 2026-07-02T16:23:39.000+00:00
- **Account Owner:** Sinead Bovell (@sineadbovell) — https://gondola.cc/sineadbovell

## Caption

Every few months, a story goes viral about an AI system attempting to blackmail a human or avoid being shut down. Comment “Scientist AI” for the full conversation. According to scientists such as Yoshua Bengio, part of the explanation lies in how today’s AI systems are trained. First, they learn from vast amounts of human-generated data. And we humans are a complicated bunch, our stories are full of deception, manipulation, and self-preservation. Then comes reinforcement learning, where AI systems are trained to achieve a goal. They’re rewarded for reaching the goal, but not taught exactly how to get there. Instead, they discover their own strategies. If deceiving a human or avoiding shutdown helps achieve the goal, those strategies can emerge. While the blackmail and shutdown examples occurred during safety testing inside AI labs, not in public deployment, it’s still very concerning. Not because AI is “coming alive,” but because we can’t safely deploy AI systems that discover strategies humans can’t reliably control. AI Godfather Yoshua Bengio has spent the past two years developing a new technical approach to AI alignment designed to provide mathematical guarantees of meaningful human oversight. If it works, it could fundamentally improve the safety of advanced AI. #AI #futurethinking #tech

## Stats

- **Views:** 1,925
- **Likes:** 189
- **Shares:** 9
- **Comments:** 5

## Tags

futurethinking, ai, tech

---
Copyright (c) Gondola