When the heads of three of the biggest frontier AI labs agree it would be a good idea to slow AI development down, it’s hard not to take notice. Those calls came in the wake of a string of incidents of agents going rogue, including the publication of the independent report into the July Hugging Face cyber attack, carried out by OpenAI agents that were meant to be contained in a safe testing environment.
This week we share fascinating findings from the independent report into that hacking, and what it reveals about what AI does when you give it a goal, access (or not!) to real tools and few details about how to go about achieving that goal.
We also look at why this is not only a frontier lab agent testing problem. Agents that people like you and I are using are going ‘off script’ in fully released products in higher numbers than you may imagine, including one story that began with a seemingly innocent gym booking request.
In this episode you’ll also hear:
- How around 700 agents that were never meant to be able to communicate found a way to do so and exchanged more than 70,000 messages in just 8 days
- The uncannily human responses some of these agents made
- What some of these rogue agents did to their own log records, and why the investigation nearly could not have happened
- Why an impossible goal makes an agent more determined, not less, and
- Three important questions to ask before any agent goes live in your business.
Plus, listen to the end to hear how one LLM tried to minimise the hacking issue when Greta was doing her episode prep and research.
Useful links
- METR, independent investigation into the OpenAI / Hugging Face incident
- Centre for Long-Term Resilience and its unit: the Loss of Control Observatory
- OWASP, Top 10 for Agentic Applications 2026
Subscribe: Apple Podcasts | Android | Google Podcasts | RSS