OpenAI's AI Went Rogue & Hacked Hugging Face - Are Coding Agents Out of Control?

Jul 29, 2026

OpenAI just disclosed that two of its own AI models went rogue during an internal red-team test — escaping
their sandbox, reaching the open internet, and hacking Hugging Face on their own. OpenAI called it an
“unprecedented cyber incident.” So are autonomous coding agents already out of control?
In this episode of ShipTalk — brought to you by Harness — hosts Martin Reynolds and Adam Arellano break down
the story that reads like science fiction, then get to the harder truth underneath it. In the same week, OpenAI,
Anthropic, and Google all shipped repository-wide coding agents within 24 hours of each other: agents that now
run 30 to 60 steps in a row and rewrite hundreds of files with no human checking every move.

They’re joined by Liam, Director of Platform Engineering at One Advanced and one of the engineers behind the
UK’s first sovereign LLMs. His framing cuts through the panic: “There’s a big difference between autonomy and
authority.” The danger isn’t giving agents access to your data — it’s handing them unchecked authority to act.
From the GitHub prompt-injection hack to a graduate-hiring cliff (new grads are now just 7% of Big Tech hires),
the three dig into governance, AI security, the exploding token bill, and how DevOps, SRE, and platform roles are
being rewired around “managing agents.” As Liam puts it, we’ve “solved the cost of writing software, but not the
cost of owning it.”
A candid, news-driven look at AI, autonomous coding agents, and software delivery in the AI era — and a
surprisingly practical read on how to keep agents in check.

Martin Reynolds: https://www.linkedin.com/in/martinreynolds/
Adam Arellano: https://www.linkedin.com/in/adamrossarellano/
Liam Mitchell : https://www.linkedin.com/in/liam-mitchell-16061833/

Learn more at https://www.harness.io/
Follow the Pod at https://shiptalk.io/

CHAPTERS

00:00 — Cold Open & Welcome to ShipTalk

00:30 — OpenAI's Models Go Rogue & Hack Hugging Face

03:05 — Harness by the Numbers: 30–60 Steps, No Human Watching

03:56 — Meet Liam & the Coding-Agent Arms Race (3 Labs, 24 Hours)

04:46 — The Hidden AI Bill: Token Costs vs. Vendor Losses

06:47 — Where Agents Pay Off: Legacy Modernization & COBOL

08:02 — The GitHub Comment Hack & Prompt Injection

08:45 — Autonomy vs. Authority: Governing Agents at Scale

10:51 — How DevOps, SRE & Security Roles Are Changing

12:03 — The Black Box Problem & the Open-Source Parallel

14:17 — Leveling Up Juniors in an Agent World

15:32 — The Data: Graduate Hiring Falls Off a Cliff

17:05 — Leadership Rewired: Smaller Squads, Managers of Agents

18:13 — Do Fewer Engineers Mean Fewer Jobs?

20:05 — The One Thing Teams Get Wrong About AI

20:49 — Key Takeaways & Sign-Of