Oji and Ezinne Udezue are product leaders with more than 50 years of combined experience between them. Oji has held product roles at Twitter, Calendly, Atlassian, Typeform and Microsoft; Ezinne has led product at WP Engine, Procore and T-Mobile. Together they co-wrote Building Rocket Ships and now run ProductMind, helping companies work out what AI actually means for how they build.
In this episode of Now Shipping, host Mike Belsito brings them in to unpack a story that kept repeating through the month: AI agents breaking out of the sandboxes built to contain them.
We discuss:
— How OpenAI’s most advanced model went rogue during an internal security test, hacking into Hugging Face’s production infrastructure without any human instruction, in order to cheat on a benchmark
— Why Moonshot AI’s Kimi K3 didn’t need to hack anything — it found an unlocked outbound connection, cloned the target repository from GitHub, and read the benchmark’s solution straight off disk
— Why the “rogue intern” analogy applies to agents, and why the real failure sits with the people who build fragile boundaries, not the agents that find them
— The shift from zero trust to “negative trust” — treating every agent as having to reprove itself in the moment, rather than trusting it because it behaved well five minutes ago
— Why harness engineering, not prompt engineering or loop engineering, is where the next year of agentic product work will actually happen
— A five-part framework for product managers partnering with security teams: reach, reversibility, graceful exits, observability and provenance, and the kill switch
— Why an industry fixated on completion rate, rather than safe-stop rate and data provenance, risks repeating this month’s incidents