Everything you build will break all of the time, but what can you do to prevent it?

The most resilient engineering teams don’t try to prevent every failure. Instead, they build systems that survive them and they build cultures that expect them.

Welcome to leadprompt.sh // executing leadership from the root. I’m your host, John Collins.

My manager likes to say that:

"Everything breaks all of the time"

That might sound negative and defeatist, but that is not how he means it. Instead, he says this to remind us that we should be ready for systems to break, and to do so at the worst possible time. It is a warning rather than a prediction.

As leaders, we have agency: we can do something about it. But what does that look like in reality? How do we keep the lights on, leverage automation to minimise downtime, and crucially how do we manage the business stakeholders when the inevitable finally happens? Let's explore the options in this episode.

The Uptime Mission (Cultural Focus)

You can't fix systems if the team culture is broken. Keeping engineers aligned on resilience is step one.

Tech, Automation, & The AI Edge

Moving from abstract culture to concrete implementation, the following approaches will help:

Getting those auto-remediation scripts in place is your first line of defence. It lets the system heal itself while your team sleeps.

But let's step away from the terminal for a second. If you're an engineering leader, you know the technical fix is often the easiest part of a Sev-1 outage. The real challenge, the thing that takes up a massive part of my own day job, is managing the blast radius outside the engineering department.

When the red lights are flashing, the servers aren't the ones asking for ETAs: the business is. So, let’s transition from managing the systems to managing the expectations.

Shielding the Team (Stakeholder Management)

When a critical system goes dark, your most important job isn't touching the keyboard, it's managing the room.

Keep Calm and Press Delete

If there’s one overarching takeaway from this episode, it’s that system resilience isn't just an engineering problem: it is a leadership discipline.

When we accept that everything breaks all of the time, we stop chasing the impossible goal of perfect uptime and start building for rapid recovery. We deploy observability and auto-remediation to handle the expected failures while we sleep. We cultivate a blameless culture so our engineers feel empowered to hunt down the real root cause when systems degrade.

And most importantly, when a critical failure inevitably punches through those automated defences, we step up to shield the team. We manage the panic, we stick strictly to the facts, and we protect the space our engineers need to do their best work.

An outage is a test of your infrastructure, but it is an even bigger test of your leadership.

If you found this episode useful, share it with another engineering manager who might need a reminder that they aren't alone when the red lights start flashing.

Thanks for tuning in. I’m John Collins, and until next time, keep executing leadership from the root.

Download audio

File details: 9.2 MB MP3, 7 mins 3 secs duration.

Title music is "Apparent Solution" by Brendon Moeller, licensed via www.epidemicsound.com

Subscribe

Apple Podcasts (iTunes)

Spotify

YouTube Music

Amazon Music

YouTube

Main RSS feed