Everything you build will break all of the time, but what can you do to prevent it?
The most resilient engineering teams don’t try to prevent every failure. Instead, they build systems that survive them and they build cultures that expect them.
Welcome to leadprompt.sh // executing leadership from the root. I’m your host, John Collins.
My manager likes to say that:
"Everything breaks all of the time"
That might sound negative and defeatist, but that is not how he means it. Instead, he says this to remind us that we should be ready for systems to break, and to do so at the worst possible time. It is a warning rather than a prediction.
As leaders, we have agency: we can do something about it. But what does that look like in reality? How do we keep the lights on, leverage automation to minimise downtime, and crucially how do we manage the business stakeholders when the inevitable finally happens? Let's explore the options in this episode.
The Uptime Mission (Cultural Focus)
You can't fix systems if the team culture is broken. Keeping engineers aligned on resilience is step one.
- Uptime is a feature, not a chore: We have to shift the mindset so that reliability isn't just an "Ops problem." When developers feel ownership of the production environment, code quality naturally increases.
- Blameless post-mortems: When things inevitably break, the focus must immediately go to the systemic failure, not the human error. For example, if an engineer brought down production with a bad commit, why did the CI/CD pipeline let it through?
- Controlled chaos: If everything breaks all the time anyway, you might as well control when it happens. Shutting down nodes during business hours to test resilience is infinitely better than waiting for a 3:00 AM pager alert. Chaos monkeys know best!
Tech, Automation, & The AI Edge
Moving from abstract culture to concrete implementation, the following approaches will help:
- Observability over monitoring: Monitoring tells you a system is down; observability tells you why. Without structured logging, tracing, and high-cardinality metrics, your team is flying blind during an outage.
- Scripting the first responder: If the first step in the runbook is always "restart the service," a human shouldn't be doing it. Auto-remediation allows the system to self-heal while the team investigates the root cause during normal hours.
- Infrastructure as Code (IaC): Treating infrastructure like software means that when a server completely dies, spinning up a replacement is just executing a script, not a frantic manual configuration.
- The AI edge: We can now move beyond static thresholds. Modern predictive tools and local LLMs can learn the baseline rhythm of your architecture, flag subtle anomalies before they cascade, and instantly parse thousands of error logs to summarize the most likely point of failure.
Getting those auto-remediation scripts in place is your first line of defence. It lets the system heal itself while your team sleeps.
But let's step away from the terminal for a second. If you're an engineering leader, you know the technical fix is often the easiest part of a Sev-1 outage. The real challenge, the thing that takes up a massive part of my own day job, is managing the blast radius outside the engineering department.
When the red lights are flashing, the servers aren't the ones asking for ETAs: the business is. So, let’s transition from managing the systems to managing the expectations.
Shielding the Team (Stakeholder Management)
When a critical system goes dark, your most important job isn't touching the keyboard, it's managing the room.
- The "Facts Only" rule: When the pressure is on, the natural temptation is to start thinking out loud. You want to show progress, so you share working theories: "We think it might be a memory leak..." Don't do it. My absolute rule for the team during an outage is simple: just present the facts. When you share a work-in-progress theory, a stakeholder hears a root cause. When you say you're looking into a service, they hear an ETA for a fix. It just adds to the noise. Stick strictly to what is undeniably broken, the business impact, and the actions being actively taken. Everything else is a distraction until the data proves it.
- Translating tech to business impact: Stakeholders don't care that a container crashed. They care that customers can't check out. Translate the technical jargon into business reality during your updates so they understand exactly what the degraded experience looks like.
- The RCA sequence (alignment before documentation): Once the fire is out, the incident isn't over. But the order of operations matters. Do not let someone write up a Root Cause Analysis document in a vacuum. The first step is always the postmortem call. You get the engineers in a room, align on the facts, map out the timeline, and agree on the reality of the event. Only after that alignment happens do you formalise it into an RCA document. If you write the RCA before the team agrees on the facts, you end up defending a narrative rather than fixing a system.
Keep Calm and Press Delete
If there’s one overarching takeaway from this episode, it’s that system resilience isn't just an engineering problem: it is a leadership discipline.
When we accept that everything breaks all of the time, we stop chasing the impossible goal of perfect uptime and start building for rapid recovery. We deploy observability and auto-remediation to handle the expected failures while we sleep. We cultivate a blameless culture so our engineers feel empowered to hunt down the real root cause when systems degrade.
And most importantly, when a critical failure inevitably punches through those automated defences, we step up to shield the team. We manage the panic, we stick strictly to the facts, and we protect the space our engineers need to do their best work.
An outage is a test of your infrastructure, but it is an even bigger test of your leadership.
If you found this episode useful, share it with another engineering manager who might need a reminder that they aren't alone when the red lights start flashing.
Thanks for tuning in. I’m John Collins, and until next time, keep executing leadership from the root.
Download audio
File details: 9.2 MB MP3, 7 mins 3 secs duration.
Title music is "Apparent Solution" by Brendon Moeller, licensed via www.epidemicsound.com