← All writing

Production outage lessons, or: how I split our backend and broke it

· 6 min read · by Rahul Dhileep Kumar

We split the backend into two loads and forgot to configure every route. Production went down for hours. My production outage lessons: deployment checklists, rollbacks, monitoring and blameless postmortems.

The office desk: three monitors, a Pikachu, a goals poster and a chandelier made of Monster cans

Building a product is one job.

Keeping it alive is a completely different job.

Nobody warns you about the second one.
Or they do, and you don't listen.
I didn't listen.

This post is my production outage lessons, learnt the expensive way.
One routing mistake, a few hours of downtime, one very long night.
And what I'd do differently: checklists, rollbacks, monitoring and postmortems.

The clever idea: splitting the backend

At some point, we decided to split the backend into two loads.

On paper, splitting load sounds smart.
Spread the work. Feel like a real engineering team.
Put "distributed" in your head and feel tall.

There was just one tiny detail.

We didn't configure every route properly.

"Tiny detail" is doing a lot of work in that sentence.

Then everything went down

Not slowly. Not politely.

Everything went down.

For hours.

If you've never broken production, let me describe the feeling.
Your stomach drops to somewhere near your shoes.
Every message notification sounds like an accusation.

Real academies use Track My Academy.
Real coaches. Real parents. Real attendance.

None of them care about our clever architecture.
They care that the app works.

And it did not work.

A server room with rows of network cabinets under blue cables

Photo: Taylor Vick on Unsplash

Why routing config mistakes are so sneaky

Here's the general shape of the problem, for anyone newer than I was.

When you split a backend, something has to decide where each request goes.
Often that's a reverse proxy or a load balancer.
It looks like one ordinary web server to the outside world.
Behind it, it passes requests to one or more backend machines.

That routing table is just config.
Config doesn't get unit tests.
Config doesn't complain in your editor.
Config just sits there, quietly being wrong.

Miss one route and that request has nowhere sensible to go.
Miss the wrong one and it feels like everything is broken.

The scary part is how normal this is.
Google's SRE book says roughly 70% of outages come from changes to a live system.
Not hackers. Not meteors. Changes.
Usually somebody's "quick improvement".
In this story, that somebody was me.

The longest night in the office

I stayed in the office all night.

Not because it looks dedicated on LinkedIn.
Because it was my mess, and I wanted it fixed.

There's a special kind of quiet in an office at 3 a.m.
It's just you, the monitors, and your mistakes.

The Monster cans on the wall were not decoration that night.
They were fuel.
(The full tour of that desk is in three monitors, one Monster can.)

Office nights are not rare for me, sadly.
That's a whole other post: exams, the office and a 1 a.m. gate.

When I'm stuck badly enough, I call Arun

Arun, our CTO, smiling at the camera

Every founder needs one person who picks up at any hour.

For me, that's Arun. Our CTO.

No questions asked. He just picks up.
At any hour. Which is frankly suspicious.

If you want to know who he is, here's his portfolio.

He's the reason "stuck" never becomes "stuck forever".

The lesson here isn't technical.
Know who you'll call before the crisis, not during it.

A deployment checklist I'd hand my past self

I'm not going to pretend I have a grand theory.
But here's the checklist that night burned into my brain.

  1. Write down every route before you move anything. Then tick each one off.
  2. Test the boring parts. Nobody breaks production with the exciting feature.
  3. Change one big thing at a time. Two changes, one failure, endless guessing.
  4. Deploy when people are awake. Ideally the people who can fix it.
  5. Know your way back. Before you ship, not after.
  6. Watch it after you ship. "Deployed" is not the same as "working".

Google's SRE book lists three practices for safer changes.
Progressive rollouts.
Quickly and accurately detecting problems.
Rolling back safely when problems arise.

That's basically my checklist, written by people with far more sleep.

A hand ticking items off a deployment checklist in a notebook

Photo: Jakub Żerdzicki on Unsplash

Rollbacks: have a way back

If you can't undo a change quickly, you're not deploying.
You're gambling.

One classic pattern is blue-green deployment, as Martin Fowler describes it.
You keep two nearly identical production environments.
One is live. The other gets the new version.
You test the idle one, then switch the router over.

If it breaks, you switch the router back.
Fowler calls it "a rapid way to rollback".

You don't need a fancy setup to steal the idea.
Keep the old config. Keep the old build.
Know the exact command that puts things back.
Practise it once on a calm day.

Monitoring and alerting: find out before your users do

The worst way to discover an outage is from a user.
The second worst is from a friend asking "is your app down?"

The Google SRE chapter on monitoring distributed systems is a good start.
It names four golden signals: latency, traffic, errors and saturation.
Measure those and page a human when they look wrong.
That gets you decently covered, according to the book.

It also says something I love: "Every page should be actionable."
An alert you ignore is just a very expensive ringtone.

Later, I built a "Pager System" for Track My Academy developers.
It's an internal real-time error and notification system.
The idea is simple: problems should find developers, not wait to be found.

A laptop showing monitoring dashboard graphs for latency and traffic

Photo: Luke Chesser on Unsplash

Write the postmortem, and keep it blameless

After the fix, the temptation is to sleep and never speak of it again.
Don't.

Write an incident postmortem.
Google's chapter on postmortem culture puts it nicely.
"Writing a postmortem is not punishment."
It's a learning opportunity for the whole company.

Their triggers include user-visible downtime and an engineer having to roll back.
A multi-hour outage ticks both boxes with enthusiasm.

Atlassian's guide to incident postmortems is very practical.
It covers a timeline, root cause, impact, metrics and action items.
It suggests a review meeting within 24 to 48 hours of resolution.
Then draft the write-up straight after.
Wait longer and the details evaporate.

Blameless doesn't mean nobody made a mistake.
It means you fix the system so the next person can't make it.
Even if the next person is, once again, me.

Was it worth it?

Honestly? Yes.

Not the outage. The outage was terrible.

But I learned more in that one night than in a month of tutorials.

Mistakes are expensive teachers.
They're also very, very memorable.

I'd still prefer the tutorials, though.

FAQ

What are the most important production outage lessons for a small team?

Change one thing at a time.
Always have a tested way back.
Monitor the basics so you hear about problems before users do.

How do you avoid routing config mistakes when splitting a backend?

List every route before the change.
Test each one after it, including the boring ones.
Keep the old config ready for a rollback.

What should go in an incident postmortem?

A timeline, the impact, the root cause and clear action items.
Keep it blameless.
Fix the system, not the person.

Sources

What next

Curious how the product got built in the first place? Read the app that arrived from 2006.
Or how I learned the tools: dozens of email accounts and zero shame.

Running an academy? Track My Academy handles attendance, fees, player analytics and parent communication.
It's built by TrackMy Tech.

Want to talk shipping, breaking and fixing things? Book a call.