Incident Response & Postmortems — DevOps

Incident Response Ketika production down, bukan saatnya panik. Incident response adalah proses terstruktur untuk menangani outage dengan cepat dan efektif. Seve

Incident Response

Ketika production down, bukan saatnya panik. Incident response adalah proses terstruktur untuk menangani outage dengan cepat dan efektif.

Severity Levels

Incident Workflow

1. DETECT
   - Alert dari monitoring (Grafana, PagerDuty)
   - User report

2. TRIAGE
   - Tentukan severity
   - Assign incident commander
   - Buat war room (Slack channel / video call)

3. MITIGATE
   - Prioritas: pulihkan service, BUKAN cari root cause
   - Rollback deployment?
   - Scale up?
   - Toggle feature flag off?
   - Failover ke backup?

4. RESOLVE
   - Confirm service recovered
   - Monitor untuk recurrence

5. FOLLOW-UP
   - Blameless postmortem (dalam 48 jam)

Blameless Postmortem Template

## Incident: [Judul]
**Date:** 2026-04-15
**Duration:** 45 menit (14:30 - 15:15 WIB)
**Severity:** SEV2
**Impact:** 30% user tidak bisa login

## Timeline
- 14:25 — Deploy v2.5.1 ke production
- 14:30 — Alert: login error rate > 10%
- 14:35 — Incident declared, war room dibuat
- 14:40 — Root cause: migration gagal, auth table locked
- 14:45 — Rollback ke v2.5.0
- 15:15 — Service fully recovered

## Root Cause
Migration ALTER TABLE users menambah kolom dengan NOT NULL tanpa default value pada tabel 2M rows, menyebabkan table lock selama 45 menit.

## Action Items
- [ ] Semua migration harus di-test di staging dengan production-size data
- [ ] Tambahkan migration timeout di CI pipeline
- [ ] Buat runbook untuk database rollback

Key Principles