Incident Response
Ketika production down, bukan saatnya panik. Incident response adalah proses terstruktur untuk menangani outage dengan cepat dan efektif.
Severity Levels
- SEV1 (Critical) — Service down, semua user terdampak. All-hands response.
- SEV2 (Major) — Degraded performance, sebagian user terdampak.
- SEV3 (Minor) — Bug yang menyebabkan inconvenience, tidak blocking.
- SEV4 (Low) — Cosmetic issue, bisa ditangani saat working hours.
Incident Workflow
1. DETECT
- Alert dari monitoring (Grafana, PagerDuty)
- User report
2. TRIAGE
- Tentukan severity
- Assign incident commander
- Buat war room (Slack channel / video call)
3. MITIGATE
- Prioritas: pulihkan service, BUKAN cari root cause
- Rollback deployment?
- Scale up?
- Toggle feature flag off?
- Failover ke backup?
4. RESOLVE
- Confirm service recovered
- Monitor untuk recurrence
5. FOLLOW-UP
- Blameless postmortem (dalam 48 jam)
Blameless Postmortem Template
## Incident: [Judul]
**Date:** 2026-04-15
**Duration:** 45 menit (14:30 - 15:15 WIB)
**Severity:** SEV2
**Impact:** 30% user tidak bisa login
## Timeline
- 14:25 — Deploy v2.5.1 ke production
- 14:30 — Alert: login error rate > 10%
- 14:35 — Incident declared, war room dibuat
- 14:40 — Root cause: migration gagal, auth table locked
- 14:45 — Rollback ke v2.5.0
- 15:15 — Service fully recovered
## Root Cause
Migration ALTER TABLE users menambah kolom dengan NOT NULL tanpa default value pada tabel 2M rows, menyebabkan table lock selama 45 menit.
## Action Items
- [ ] Semua migration harus di-test di staging dengan production-size data
- [ ] Tambahkan migration timeout di CI pipeline
- [ ] Buat runbook untuk database rollback
Key Principles
- Blameless culture — Fokus pada sistem, bukan orang
- Mitigate first — Pulihkan dulu, investigasi nanti
- Communicate — Status page, Slack updates, customer notification
- Learn — Setiap incident adalah kesempatan memperbaiki sistem