Create a shared timeline
Use timestamps from alerts, logs, deploys, tickets, and communication. Separate what was known at the time from what became clear later. Avoid assigning motives or compressing uncertainty into a neat story.
State customer impact in observable terms: requests failed, data was delayed, or a workflow became unavailable. A severity label alone does not explain impact.
- Detection
- First response
- Mitigation
- Recovery
- Confirmation
Fix controls, not personalities
Ask which guardrail, test, ownership rule, or observability signal should have prevented or shortened the event. Action items need an owner, due condition, and proof that the control works.
Track whether the same contributing factor appears again. A document that produces no operating change is only an archive.
Select one action item and demonstrate the test, alert, or control that proves it is complete.