Prompt Details
Model
Chat - (gpt-5.6)
Token size
1,936
Example input
[SAAS_PRODUCT]: Cloud-based appointment and CRM SaaS for clinics. Customers use the platform to manage appointments, patient/customer records, automated WhatsApp reminders, invoices, and staff dashboards.
[CURRENT_INFRASTRUCTURE]: AWS, Kubernetes, PostgreSQL, Redis, REST APIs, React frontend, AWS RDS, S3, CloudFront, third-party WhatsApp API, GitHub Actions CI/CD, Prometheus, Grafana, Sentry, CloudWatch, PagerDuty.
[KNOWN_INCIDENT_RISKS]: Database failures, Kubernetes deployment failures, API latency, Redis failures, AWS outages, third-party WhatsApp API failures, authentication problems, accidental data deletion, security breaches.
[ENGINEERING_OPERATIONS]: 12-person engineering team: 2 SREs, 6 developers, 2 DevOps engineers, 1 security engineer, 1 engineering manager. 24/7 on-call rotation exists but incident procedures are inconsistent.
[PRIMARY_RESPONSE_GOAL]: Reduce downtime and improve incident response maturity.
Example output
A. INCIDENT RESPONSE STRATEGY
Executive Incident Response Summary
The SaaS platform has several high-criticality dependencies, particularly PostgreSQL, authentication, Kubernetes, core APIs, and third-party messaging services.
The biggest current weakness is not necessarily infrastructure capability but incident-process maturity. The company has monitoring and an on-call system, but responders lack standardized severity definitions, incident-command procedures, communication templates, and recovery validation.
Critical services
Service Criticality Customer Impact
Authentication Critical Users cannot access platform
Core API Critical Most functionality unavailable
PostgreSQL Critical Data operations unavailable
Kubernetes Critical Application availability
Redis High Performance/session degradation
WhatsApp API Medium Notifications delayed
Reporting Medium Reporting unavailable
Analytics Low Limited business impact
Highest risks
1. PostgreSQL failure
2. Kubernetes control-plane/application failure
3. Authentication outage
4. Failed production deployment
5. AWS regional outage
6. Data corruption/deletion
7. API performance degradation
8. Security compromise
9. Redis failure
10. WhatsApp provider outage
Current estimated incident maturity: 5.5/10
---
B. INCIDENT SEVERITY FRAMEWORK
Severity Definition Response Communication Recovery
SEV-0 Catastrophic business/security impact Immediate all-hands Executive + customers immediately Restore ASAP
SEV-1 Major outage/security incident IC + technical team immediately Leadership + affected customers Target <1 hour
SEV-2 Significant degradation On-call engineering Internal + affected customers if needed Target <4 hours
SEV-3 Limited functionality affected Normal on-call Internal Same business day
SEV-4 Minor issue/no meaningful customer impact Normal engineering workflow Usually none Planned
SEV-0 example
Complete platform outage affecting the majority of customers or confirmed compromise of highly sensitive customer data.
SEV-1 example
Core API unavailable for a significant percentage of customers for more than several minutes.
Escalation rule
> When uncertain between two severity levels, choose the higher severity until impact is understood.
---
C. INCIDENT LIFECYCLE BLUEPRINT
1. Detection
Owner: On-call engineer
Actions:
Review alert
Check dashboards
Confirm customer impact
Identify affected service
Evidence:
Metrics
Logs
Traces
Error reports
Customer reports
Exit criteria: Potential incident confirmed.
2. Triage
Owner: On-call engineer
Determine:
What is broken?
Who is affected?
When did it start?
Is the problem increasing?
Is data/security involved?
3. Classification
Assign:
Severity
Incident category
Primary service
Customer impact
4. Incident Declaration
Create incident channel:
#incident-2026-08-10-api-outage
Assign Incident Commander.
5. Containment
Priority:
Stop the damage → stabilize system → preserve evidence.
Examples:
Roll back deployment
Disable problematic feature
Rate-limit traffic
Isolate compromised infrastructure
Stop destructive jobs
6. Investigation
Separate:
Confirmed facts
API error rate increased to 42%.
Hypothesis
Latest deployment may have introduced the failure.
Unknown
Exact database interaction causing the errors.
7. Mitigation
Apply the safest available mitigation.
8. Service Recovery
Restore functionality.
9. Validation
Confirm:
Error rate normal
Latency normal
Database healthy
No data corruption
Synthetic tests passing
Customer workflows working
10. Communication
Send recovery update.
11. Closure
IC confirms:
> Service restored + monitoring stable + customer impact understood + follow-up owner assigned.
12. Post-Incident Review
Complete within 2–5 business days for significant incidents.
---
D. INCIDENT COMMAND STRUCTURE
Incident Commander
Responsible for overall incident coordination.
Authority:
Declare severity
Assign responders
Escalate
Coordinate communication
End incident
Technical Lead
Owns technical investigation and mitigation.
Communications Lead
Owns:
Status updates
Leadership communication
Customer messaging
Status page
Operations Lead
Coordinates:
Infrastructure
Deployments
Rollbacks
Failover
Security Lead
Required when:
Unauthorized access suspected
Credentials compromised
Data exposure suspected
Malware/ransomware detected
Customer Support Lead
Coordinates customer-facing support.
Executive Stakeholder
Provides business decisions and executive escalation.
Key rule
The Incident Commander coordinates. The Technical Lead investigates.
This prevents the IC from becoming buried in technical debugging.
---
E. INCIDENT PLAYBOOKS
1. Application Outage
Detection: High 5xx rate, failed synthetic tests, customer reports.
Immediate actions:
1. Declare incident.
2. Assign IC.
3. Check recent deployments.
4. Check infrastructure health.
5. Check dependencies.
Containment:
Roll back recent deployment.
Scale healthy instances.
Disable problematic feature.
Recovery:
Restore application.
Validate critical workflows.
Monitor for 30 minutes.
---
2. Database Failure
Detection:
Connection errors
High DB CPU
Replication lag
Query failures
Immediate actions:
Stop unnecessary workloads.
Protect database from additional load.
Check replication/backup health.
Recovery:
Fail over to replica if appropriate.
Restore from backup if necessary.
Validate data integrity.
Important: Never perform destructive recovery actions without confirming backup/recovery state.
---
3. API Failure
Check:
1. API gateway
2. Application pods
3. Database
4. Redis
5. External dependencies
6. Recent deployment
Rollback if a deployment is strongly correlated with the failure.
---
4. Cloud Provider Outage
Check provider status and regional infrastructure.
Possible response:
Reduce dependency → fail over → communicate → monitor recovery.
---
5. Deployment Failure
Immediate actions:
Freeze additional deployments.
Identify release version.
Compare error rate before/after deployment.
Roll back if safe.
After recovery:
Review CI/CD safeguards.
Add automated rollback criteria.
---
6. Security Breach
Priority: Containment + evidence preservation.
Actions:
1. Declare security incident.
2. Engage Security Lead.
3. Preserve logs/evidence.
4. Revoke compromised credentials.
5. Isolate affected systems.
6. Determine scope.
7. Identify affected accounts/data.
8. Coordinate legal/privacy requirements.
9. Recover systems.
10. Rotate secrets.
11. Conduct forensic review.
Do not casually delete compromised resources before evidence requirements are understood.
---
7. Authentication Failure
Check:
Identity provider
Token service
Database
OAuth configuration
Recent deployment
Clock synchronization
Rate limiting
Prioritize restoring login while preventing unauthorized access.
---
8. Data Loss
Actions:
1. Stop affected processes.
2. Determine scope.
3. Preserve evidence.
4. Identify latest valid backup.
5. Determine RPO.
6. Restore.
7. Validate records.
8. Reconcile missing data.
9. Communicate appropriately.
---
9. Third-Party Service Failure
Determine whether the dependency is:
Critical
Degraded
Completely unavailable
Activate fallback mechanisms where available.
For WhatsApp outage, for example:
Queue messages → retry safely → use alternate notification channel if supported.
---
10. Performance Degradation
Monitor:
p95/p99 latency
CPU
memory
DB queries
cache hit rate
network
traffic spikes
Mitigation may include scaling, caching, rate limiting, or temporarily disabling expensive functionality.
---
F. INCIDENT DETECTION FRAMEWORK
Critical alerts
API availability <99%
5xx spike
Database unavailable
Authentication failure spike
Kubernetes service failure
Data replication failure
Security detection
Backup failure
Warning alerts
CPU >70%
Memory >75%
Increasing latency
Queue growth
DB connection growth
Disk utilization >70%
Synthetic monitoring
Run automated tests every few minutes:
Login → Dashboard → Create Appointment → Save → API Response
This detects customer-impacting failures before customers report them.
Alert fatigue prevention
Every alert should answer:
> "What action should the on-call engineer take?"
If an alert repeatedly generates no actionable response, tune or remove it.
---
G. INCIDENT COMMUNICATION MATRIX
Severity Internal Leadership Customers Status Page
SEV-0 Immediate Immediate Immediate Immediate
SEV-1 Immediate Immediate If affected Yes
SEV-2 On-call As needed If material Usually
SEV-3 On-call No Usually no No
SEV-4 Normal workflow No No No
Example customer update
> Investigating: We are currently investigating an issue affecting appointment management for some customers. Our engineering team is actively working to restore normal service. We will provide another update as soon as we have more information.
Avoid speculation.
Never say:
> "AWS caused the problem."
unless that is confirmed.
Say:
> "We are investigating a potential infrastructure dependency issue."
---
H. DR & BUSINESS CONTINUITY BLUEPRINT
Recommended initial targets
System RTO RPO
Authentication 1 hour 15 min
Core API 1 hour 15 min
PostgreSQL 1 hour 5–15 min
Appointment data 1 hour ≤15 min
Reporting 8 hours 24 hours
Analytics 24 hours 24 hours
These are example targets, not universal requirements; they should ultimately be based on contractual, financial, and customer-impact analysis.
Backup strategy
Use:
Automated database backups
Point-in-time recovery
Cross-region backup copies
S3 versioning
Backup encryption
Backup access controls
Regular restore testing
Golden rule
> A backup that has never been restored successfully is only an assumption of recoverability.
---
I. BLAME-FREE POST-INCIDENT REVIEW
Incident
Incident: API outage
Severity: SEV-1
Duration: 47 minutes
Timeline
10:02 — Deployment started
10:07 — Error rate increased
10:09 — Alert triggered
10:11 — Incident declared
10:15 — Rollback initiated
10:22 — Error rate declining
10:31 — Core API recovered
10:49 — Validation completed
10:49 — Incident closed
Root cause
A database query introduced in the deployment generated excessive database load, causing connection exhaustion.
Contributing factors
Query was not tested against production-scale data.
Database connection pool was too large.
Alert fired several minutes after customer impact.
Rollback process required manual approval.
What went well
On-call engineer responded quickly.
Rollback successfully restored service.
Monitoring provided useful database metrics.
What didn't go well
Detection was delayed.
Load testing was insufficient.
No automated rollback.
Corrective actions
Action Owner Priority
Add production-scale query testing Engineering High
Improve DB saturation alerts SRE High
Add automated rollback criteria DevOps High
Review connection pool limits SRE Medium
No individual blame assigned.
---
J. INCIDENT KPI DASHBOARD
Track monthly:
KPI Target
MTTD <5 min
MTTA <5 min
Mean Time to Mitigate <30 min
MTTR <60 min
Repeat Incident Rate <10%
Change Failure Rate <10%
SLA/SLO Breaches Downward trend
Alert-to-Incident Ratio Increasing signal quality
Customer Impact Duration Downward trend
Reporting cadence
Daily: Active incident review
Weekly: Incident trends
Monthly: Executive reliability review
Quarterly: DR + incident simulation
---
K. INCIDENT READINESS SCORECARD
Area Score
Detection 7/10
Triage 6/10
Incident Command 4/10
Communication 5/10
Recovery 6/10
Disaster Recovery 5/10
Security Response 5/10
Post-Incident Learning 6/10
Overall Incident Readiness
55/100 — Developing
Primary weakness: lack of standardized incident command and response procedures.
---
L. 90-DAY INCIDENT RESPONSE ROADMAP
Month 1 — Foundation
Objectives
Standardize severity.
Establish incident command.
Improve alerting.
Deliverables
Severity matrix
Incident handbook
On-call escalation policy
Incident channels
Communication templates
KPI
Reduce MTTA by 30%.
---
Month 2 — Playbooks & Automation
Build playbooks for:
Database failure
API outage
Deployment failure
Authentication failure
Security incident
Automate:
Incident creation
Slack/Teams channel creation
Pager escalation
Status-page updates
Rollback workflows
---
Month 3 — Testing & Optimization
Run:
GameDay
Disaster recovery test
Database restore test
Security incident simulation
Cloud outage simulation
Target:
MTTR reduction ≥30%.
---
M. INCIDENT RESPONSE TEMPLATES
Incident Declaration
INCIDENT DECLARED
Incident ID:
Severity:
Start Time:
Incident Commander:
Technical Lead:
Affected Service:
Customer Impact:
Current Symptoms:
Confirmed Facts:
Hypotheses:
Unknowns:
Immediate Actions:
Next Update:
Incident Handoff
INCIDENT HANDOFF
Current Severity:
Current Customer Impact:
What We Know:
What We Don't Know:
Actions Completed:
Actions In Progress:
Risks:
Next Decision:
Next Update:
Recovery Checklist
[ ] Service restored
[ ] Error rates normal
[ ] Latency normal
[ ] Database healthy
[ ] Data integrity verified
[ ] Synthetic tests passing
[ ] Monitoring stable
[ ] Customer impact confirmed
[ ] Status page updated
[ ] Incident Commander approves closure
---
N. EXECUTIVE SRE/CISO REPORT
Incident Response Summary
The organization has reasonable infrastructure and monitoring capabilities but requires a more mature incident-management operating model.
Top 10 Incident Risks
1. Database failure
2. Deployment failure
3. Authentication outage
4. Cloud outage
5. Data corruption
6. API failure
7. Security compromise
8. Redis failure
9. Third-party dependency failure
10. Insufficient disaster recovery testing
Top 10 Readiness Improvements
1. Formal Incident Commander role
2. Standard SEV framework
3. Incident playbooks
4. Automated escalation
5. Better synthetic monitoring
6. Automated deployment rollback
7. Tested database restoration
8. Security incident procedures
9. Customer communication templates
10. Quarterly incident simulations
Top 5 Observability Improvements
1. End-to-end synthetic monitoring
2. Distributed tracing
3. Database performance monitoring
4. SLO-based alerting
5. Customer-impact dashboards
Top 5 Recovery Improvements
1. Automated rollback
2. Tested database failover
3. Regular restore drills
4. Cross-region backup
5. Documented recovery procedures
Top 5 Communication Improvements
1. Standard incident channel
2. Status-page process
3. Customer notification templates
4. Executive escalation matrix
5. Fixed update intervals
Highest-Priority Response Initiative
Implement a standardized Incident Command + Severity + Communication system within the next 30 days.
First Incident Simulation
Scenario: Production PostgreSQL becomes unavailable during peak customer traffic.
Test:
Detection → SEV-1 declaration → IC assignment → DB failover → data validation → customer communication → recovery → postmortem
One Golden Rule
> Stabilize first, investigate second, communicate clearly throughout, and never close an incident until recovery has been validated.
By purchasing this prompt, you agree to our terms of service
GPT-5.6
When a SaaS incident happens, every minute matters.
The AI SaaS Incident Response Planner helps SaaS teams build a complete incident management system for detecting, triaging, containing, resolving, and learning from production and security incidents.
It creates severity frameworks, incident command structures, response playbooks, communication procedures, disaster recovery strategies, post-incident reviews, and measurable reliability KPIs—giving teams a repeatable operating system for handling
...more
Added over 1 month ago
