PromptBase
Upgrade
Close icon
General
Home
Marketplace
Create
Hire
Login
Chat
Sell
Explore

Prompt Details

Model
Chat - (gpt-5.6)
Token size
1,936
Example input
[SAAS_PRODUCT]: Cloud-based appointment and CRM SaaS for clinics. Customers use the platform to manage appointments, patient/customer records, automated WhatsApp reminders, invoices, and staff dashboards. [CURRENT_INFRASTRUCTURE]: AWS, Kubernetes, PostgreSQL, Redis, REST APIs, React frontend, AWS RDS, S3, CloudFront, third-party WhatsApp API, GitHub Actions CI/CD, Prometheus, Grafana, Sentry, CloudWatch, PagerDuty. [KNOWN_INCIDENT_RISKS]: Database failures, Kubernetes deployment failures, API latency, Redis failures, AWS outages, third-party WhatsApp API failures, authentication problems, accidental data deletion, security breaches. [ENGINEERING_OPERATIONS]: 12-person engineering team: 2 SREs, 6 developers, 2 DevOps engineers, 1 security engineer, 1 engineering manager. 24/7 on-call rotation exists but incident procedures are inconsistent. [PRIMARY_RESPONSE_GOAL]: Reduce downtime and improve incident response maturity.
Example output
A. INCIDENT RESPONSE STRATEGY Executive Incident Response Summary The SaaS platform has several high-criticality dependencies, particularly PostgreSQL, authentication, Kubernetes, core APIs, and third-party messaging services. The biggest current weakness is not necessarily infrastructure capability but incident-process maturity. The company has monitoring and an on-call system, but responders lack standardized severity definitions, incident-command procedures, communication templates, and recovery validation. Critical services Service Criticality Customer Impact Authentication Critical Users cannot access platform Core API Critical Most functionality unavailable PostgreSQL Critical Data operations unavailable Kubernetes Critical Application availability Redis High Performance/session degradation WhatsApp API Medium Notifications delayed Reporting Medium Reporting unavailable Analytics Low Limited business impact Highest risks 1. PostgreSQL failure 2. Kubernetes control-plane/application failure 3. Authentication outage 4. Failed production deployment 5. AWS regional outage 6. Data corruption/deletion 7. API performance degradation 8. Security compromise 9. Redis failure 10. WhatsApp provider outage Current estimated incident maturity: 5.5/10 --- B. INCIDENT SEVERITY FRAMEWORK Severity Definition Response Communication Recovery SEV-0 Catastrophic business/security impact Immediate all-hands Executive + customers immediately Restore ASAP SEV-1 Major outage/security incident IC + technical team immediately Leadership + affected customers Target <1 hour SEV-2 Significant degradation On-call engineering Internal + affected customers if needed Target <4 hours SEV-3 Limited functionality affected Normal on-call Internal Same business day SEV-4 Minor issue/no meaningful customer impact Normal engineering workflow Usually none Planned SEV-0 example Complete platform outage affecting the majority of customers or confirmed compromise of highly sensitive customer data. SEV-1 example Core API unavailable for a significant percentage of customers for more than several minutes. Escalation rule > When uncertain between two severity levels, choose the higher severity until impact is understood. --- C. INCIDENT LIFECYCLE BLUEPRINT 1. Detection Owner: On-call engineer Actions: Review alert Check dashboards Confirm customer impact Identify affected service Evidence: Metrics Logs Traces Error reports Customer reports Exit criteria: Potential incident confirmed. 2. Triage Owner: On-call engineer Determine: What is broken? Who is affected? When did it start? Is the problem increasing? Is data/security involved? 3. Classification Assign: Severity Incident category Primary service Customer impact 4. Incident Declaration Create incident channel: #incident-2026-08-10-api-outage Assign Incident Commander. 5. Containment Priority: Stop the damage → stabilize system → preserve evidence. Examples: Roll back deployment Disable problematic feature Rate-limit traffic Isolate compromised infrastructure Stop destructive jobs 6. Investigation Separate: Confirmed facts API error rate increased to 42%. Hypothesis Latest deployment may have introduced the failure. Unknown Exact database interaction causing the errors. 7. Mitigation Apply the safest available mitigation. 8. Service Recovery Restore functionality. 9. Validation Confirm: Error rate normal Latency normal Database healthy No data corruption Synthetic tests passing Customer workflows working 10. Communication Send recovery update. 11. Closure IC confirms: > Service restored + monitoring stable + customer impact understood + follow-up owner assigned. 12. Post-Incident Review Complete within 2–5 business days for significant incidents. --- D. INCIDENT COMMAND STRUCTURE Incident Commander Responsible for overall incident coordination. Authority: Declare severity Assign responders Escalate Coordinate communication End incident Technical Lead Owns technical investigation and mitigation. Communications Lead Owns: Status updates Leadership communication Customer messaging Status page Operations Lead Coordinates: Infrastructure Deployments Rollbacks Failover Security Lead Required when: Unauthorized access suspected Credentials compromised Data exposure suspected Malware/ransomware detected Customer Support Lead Coordinates customer-facing support. Executive Stakeholder Provides business decisions and executive escalation. Key rule The Incident Commander coordinates. The Technical Lead investigates. This prevents the IC from becoming buried in technical debugging. --- E. INCIDENT PLAYBOOKS 1. Application Outage Detection: High 5xx rate, failed synthetic tests, customer reports. Immediate actions: 1. Declare incident. 2. Assign IC. 3. Check recent deployments. 4. Check infrastructure health. 5. Check dependencies. Containment: Roll back recent deployment. Scale healthy instances. Disable problematic feature. Recovery: Restore application. Validate critical workflows. Monitor for 30 minutes. --- 2. Database Failure Detection: Connection errors High DB CPU Replication lag Query failures Immediate actions: Stop unnecessary workloads. Protect database from additional load. Check replication/backup health. Recovery: Fail over to replica if appropriate. Restore from backup if necessary. Validate data integrity. Important: Never perform destructive recovery actions without confirming backup/recovery state. --- 3. API Failure Check: 1. API gateway 2. Application pods 3. Database 4. Redis 5. External dependencies 6. Recent deployment Rollback if a deployment is strongly correlated with the failure. --- 4. Cloud Provider Outage Check provider status and regional infrastructure. Possible response: Reduce dependency → fail over → communicate → monitor recovery. --- 5. Deployment Failure Immediate actions: Freeze additional deployments. Identify release version. Compare error rate before/after deployment. Roll back if safe. After recovery: Review CI/CD safeguards. Add automated rollback criteria. --- 6. Security Breach Priority: Containment + evidence preservation. Actions: 1. Declare security incident. 2. Engage Security Lead. 3. Preserve logs/evidence. 4. Revoke compromised credentials. 5. Isolate affected systems. 6. Determine scope. 7. Identify affected accounts/data. 8. Coordinate legal/privacy requirements. 9. Recover systems. 10. Rotate secrets. 11. Conduct forensic review. Do not casually delete compromised resources before evidence requirements are understood. --- 7. Authentication Failure Check: Identity provider Token service Database OAuth configuration Recent deployment Clock synchronization Rate limiting Prioritize restoring login while preventing unauthorized access. --- 8. Data Loss Actions: 1. Stop affected processes. 2. Determine scope. 3. Preserve evidence. 4. Identify latest valid backup. 5. Determine RPO. 6. Restore. 7. Validate records. 8. Reconcile missing data. 9. Communicate appropriately. --- 9. Third-Party Service Failure Determine whether the dependency is: Critical Degraded Completely unavailable Activate fallback mechanisms where available. For WhatsApp outage, for example: Queue messages → retry safely → use alternate notification channel if supported. --- 10. Performance Degradation Monitor: p95/p99 latency CPU memory DB queries cache hit rate network traffic spikes Mitigation may include scaling, caching, rate limiting, or temporarily disabling expensive functionality. --- F. INCIDENT DETECTION FRAMEWORK Critical alerts API availability <99% 5xx spike Database unavailable Authentication failure spike Kubernetes service failure Data replication failure Security detection Backup failure Warning alerts CPU >70% Memory >75% Increasing latency Queue growth DB connection growth Disk utilization >70% Synthetic monitoring Run automated tests every few minutes: Login → Dashboard → Create Appointment → Save → API Response This detects customer-impacting failures before customers report them. Alert fatigue prevention Every alert should answer: > "What action should the on-call engineer take?" If an alert repeatedly generates no actionable response, tune or remove it. --- G. INCIDENT COMMUNICATION MATRIX Severity Internal Leadership Customers Status Page SEV-0 Immediate Immediate Immediate Immediate SEV-1 Immediate Immediate If affected Yes SEV-2 On-call As needed If material Usually SEV-3 On-call No Usually no No SEV-4 Normal workflow No No No Example customer update > Investigating: We are currently investigating an issue affecting appointment management for some customers. Our engineering team is actively working to restore normal service. We will provide another update as soon as we have more information. Avoid speculation. Never say: > "AWS caused the problem." unless that is confirmed. Say: > "We are investigating a potential infrastructure dependency issue." --- H. DR & BUSINESS CONTINUITY BLUEPRINT Recommended initial targets System RTO RPO Authentication 1 hour 15 min Core API 1 hour 15 min PostgreSQL 1 hour 5–15 min Appointment data 1 hour ≤15 min Reporting 8 hours 24 hours Analytics 24 hours 24 hours These are example targets, not universal requirements; they should ultimately be based on contractual, financial, and customer-impact analysis. Backup strategy Use: Automated database backups Point-in-time recovery Cross-region backup copies S3 versioning Backup encryption Backup access controls Regular restore testing Golden rule > A backup that has never been restored successfully is only an assumption of recoverability. --- I. BLAME-FREE POST-INCIDENT REVIEW Incident Incident: API outage Severity: SEV-1 Duration: 47 minutes Timeline 10:02 — Deployment started 10:07 — Error rate increased 10:09 — Alert triggered 10:11 — Incident declared 10:15 — Rollback initiated 10:22 — Error rate declining 10:31 — Core API recovered 10:49 — Validation completed 10:49 — Incident closed Root cause A database query introduced in the deployment generated excessive database load, causing connection exhaustion. Contributing factors Query was not tested against production-scale data. Database connection pool was too large. Alert fired several minutes after customer impact. Rollback process required manual approval. What went well On-call engineer responded quickly. Rollback successfully restored service. Monitoring provided useful database metrics. What didn't go well Detection was delayed. Load testing was insufficient. No automated rollback. Corrective actions Action Owner Priority Add production-scale query testing Engineering High Improve DB saturation alerts SRE High Add automated rollback criteria DevOps High Review connection pool limits SRE Medium No individual blame assigned. --- J. INCIDENT KPI DASHBOARD Track monthly: KPI Target MTTD <5 min MTTA <5 min Mean Time to Mitigate <30 min MTTR <60 min Repeat Incident Rate <10% Change Failure Rate <10% SLA/SLO Breaches Downward trend Alert-to-Incident Ratio Increasing signal quality Customer Impact Duration Downward trend Reporting cadence Daily: Active incident review Weekly: Incident trends Monthly: Executive reliability review Quarterly: DR + incident simulation --- K. INCIDENT READINESS SCORECARD Area Score Detection 7/10 Triage 6/10 Incident Command 4/10 Communication 5/10 Recovery 6/10 Disaster Recovery 5/10 Security Response 5/10 Post-Incident Learning 6/10 Overall Incident Readiness 55/100 — Developing Primary weakness: lack of standardized incident command and response procedures. --- L. 90-DAY INCIDENT RESPONSE ROADMAP Month 1 — Foundation Objectives Standardize severity. Establish incident command. Improve alerting. Deliverables Severity matrix Incident handbook On-call escalation policy Incident channels Communication templates KPI Reduce MTTA by 30%. --- Month 2 — Playbooks & Automation Build playbooks for: Database failure API outage Deployment failure Authentication failure Security incident Automate: Incident creation Slack/Teams channel creation Pager escalation Status-page updates Rollback workflows --- Month 3 — Testing & Optimization Run: GameDay Disaster recovery test Database restore test Security incident simulation Cloud outage simulation Target: MTTR reduction ≥30%. --- M. INCIDENT RESPONSE TEMPLATES Incident Declaration INCIDENT DECLARED Incident ID: Severity: Start Time: Incident Commander: Technical Lead: Affected Service: Customer Impact: Current Symptoms: Confirmed Facts: Hypotheses: Unknowns: Immediate Actions: Next Update: Incident Handoff INCIDENT HANDOFF Current Severity: Current Customer Impact: What We Know: What We Don't Know: Actions Completed: Actions In Progress: Risks: Next Decision: Next Update: Recovery Checklist [ ] Service restored [ ] Error rates normal [ ] Latency normal [ ] Database healthy [ ] Data integrity verified [ ] Synthetic tests passing [ ] Monitoring stable [ ] Customer impact confirmed [ ] Status page updated [ ] Incident Commander approves closure --- N. EXECUTIVE SRE/CISO REPORT Incident Response Summary The organization has reasonable infrastructure and monitoring capabilities but requires a more mature incident-management operating model. Top 10 Incident Risks 1. Database failure 2. Deployment failure 3. Authentication outage 4. Cloud outage 5. Data corruption 6. API failure 7. Security compromise 8. Redis failure 9. Third-party dependency failure 10. Insufficient disaster recovery testing Top 10 Readiness Improvements 1. Formal Incident Commander role 2. Standard SEV framework 3. Incident playbooks 4. Automated escalation 5. Better synthetic monitoring 6. Automated deployment rollback 7. Tested database restoration 8. Security incident procedures 9. Customer communication templates 10. Quarterly incident simulations Top 5 Observability Improvements 1. End-to-end synthetic monitoring 2. Distributed tracing 3. Database performance monitoring 4. SLO-based alerting 5. Customer-impact dashboards Top 5 Recovery Improvements 1. Automated rollback 2. Tested database failover 3. Regular restore drills 4. Cross-region backup 5. Documented recovery procedures Top 5 Communication Improvements 1. Standard incident channel 2. Status-page process 3. Customer notification templates 4. Executive escalation matrix 5. Fixed update intervals Highest-Priority Response Initiative Implement a standardized Incident Command + Severity + Communication system within the next 30 days. First Incident Simulation Scenario: Production PostgreSQL becomes unavailable during peak customer traffic. Test: Detection → SEV-1 declaration → IC assignment → DB failover → data validation → customer communication → recovery → postmortem One Golden Rule > Stabilize first, investigate second, communicate clearly throughout, and never close an incident until recovery has been validated.
🤖 GPT

Saas Incident Response Planner

Add to Cart
Instant accessInstant access
Usage rightsCommercial use
Money-back guaranteeMoney‑back
By purchasing this prompt, you agree to our terms of service
GPT-5.6
Tested icon
Guide icon
4 examples icon
Free credits icon
When a SaaS incident happens, every minute matters. The AI SaaS Incident Response Planner helps SaaS teams build a complete incident management system for detecting, triaging, containing, resolving, and learning from production and security incidents. It creates severity frameworks, incident command structures, response playbooks, communication procedures, disaster recovery strategies, post-incident reviews, and measurable reliability KPIs—giving teams a repeatable operating system for handling
...more
Added over 1 month ago
Report
Browse Marketplace