Production Ops
Production Operations Weekly Review
Jul 18 to Jul 26, 2026 / cadence: weekly
Overall statusGREEN
Degraded days0
RED components0
AMBER components0
Actions needing disposition0
Daily Production Health
No daily health artifacts were available for this period.
Latest daily health issues
- No daily health issues listed on the latest collected day.
Daily Findings
No chartable points.
Reliability Movement
Customer visible failures+3
3 current / 0 previous
Observed incidents+3
3 current / 0 previous
Application alerts+21
112 current / 91 previous
Infrastructure alerts-69
0 current / 69 previous
Pingdom events+3
3 current / 0 previous
Latency Trend
No chartable points.
Component Heatmap
No component trend data was available.
Production Dependency Map
| Component | Status | Current period | Source | References | Disposition ID |
|---|---|---|---|---|---|
| Customer Edge4 components | |||||
| Pingdom public checks | GREEN | No degraded signal in collected evidence | pingdom | component-pingdom-public-checks | |
| DNS resolution | GREEN | No degraded signal in collected evidence | dns_issue_check | component-dns-resolution | |
| Ingress / Traefik | GREEN | No degraded signal in collected evidence | service_5xx_rate_pct, service_5xx_rps, service_p95_top | component-ingress-traefik | |
| Public 5xx and latency | GREEN | No degraded signal in collected evidence | service_5xx_rate_pct, latency_triage, latency_signal | component-public-5xx-and-latency | |
| Application Runtime5 components | |||||
| PHP-FPM | GREEN | No degraded signal in collected evidence | web_probe_failures_5m, web_restarts_5m, latency_triage | component-php-fpm | |
| API/app pods | GREEN | No degraded signal in collected evidence | web_ready_pods, web_running_pods, web_unavailable_replicas | component-api-app-pods | |
| Worker pods | GREEN | No degraded signal in collected evidence | top_restarts_24h, top_memory_containers_24h | component-worker-pods | |
| Runtime memory and restarts | GREEN | No degraded signal in collected evidence | top_memory_containers_24h, top_restarts_24h | component-runtime-memory-and-restarts | |
| Slow request families | GREEN | No degraded signal in collected evidence | latency_triage, slow_traces | component-slow-request-families | |
| Data Layer5 components | |||||
| MySQL catalog | GREEN | No degraded signal in collected evidence | rds:mysql-catalog, slowquery:mysql-catalog | component-mysql-catalog | |
| MySQL catalog2 | GREEN | No degraded signal in collected evidence | rds:mysql-catalog2, slowquery:mysql-catalog2 | component-mysql-catalog2 | |
| MySQL master | GREEN | No degraded signal in collected evidence | rds:mysql-master, slowquery:mysql-master | component-mysql-master | |
| Postgres billing | GREEN | No degraded signal in collected evidence | rds:postgres-billing, slowquery:postgres-billing | component-postgres-billing | |
| Slow queries | GREEN | No degraded signal in collected evidence | slowquery | component-slow-queries | |
| Cache & Messaging3 components | |||||
| Redis / Sentinel | GREEN | No degraded signal in collected evidence | redis_issue_check, redis_evicted_keys_5m, redis_rejected_connections_5m | component-redis-sentinel | |
| RabbitMQ / queues | GREEN | No degraded signal in collected evidence | rabbitmq, queue_depth, consumer_lag | component-rabbitmq-queues | |
| Consumer lag or delayed processing | GREEN | No degraded signal in collected evidence | consumer_lag, cron_active | component-consumer-lag-or-delayed-processing | |
| Batch & Scheduled Work3 components | |||||
| Watched CronJobs | GREEN | No degraded signal in collected evidence | cron_active | component-watched-cronjobs | |
| Kubernetes Jobs | GREEN | No degraded signal in collected evidence | KubeJobFailed, job_failed | component-kubernetes-jobs | |
| Long-running scheduled tasks | GREEN | No degraded signal in collected evidence | cron_active, slow_traces | component-long-running-scheduled-tasks | |
| Kubernetes & Capacity5 components | |||||
| Readiness | GREEN | No degraded signal in collected evidence | web_ready_pods, web_unavailable_replicas | component-readiness | |
| Probe failures | GREEN | No degraded signal in collected evidence | web_probe_failures_5m | component-probe-failures | |
| HPA saturation | GREEN | No degraded signal in collected evidence | hpa_current, hpa_max | component-hpa-saturation | |
| Pod/node churn | GREEN | No degraded signal in collected evidence | top_restarts_24h | component-pod-node-churn | |
| Top memory containers | GREEN | No degraded signal in collected evidence | top_memory_containers_24h | component-top-memory-containers | |
| Observability & Alerting5 components | |||||
| Grafana / Prometheus | GREEN | No degraded signal in collected evidence | prometheus | component-grafana-prometheus | |
| Loki | GREEN | No degraded signal in collected evidence | loki, latency_triage | component-loki | |
| Tempo | GREEN | No degraded signal in collected evidence | slow_traces | component-tempo | |
| Alert rule health/noise | GREEN | No degraded signal in collected evidence | slack_alerts, alert_table | component-alert-rule-health-noise | |
| AWS alarm emails | GREEN | No degraded signal in collected evidence | aws_email_alerts, email_table | component-aws-alarm-emails | |
| External Dependencies5 components | |||||
| Email provider | GREEN | No degraded signal in collected evidence | email_provider | component-email-provider | |
| SMS provider | GREEN | No degraded signal in collected evidence | sms_provider | component-sms-provider | |
| Payment provider | GREEN | No degraded signal in collected evidence | payment_provider | component-payment-provider | |
| Identity/login provider | GREEN | No degraded signal in collected evidence | identity_provider | component-identity-login-provider | |
| Webhooks / third-party APIs | GREEN | No degraded signal in collected evidence | third_party_api | component-webhooks-third-party-apis | |
Previous period
| Metric | Current | Previous | Delta |
|---|---|---|---|
| active_aws_alarms | 0 | 0 | 0 |
| application_alerts | 112 | 91 | +21 |
| customer_incidents_confirmed | 0 | 0 | 0 |
| customer_incidents_observed | 3 | 0 | +3 |
| customer_visible_failures | 3 | 0 | +3 |
| impacted_services | 10 | 20 | -10 |
| infrastructure_alerts | 0 | 69 | -69 |
| pingdom_downtime_minutes | 3 | 0 | +3 |
| pingdom_events | 3 | 0 | +3 |
ADS Action Queue
Missing dispositions: none
| Status | Action | Domain | ID |
|---|
Source Coverage
| Source | Status | Detail |
|---|---|---|
| Daily health JSON | missing | 0 daily artifact(s) found |
| Production reliability dashboard | ok | Weekly alerts/reliability artifact |
| Team weekly report | ok | Delivery/deploy evidence |
| Engineering council test report | ok | Test/smoke evidence |
| AWS posture evidence | ok | Cost/security/recommendation evidence |
| Action register | warning | ADS/accepted-risk/false-positive dispositions |
Evidence References
Source reports
Pingdom
Slack
Slack alert: TraefikServiceHighErrorRateSlack alert: TraefikServiceHighErrorRateSlack alert: TraefikServiceHighErrorRateSlack alert: TraefikServiceHighLatencyCriticalSlack alert: CriticalPagerDutyTestSlack alert: KubeAggregatedAPIDownSlack alert: WatchdogSlack alert: KubeletServerCertificateExpirationSlack alert: KubeAPIErrorBudgetBurnSlack alert: RESOLVED - KubeAPIErrorBudgetBurnSlack alert: RESOLVED - KubeletServerCertificateExpirationSlack alert: RESOLVED - KubeClientErrors
Reliability
- application alerts
- 112
- infrastructure alerts
- 0
- active aws alarms
- 0
- degraded days
- 0
- red components
- 0
- amber components
- 0
Customer Impact
- customer visible failures
- 3
- customer incidents observed
- 3
- customer incidents confirmed
- 0
- pingdom events
- 3
- pingdom downtime minutes
- 3
Delivery Health
- production bugs closed
- 17
- delivery items
- 28
- deployments
- 8
- test runs
- 8
- test pass rate
- 0
- smoke attempts
- 12
- smoke failed
- 0
Cost
- total
- 0
- currency
- USD
- forecast
- 0
Security
- security hub score
- 0
- critical findings
- 0
- high findings
- 0
- guardduty findings
- 1
- inspector critical
- 0
- iam external access
- 0
Aws Recommendations
- trusted advisor red
- 0
- trusted advisor yellow
- 0
- compute optimizer savings
- 0
- cost optimization savings
- 0
- well architected high risk issues
- 0
- well architected medium risk issues
- 0
Backup
- failed jobs
- 0
- protected resources
- 0