Ops Briefing Surface

Production Reliability Dashboard

Generated 2026-07-19 22:19 for 2026-07-11 07:00 to 2026-07-18 07:00 from Pingdom checks, Slack #_alerts_prod, and AWS SNS alerts.

All sources Pingdom customer checks Slack alert families AWS alarm emails
Email-confirmed customer incidents0Pingdom down/slow events confirmed by inbox alertsObserved in Pingdom: 1
Impacted services15Mapped from Slack and Pingdom evidence
AWS alarms in ALARM0Still alarming at window end
Latest observed signal2026-07-17 23:05Most recent cross-source activity
Executive summary

What needs attention

Bottom line: Pingdom observed recent customer-facing glitches (email unconfirmed) and application-level critical paths are present.

Pingdom customer impact

External signal
No criticalActive: 1Total seen: 1

1 active item(s) in this window.

Slack impacted services

Application signal
5 criticalActive: 14Total seen: 14

5 critical, 9 non-critical active item(s).

AWS alarms

Infrastructure signal
No criticalActive: 0Total seen: 0

No AWS alarm emails were captured in this window.

No active issue listed in this category.

What to do next

  1. NextUse Pingdom check Adservio Ro as churn/degradation evidence, and escalate to customer-facing incident priority only after inbox confirmation or strong cross-source correlation

    Use Pingdom check Adservio Ro as churn/degradation evidence, and escalate to customer-facing incident priority only after inbox confirmation or strong cross-source correlation.

    Executive recommendation
  2. ThenTreat critical services accommodations-api, grafana, core-grafana-80, web-80, etcd as the primary application investigation set

    Treat critical services accommodations-api, grafana, core-grafana-80, web-80, etcd as the primary application investigation set. Use TraefikServiceHighErrorRate as the leading category, reproduce the failing paths on accommodations-api, grafana, core-grafana-80, web-80, etcd, compare against the latest deploy or config change, and do not close it until both error rate and latency flatten.

    Executive recommendation
Global Evidence Explorer

Global Evidence Explorer

Report-wide charts and tables stay here, separate from the active investigation scope.

Application + Infrastructure Alerts by Day

Pingdom latency + downtime by source

Customer View

Pingdom Checks

Pingdom CheckStatusEventsDowntimeLast SeenLikely ServicesCorrelated Evidence
Adservio RoRecovered recently10m2026-07-17 18:05unclassified

Pingdom rows show externally visible signal first. The correlated evidence column helps tie the failing check back to services, Slack alert families, or AWS alarms when those links exist.

Application View

Slack Impacted Service / Resource View

This view attributes alerts to the workload or resource named in the alert text. Grafana, Loki, and Tempo are treated as observability components and are excluded when a more specific impacted target is also present.

Impacted Service / ResourceHighest SeverityCountLast SeenStatusTop Alert TypesDiscussion SignalLatest Thread Note
grafanaCritical172026-07-17 23:05Seen todayPlatformLatencyP95Critical1s (1)PagerDutyTest (1)Watchdog (6)KubeHpaMaxedOut (2)PlatformLatencyP95Warning400ms (1)NoneNo thread note
accommodations-api
Grouped 4 variantsVariant mentions 14Active variants 4
Critical142026-07-16 01:04Recent (72h)TraefikServiceHighErrorRate (10)KubePodContainerRestartingFrequently (2)KubePodCrashLooping (1)KubeDeploymentReplicasMismatch (1)General investigation
Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a…
etcdCritical72026-07-15 18:50Recent (72h)etcdInsufficientMembers (1)TargetDown (3)etcdMembersDown (3)Release / migration issueGeneral investigation
tuiasi is having metrics server issues and these are all false alarms . will fix it with a new ticket
web-80Critical22026-07-16 01:00Recent (72h)TraefikServiceHighErrorRate (1)TraefikServiceHighLatency (1)NoneNo thread note
core-grafana-80Critical12026-07-16 12:21Recent (72h)TraefikServiceHighErrorRate (1)General investigation
This should go away caused due codex agent running queries
admission-apiWarning192026-07-17 22:31Seen todayTraefikServiceHighLatency (18)RESOLVED - TraefikServiceHighLatency (1)NoneNo thread note
core-grafanaWarning22026-07-17 16:23Seen todayKubePodContainerRestartingFrequently (2)NoneNo thread note
library-api
Grouped 5 variantsVariant mentions 8Active variants 5
Warning52026-07-16 01:12Recent (72h)KubePodCrashLooping (2)KubePodContainerRestartingFrequently (2)KubeDeploymentReplicasMismatch (1)General investigation
Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a…
rooms-api
Grouped 2 variantsVariant mentions 3Active variants 2
Warning42026-07-16 01:04Recent (72h)KubePodContainerRestartingFrequently (2)KubePodCrashLooping (1)KubeDeploymentReplicasMismatch (1)General investigation
Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a…
send-codes
Grouped 2 variantsVariant mentions 3Active variants 2
Warning32026-07-16 05:25Recent (72h)KubeJobFailed (2)KubePodContainerRestartingFrequently (1)NoneNo thread note
metrics-serverWarning32026-07-15 18:49Recent (72h)KubeAggregatedAPIDown (3)NoneNo thread note
colecteaza-sms-note-absWarning12026-07-16 01:04Recent (72h)KubePodContainerRestartingFrequently (1)NoneNo thread note
download-albumWarning12026-07-16 01:04Recent (72h)KubePodContainerRestartingFrequently (1)NoneNo thread note
rezumatWarning12026-07-16 01:04Recent (72h)KubePodContainerRestartingFrequently (1)NoneNo thread note
Evidence

Slack Alert Families

AlertSeverityCountLast SeenStatusThreadsTop Impacted ServicesDiscussion SignalLatest Thread Note
TraefikServiceHighErrorRateCritical122026-07-16 12:21Recent (72h)1accommodations-api (10)web-80 (1)core-grafana-80 (1)General investigation
This should go away caused due codex agent running queries
PlatformLatencyP95Critical1sCritical12026-07-16 01:05Recent (72h)0grafana (1)None
PagerDutyTestCritical12026-07-15 10:41Recent (72h)0grafana (1)None
etcdInsufficientMembersCritical12026-07-15 10:40Recent (72h)1etcd (1)Release / migration issue
Looks good i did a helm deployed on tuiasi that rolled out alert manger pod and these alerts are triggered all good on tuiasi cluster . | false positive alerts.
TraefikServiceHighLatencyWarning192026-07-17 22:26Seen today0admission-api (18)web-80 (1)None
WatchdogWarning62026-07-17 23:05Seen today0grafana (6)None
KubePodContainerRestartingFrequentlyWarning42026-07-17 16:23Seen today0accommodations-api (2)library-api (2)rooms-api (2)core-grafana (2)colecteaza-sms-note-abs (1)None
RESOLVED - TraefikServiceHighLatencyWarning12026-07-17 22:31Seen today0admission-api (1)None
TargetDownWarning32026-07-15 18:50Recent (72h)1etcd (3)General investigation
tuiasi is having metrics server issues and these are all false alarms . will fix it with a new ticket
etcdMembersDownWarning32026-07-15 18:50Recent (72h)0etcd (3)None
KubeAggregatedAPIDownWarning32026-07-15 18:49Recent (72h)0metrics-server (3)None
KubeJobFailedWarning22026-07-16 05:25Recent (72h)0send-codes (2)None
KubePodCrashLoopingWarning22026-07-16 01:12Recent (72h)1library-api (2)accommodations-api (1)rooms-api (1)General investigation
Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a…
KubeHpaMaxedOutWarning22026-07-16 01:12Recent (72h)0grafana (2)None
PlatformLatencyP95Warning400msWarning12026-07-16 01:06Recent (72h)0grafana (1)None
KubeAPIErrorBudgetBurnWarning52026-07-14 11:53Seen this week0grafana (5)None
AlertmanagerFailedToSendAlertsWarning12026-07-13 07:42Seen this week0grafana (1)None
KubeDeploymentReplicasMismatchWarning12026-07-13 07:41Seen this week1accommodations-api (1)library-api (1)rooms-api (1)General investigation
Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a…

Status is heuristic. Slack rarely posts explicit resolutions, so “Seen today” or “Recent” means the alert family still appeared in production recently, not that it is definitely unresolved.

AWS Email Alarm Families

AWS AlarmEmailsALARMOKState FlipsFirst SeenLast SeenLatest StateStatus

“Flapping, latest OK” means the most recent email was an OK, but the alarm toggled repeatedly and is still a reliability concern.

Global Discussion-Derived Signal

Thread DateAlertSeverityServicesSignalKey Notes
2026-07-16 12:21TraefikServiceHighErrorRateCriticalcore-grafana-80General investigation
This should go away caused due codex agent running queries
2026-07-15 14:45TargetDownWarningetcdGeneral investigation
tuiasi is having metrics server issues and these are all false alarms . will fix it with a new ticket
2026-07-15 10:40etcdInsufficientMembersCriticaletcdRelease / migration issue
Looks good i did a helm deployed on tuiasi that rolled out alert manger pod and these alerts are triggered all good on tuiasi cluster . | false positive alerts.
2026-07-13 07:41KubePodCrashLoopingWarningaccommodations-api, library-api, rooms-apiGeneral investigation
Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a…
2026-07-13 07:41KubeDeploymentReplicasMismatchWarningaccommodations-api, library-api, rooms-apiGeneral investigation
Both pods are fine now accommodations-api , library-api and rooms api recovered on their own about 1.5–2h ago and have been stable since (a…