Always on watch
Every alert source, 24×7. It acknowledges before your phone lights up.
AI SRE employee
EdgeX11 Nightwatch watches your alerts around the clock, owns the incident when one fires, and does the RCA and five-whys before you're awake.
Root cause in minutes. Not hours.
No credit card · connect Alertmanager in minutes.
Alertmanager · checkout-service — error rate 43% Alert
Nightwatch acknowledged. Nobody was paged. Owned
Seen before: INC-2214, 3 months ago — OOMKilled, cache never evicts. Memory
Mitigated. Pods recycled, error rate normal. Resolved
Post-mortem in your inbox — RCA, five whys, the permanent fix. For you
Reads alerts, metrics and logs from
Always on watch
What it does
Every alert source, 24×7. It acknowledges before your phone lights up.
Metrics, Loki logs, K8s state and your cloud — AWS, Azure or GCP — read together, in seconds.
Déjà vu in seconds — it recalls the fix from three months ago.
Not just what broke — why, five levels deep, in a chain you can read.
Safe mitigations run unattended. Production changes wait for your approval.
Timeline, root cause and the permanent fix — delivered, and remembered.
Root-cause analysis
In minutes. Not hours. And it never forgets the answer.
One incident, start to finish
The alert fires. Nightwatch takes it.
Metrics, logs, deploys — and every past incident.
Inside guardrails. Prod changes wait for you.
RCA, five whys, the fix — over coffee.
Human control. AI execution.
FAQ
From the monitoring you already run: Alertmanager and Prometheus natively, plus Datadog, PagerDuty, Loki and plain webhooks — and it reads Kubernetes and your cloud (AWS, Azure or GCP) directly during an investigation. When an alert fires, Nightwatch acknowledges and investigates instead of paging you.
Only inside the guardrails you set. Safe, reversible mitigations run unattended; production rollbacks, config edits and infra changes are prepared and wait for your approval. Everything is logged, and a kill switch stops it all at once.
The incident timeline, the root-cause analysis, the written five-whys chain, what was mitigated, and the permanent fix it recommends. In your inbox by morning — and saved to memory forever.
It works from your organizational memory. Every incident links to the services, deploys and fixes around it — so when something fires that looks like three months ago, it knows in seconds.
Only when a remediation needs your approval, or when something is genuinely new. Everything else becomes a morning read instead of a 2:47 AM page.
It runs as a Reliability employee on your plan — priced per employee, never per seat. See pricing.