Skip to content
downpipes docs

Backup health for your monitoring stack

Two different questions get asked of the same engine. A security or compliance reader asks who did what, from where, and whether the record can be trusted: that is the audit feed, covered in wiring the audit feed into your SIEM and forwarding the audit log to your SIEM (push). An on-call engineer asks a narrower, more urgent question: is this backup healthy right now, and will something tell me the moment it stops being healthy. That second question is what this page, and the three pages under it, answer.

downpipes treats the two as separate surfaces read off the same engine state, not one feed wearing two labels. The audit feed is a hash-chained record of operator action: who signed in, who approved a restore, who changed a role. Monitoring is a live read of backup outcomes: whether the last run succeeded, how long it took, how large the archive was, and whether each destination is reachable. Wiring one does nothing for the other, and neither substitutes for it. A self-hoster running a real monitoring stack alongside downpipes typically wants both, for two different readers.

The two audiences, side by side

ReaderQuestionSurfaceDocumented in
Security or complianceWho did what, from where, and can I prove the record is unalteredThe hash-chained audit feed, by pull or by pushWiring the audit feed into your SIEM, forwarding the audit log to your SIEM (push)
SRE or operationsIs this backup healthy right now, and am I paged the moment it is notBackup-health metrics, pushed or scraped, and the incident channels behind an alertThis page’s three children, plus notifications and alert events

The two surfaces do not share fields by design. The audit feed’s events carry member emails, source IPs, roles and approver emails on purpose, because attributing who did what is the point of an audit trail, so the feed is not free of personal data. The monitoring surfaces below carry none of that: a metric or a push body names a downpipe, a destination, a duration and a byte count, never a person.

Built on the notification channels

Monitoring builds on notifications. PagerDuty, the generic webhook, Slack, Microsoft Teams and email are channels there. The rules that decide which event reaches which channel, the digesting of success-class events, and the delivery history all work as that page describes. Monitoring uses two more channel kinds on the same architecture: Jira Service Management and Opsgenie share one client, and ServiceNow Event Management is the other. A recovered backup emits a matching recovery event, so a channel that can close what it opened closes it. Incident channels covers all of that, including which channels can close an incident and which are fire-only by design.

The metrics surfaces work the same way. The Prometheus /metrics endpoint and OTLP metrics push both read the identical run-history and replication state that drives the backup-failure and backup-stale alerts described in alert events; they are a second way to observe the same facts, not a parallel source of truth.

The auto-parse effort, in three tiers

Standing up a monitoring stack usually means teaching it a new shape: a custom parser, a field mapping, a dashboard built from scratch. downpipes keeps that work small, and the effort falls into three tiers.

TierWhat it costs youWhere it applies here
1: zero setupThe wire format needs no mapping: any Prometheus-compatible tool already parses it natively. Grafana Cloud’s hosted scraper is the one agentless case. Every other tool still means pointing your own agent or collector at the endpoint, at effort that varies by vendor from a few config fields to a YAML mapping file (the per-vendor breakdown).The Prometheus /metrics scrape surface.
2: one formFill in one short form in the console: a URL, a credential, or both.OTLP push’s endpoint URL, auth header name and auth secret; PagerDuty’s routing key; a generic webhook URL; the Jira Service Management/Opsgenie channel’s URL and token; the ServiceNow channel’s URL, username and password.
3: one-time mappingMap a handful of fields once, in the receiving platform’s own console.Pointing the generic webhook at Splunk On-Call, Grafana OnCall/IRM or incident.io, none of which share our field names natively.

Tier 1 is the most valuable of the three, because a Prometheus-compatible scraper needs no translation at all: the metric names, types and labels on the wire are the ones any of these tools already expects from any exporter. It is also the one tier with a catch to know before you commit a scrape config to it. The metrics page sets out that catch.

Every surface below is set up from the console: the metrics credential and the OTLP push destination live on the Integrations screen, each on its own vendor tile, and the Jira Service Management/Opsgenie and ServiceNow incident channels are kinds you add on the Notifications screen. Settings keeps only the diagnostics credential, which is a support-access concern rather than an integration. Each page leads with that console path and gives the admin API call underneath it, for when you would rather script the setup.

Where this fits

For the channel and rule mechanics every incident channel sits on top of, read notifications. For the catalogue of the events the engine emits, read alert events. For the audit-feed side of the same engine, read wiring the audit feed into your SIEM and forwarding the audit log to your SIEM (push).

Last updated .