Incident channels: PagerDuty, JSM/Opsgenie, ServiceNow and the rest
Every channel on this page sits on the architecture notifications documents: a channel is one delivery destination, a rule decides which event reaches it, and delivery is fail-open so a channel problem can never block a backup. This page covers the Jira Service Management/Opsgenie and ServiceNow Event Management channel kinds built on it, the recovery event that tells PagerDuty and those two channels to close what the matching alert opened, and mapping guidance for three more platforms that do not read our body shape natively.
How a recovery closes an incident
Four alert conditions can also clear: a failing or stale downpipe that returns to health, a downpipe whose replication falls below its configured copy count and then recovers, and a run that was about to age out of history while a replica lacked it and is no longer at risk.
The engine detects the falling edge of each of those four conditions from the cooldown state it keeps to avoid re-paging on every tick, and emits a second, matching event on the first alert pass after the condition clears: the identical event name and downpipe as the original alert, marked as a recovery, with the same severity the original alert used (so the recovery clears the same rule threshold the trigger did). The engine’s cron runs that pass every 15 minutes. That one field, and the stable pairing of an event with its downpipe, is the entire mechanism. A channel that has a concept of closing something it opened reads that pairing and closes it; a channel that does not just delivers another line.
When the cron has not run for more than 45 minutes, the engine’s own alarm starts a backstop alert sweep. That sweep cannot send a notification. Before engine 0.3.6, if the sweep found that a failing or stale downpipe had recovered, the resolve was lost. From engine 0.3.6, the backstop keeps the resolve owed, and the first cron pass after the cron starts again sends it.
| Channel | Recognises a recovery? | What happens |
|---|---|---|
| PagerDuty | Yes | Sends event_action: "resolve" against the same dedup_key the trigger used, closing the matching incident. |
| Jira Service Management / Opsgenie | Yes | Closes the alert by the same alias the trigger’s create used. |
| ServiceNow Event Management | Yes, at the Alert layer | Sends severity 0 (Clear) against the same message_key; see the warning below for what that does and does not do at the Incident layer. |
| Generic webhook | Yes, from engine 0.3.6 | Delivers the recovery with recovered: true and the same dedupKey the trigger carried. Your receiver maps that to a close; see the mapping recipes below. Before engine 0.3.6, only the detail text marks a recovery. |
| Slack, Microsoft Teams, email | No distinct signal | Delivers the recovery as another message. The detail text says the condition cleared, in plain English; there is no boolean field a receiver can key on. |
PagerDuty: trigger and resolve
You add a PagerDuty channel as notifications describes: a routing key, nothing else. A recovery then closes the incident its trigger opened. Every emission carries a dedup_key from downpipe id and event name (or the event name alone for an account-level event like a canary), and a recovery reuses that key with event_action: "resolve" instead of "trigger". PagerDuty’s own Events API v2 correlates the two, so the incident a failure opened is the exact incident its later recovery closes; there is nothing to configure on the PagerDuty side beyond the integration you already have.
Jira Service Management and Opsgenie
JSM carries Opsgenie’s alerting engine, and both accept the same Alert API shape. Opsgenie itself is sunsetting in April 2027, so this channel is built against the JSM endpoint contract, which both providers honour identically for the life of that migration. One channel client covers either provider; which one you are pointed at is simply whichever alert-create URL you configure.
Unlike PagerDuty’s single trigger-or-resolve envelope, the Alert API is a create, then a separate close-by-alias: a new or continuing condition posts a create (a re-post with the same alias de-duplicates onto the still-open alert rather than opening a second one), and a recovery posts to that alias’s close endpoint. The alias is the same downpipe-plus-event pairing PagerDuty’s dedup_key uses, so an operator running both channels sees identical correlation behaviour across them.
| Our severity | JSM/Opsgenie priority |
|---|---|
| critical | P1 |
| warning | P3 |
| info | P5 |
P2 and P4 are left as headroom for your own hand-raised alerts to sit above or below ours without collision.
Adding, replacing or deleting a channel needs a fresh step-up re-authentication, and so does adding, replacing or deleting a rule. Both directions are gated deliberately: the confidentiality argument does not apply, because a notify emission is redaction-safe and repointing a channel steals nothing, but silencing an account is the move an attacker makes before an attempt rather than after it, and suppression by replacement needs no delete at all. The console asks for a fresh passkey assertion and retries with the token it gets back (step-up re-authentication). For a session that Cloudflare Access signed in, the engine asks for a fresh Access sign-in instead.
Add the channel on the Notifications screen, the same Channels tab notifications covers. Pick the kind “Jira Service Management / Opsgenie (Alert API)”, give it a name, the HTTPS alert-create URL (for Opsgenie, https://api.opsgenie.com/v2/alerts; for a JSM Cloud tenant, its own alert-api base), and the GenieKey API token. The token field is write-only: it is sent once over the authenticated channel, then sealed by your engine and never re-displayed, so on a later edit you leave it blank to keep the current token. The engine keeps the token only while the URL is unchanged. Paste a new one to rotate it or when you change the URL. Once saved, the channel is selectable in a rule exactly like any other, on the Rules tab.
Or add it against the API
Anyone holding notify.config, which is Operator, Approver and Owner, can create the same channel directly, the route the console calls:
POST /admin/notify/channels HTTP/1.1
Host: console.example.com
Content-Type: application/json
{
"kind": "jsm",
"name": "Opsgenie: platform team",
"url": "https://api.opsgenie.com/v2/alerts",
"apiKey": "<your GenieKey API token>",
"enabled": true
}The GenieKey token rides only in the Authorization header on send, never in the alert body. At rest, the engine seals it under its configuration wrap key with the token’s own domain-separation label.
ServiceNow Event Management
ServiceNow’s Event Management module ingests a POST to its em_event table API, and its correlation engine groups events sharing a message_key into one Alert. A later event on the same key at a lower severity updates that Alert, and severity 0 clears it. There is no separate close endpoint the way PagerDuty and JSM have one: closing is posting another event with the same key at severity 0, so create and recover share the identical shape here, varying only severity and description.
| Our severity or state | ServiceNow severity |
|---|---|
| critical | 1 |
| warning | 4 |
| info | 5 |
| recovered (any original severity) | 0 (Clear) |
Severities 2 (Major) and 3 (Minor) are left unused, reserved for a human to hand-adjust an Alert’s severity inside ServiceNow itself. Authentication is HTTP Basic: a username (not a secret, always resupplied) and a password, sealed at rest the same way the JSM token is, under a different domain-separation label so one channel’s credential can never be opened as the other’s.
Severity 0 clears the Alert. Whether the Incident auto-closes is yours to configure
Reaching severity 0 clears the ServiceNow Alert this channel creates. Whether that Alert’s associated Incident (if Event Management is configured to open one) also auto-closes depends entirely on your own ITOM or Event Management business rules, the “auto-close incident on alert-clear” property that lives in your ServiceNow instance. This is a setting in your ServiceNow instance, and downpipes does not control it. If you want a recovered backup to close an Incident, not only the Alert beneath it, set that up on the ServiceNow side; this channel’s part of the contract ends at clearing the Alert.
Add it the same way as JSM, on the Notifications Channels tab, choosing the kind “ServiceNow Event Management”. It carries the em_event table URL, the Basic-auth username (not a secret, so it is prefilled on an edit and always resupplied), and the password. The password is write-only like the JSM token: sent once, then sealed, and left blank on a later edit to keep the current one. Because the password is kept only while the URL is unchanged, repointing the URL means resupplying it.
Or add it against the API
The equivalent direct call:
POST /admin/notify/channels HTTP/1.1
Host: console.example.com
Content-Type: application/json
{
"kind": "servicenow",
"name": "ServiceNow: platform events",
"url": "https://yourinstance.service-now.com/api/now/table/em_event",
"username": "downpipes-integration",
"apiKey": "<the account's password>",
"enabled": true
}Generic webhook: mapping recipes for three more platforms
A generic webhook channel posts a small, versioned JSON body to your own URL on every matching event (the full security model is described here):
{
"kind": "downpipe-event-v1",
"at": "2026-07-05T02:14:08.221Z",
"event": "backup-failure",
"severity": "critical",
"downpipe": { "id": "dp_8f2a1", "name": "Prod KV" },
"detail": "Prod KV last run failed",
"dedupKey": "downpipe:dp_8f2a1:backup-failure"
}
From engine 0.3.6, two fields in this body matter for mapping it into another platform. The first is dedupKey: downpipe:<id>:<event>, or account:<event> for an account-level event. A trigger and its later recovery carry the same value. PagerDuty’s dedup_key, JSM’s alias and ServiceNow’s message_key use the same key. Point a platform that lets you choose its grouping or deduplication key at this field.
The second is recovered. A recovery carries recovered: true, and a trigger does not carry the field. A recovery is otherwise another POST of the identical event, severity and downpipe.id, so key the close on recovered and not on those fields.
These two fields close the four recovering events in the table below only. They do not close a canary alert. canary-recovered is a separate event with no recovered field, and its dedupKey is account:canary-recovered, not account:canary-dead. The recipes below therefore do not close a canary-dead alert, and they send canary-recovered as a new alert.
To close a canary alert, close it by hand, or add a rule to your mapping: for the event canary-recovered, resolve the alert with the key account:canary-dead. The canary uses one key for all destinations. That rule closes the alert when one destination recovers, even if another destination is still dead.
An account-level event has one key for each event name. Repeated alerts of one account-level event, for example two recovery-code-used events, group onto one open alert, and no recovery closes it.
Before engine 0.3.6, the body has neither field. Build the key from event plus downpipe.id, and tell a recovery only by its detail wording, which varies per event:
| Event | Recovery detail reads |
|---|---|
backup-failure | “{name}: the last run succeeded; the previous failure has cleared” |
backup-stale | “{name}: a recent run succeeded; the staleness has cleared” |
replication-degraded | “{name}: replication recovered; N of M copies proven, every destination reachable” |
run-at-risk-eviction | “{name}: the at-risk run is no longer exposed to eviction; N of M copies proven, every destination reachable” |
Two more recovery wordings apply to either replication event. A downpipe edited back to one destination sends “{name}: now configured for a single destination; the replication alert no longer applies and is cleared”. A downpipe with no recent successful run sends “{name}: no recent successful run to assess replication; the prior replication alert is cleared (backup health is reported separately)”.
This wording is prose, not a versioned field. On engine 0.3.6 or later, use recovered instead.
incident.io
Its custom HTTP alert source is the cleanest of the three to wire up, because its transform expression runs directly against whatever JSON you send it, no relay required. Point the source at this webhook’s URL, and in the transform (plain JavaScript, the incoming body available as $) compute a firing-or-resolved status from the fields above:
return {
title: $.detail,
status: $.recovered === true ? "resolved" : "firing",
description: $.detail,
metadata: { event: $.event, downpipe_id: $.downpipe ? $.downpipe.id : null },
};
Set the deduplication key to the dedupKey field. An alert in incident.io resolves when a later event with the same deduplication key reports status: "resolved". The expression above produces that status for a recovery. Before engine 0.3.6, compute the status from the detail wording in the table above, and use a deduplication key unique per event-and-downpipe pair.
Grafana OnCall / IRM
Its Formatted Webhook integration recognises a fixed field set: alert_uid (the grouping key), title, state (ok or alerting, the field that drives auto-resolution), and message. Our webhook body does not use those field names, so reaching this path means an intermediary you control (a small relay, for example a Cloudflare Worker, that receives our downpipe-event-v1 body and re-posts Grafana’s shape), mapping alert_uid from dedupKey, state to ok when recovered is true (otherwise alerting), and message from detail. Before engine 0.3.6, map alert_uid from event plus downpipe.id and state from the recovery-wording check above. Grafana’s raw “Webhook” integration type accepts arbitrary JSON directly with no relay, but check Grafana’s own current documentation for whether its Jinja2 templating can compute a resolve signal from an arbitrary field on that path before relying on it, or use the Formatted Webhook path above.
Splunk On-Call
Its REST endpoint integration expects its own fixed fields too: message_type (one of CRITICAL, WARNING, ACKNOWLEDGEMENT, INFO or RECOVERY), entity_id (the field that must stay identical across an incident’s whole lifecycle), entity_display_name and state_message. As with Grafana’s Formatted Webhook, our body’s field names do not match, so reaching auto-resolve here needs the same kind of relay: map entity_id from dedupKey, message_type to RECOVERY when recovered is true and otherwise from severity (critical maps to CRITICAL, warning to WARNING, info to INFO), entity_display_name from downpipe.name, and state_message from detail. A RECOVERY message against the same entity_id closes the incident Splunk On-Call opened for it. Before engine 0.3.6, map entity_id from event plus downpipe.id and emit RECOVERY when the detail-wording check above matches.
Where a recipe above needs a relay, it is a small piece of code you own and can inspect, not a dependency this product adds. The alternative, simpler answer for any of these three platforms is to run the webhook fire-only, the same as Slack and Teams below, and let a human close what it opens.
Slack and Teams: notification fan-out, not an incident system
Slack and Microsoft Teams are what notifications describes: a message to a channel or an incoming-webhook connector, built from the same redaction-safe detail line. Neither carries any notion of an open or closed incident, so a recovery reaches them as a second message, worded as a recovery, and nothing closes on either platform’s side. Treat both as where a human reads that something happened, never as the system of record for whether it is still happening. Where you want an actual incident that a recovery can close, wire one of the channels above instead, or fan out to Slack or Teams alongside it for visibility.
Where this fits
For adding any of the seven channel kinds ChannelKind allows (engine/src/notify/types.ts), including the Jira Service Management/Opsgenie and ServiceNow kinds this page covers, writing the rules that route events to them, and the delivery-history view, read notifications. For the catalogue of the events the engine emits, including the four whose recovery this page covers, read alert events. For the egress screening every URL-bearing channel here inherits, read securing notification webhooks. For the two-audience framing this page sits under, read the monitoring overview.
Last updated .