Skip to content
downpipes docs

Engine error and status codes

This page is the lookup for the HTTP outcomes the engine admin API returns: the status codes, the shape of each error body, what causes it, and what to do about it. It is for a developer integrating with the engine directly rather than through the console.

Hold this first: the content type is part of the contract, not an accident, and it decides an operator’s day at the 401. A plaintext body carrying unauthorised is a genuine authentication failure and the caller should sign in again. A JSON body carrying stepUpRequired at that same 401 status is a gated route asking a still-valid session for a fresh identity proof, and a client that reads it as a lapsed session signs out an operator who was never signed out. A failed capability check is a JSON 403. This page is the HTTP-channel view of the admin surface; for how a caller authenticates and the full role-by-capability matrix, see authentication and authorisation, and for the offline reader’s process exit codes (a separate contract) see the CLI exit codes.

The status channels at a glance

The channels below carry the admin API. A few statuses sit outside this table. The /support/diagnostics and /support/audit-feed pull routes answer 503 when the engine cannot read their grant, seal the diagnostics bundle or read the audit log. The SCIM endpoint answers 503 when SCIM is not enabled. Some admin routes answer 502 when the engine cannot read its own scheduler state. A request over TLS 1.1 or older gets a plaintext 426 before any route runs.

StatusBodyContent typeWhat it means
200JSON resultapplication/jsonThe request succeeded. Some routes that report rather than mutate are always 200, with a payload that states the miss when there is one
202{ ownerActionQueued: true, id, status: "pending" }application/jsonAn owner action was recorded for a second owner’s approval and has not been carried out yet. This is a not-yet-acted outcome, never a failure
204emptynoneThe CORS preflight (OPTIONS) answer. No body
400{ error: "<reason>" }application/jsonThe request was malformed or the input was rejected. The reason names the problem
401unauthorisedtext/plainAuthentication failed or no usable credential was presented. This is the sign-in-again channel
401{ error: "step-up required", stepUpRequired: true }application/jsonA step-up gated route wants a fresh identity proof. The session is still valid, so this must never route the caller to a sign-in screen. It is the opening move of a re-verification, covered on step-up re-authentication
403JSON, three shapes at the auth boundary (below)application/jsonThe caller authenticated but is not allowed to proceed: a capability gate, a cross origin write rejection, or a failed double-submit token check
404{ error: "<reason>" } or not foundmixedThe resource or route does not exist on this engine
409{ error: "<reason>" }application/jsonThe request conflicts with current state, for example a colliding config write or an update that already awaits verification. One attended-verification session may be live per account at a time. That refusal carries the existing session’s sessionId so a caller can resume it rather than start a second
429RFC 9457 problem+json: { type, title, status, error: "rate limited" }, plus a Retry-After headerapplication/problem+jsonA rate limit was met. Never an authentication failure: a 429 says to slow down, not to sign in again. A dual-control approve sent too soon after its request gets a plain JSON 429 instead, described under rate limiting
500RFC 9457 problem+json: { type, title, status, error: "internal error", requestId }application/problem+jsonAn unhandled fault inside the engine’s own dispatch, not a refusal of anything you sent
501{ error: "passkey_not_configured" } from the step-up routes, or { ok: false, reason: "passkey_not_configured" } from the sign-in ceremonyapplication/jsonA passkey or step-up ceremony was asked for on an engine with no CONSOLE_ORIGIN set, so there is no origin to bind the ceremony to. A configuration gap, not a credential problem, and retrying never clears it

A request whose method and path match no route at all returns a plain not found with status 404 (engine/src/admin/router.ts). The CORS allowlist is exact: the engine reflects Access-Control-Allow-Origin only for the configured console origin, so a call from any other origin gets no CORS headers back rather than a wildcard (corsHeaders, engine/src/index.ts). The one exception is the unauthenticated GET /admin/health probe, which answers with a wildcard origin and no credentials.

The 401 channel: two bodies that mean opposite things

A 401 is the one status on this surface that does not identify its own meaning. Read the body before you act on it.

The plaintext unauthorised is the authentication failure. It carries no JSON envelope. The engine returns it from the auth boundary, logging a coarse failure line that never carries a token, an assertion, or an email (logAuthFailure, engine/src/admin/router-core.ts).

The JSON { "error": "step-up required", "stepUpRequired": true } at the same status is not an authentication failure at all (requireStepUp, engine/src/admin/router-core.ts). The caller is authenticated, their session is live, and a step-up gated route is asking for a fresh proof of presence before it acts. The two are covered separately below, because a client that handles them the same way has a bug that only shows up on the sensitive routes.

StatusCodeWhenFix
401unauthorisedNo usable credential, or a presented credential did not verifyPresent a valid credential: a Cloudflare Access assertion, a signed session cookie, or the ADMIN_TOKEN bearer
401unauthorisedA higher-assurance method was present but invalid (a bad Access JWT, or a tampered or expired session cookie)Fix the failing credential. The engine never downgrades a present-but-invalid higher method to a weaker one, so a broken cookie does not silently fall through to the token
401unauthorisedAn Access assertion verified but carried no usable email or no stable subject (for example a Cloudflare Access service token)Use a credential that carries a verified email and subject. Only the bare ADMIN_TOKEN is exempt, because it is the email-less break-glass by design

The anti-downgrade rule is the property to note here. Precedence is Access, then the session cookie, then the token. A present-but-invalid step is rejected outright rather than retried at a lower tier (authorise, engine/src/admin/auth.ts). If you see a 401 with what you believe is a valid token, check that you did not also send an invalid Access header or a stale cookie, because the higher method is consulted first and its failure stops the request.

The JSON 401: step-up required

StatusCodeWhenFix
401step-up required (with stepUpRequired: true)A first-party cookie session called a step-up gated route without a fresh enough authentication and without a valid single-use step-up tokenRun the step-up ceremony at POST /admin/stepup/begin and POST /admin/stepup/finish, then re-send the original request carrying the returned token in the x-downpipes-stepup header. Do not sign the caller out

Two properties of this row matter more than the row. The first is that a bare ADMIN_TOKEN caller never sees this 401, because the bare token is exempt from step-up. A Cloudflare Access caller sees it only when its Access sign-in is more than five minutes old, so the body then also carries reauth: "access" and the fix is a fresh Access sign-in, not the passkey ceremony. The second is that the gate fails closed: if the engine cannot reach the component that answers the freshness question, it returns this same body rather than admitting the request. A client cannot distinguish those two from the wire, which is why the fleet-wide version of the symptom is diagnosed by blast radius rather than by the response, on step-up re-authentication.

The 403 channel: three different rejections

A 403 is always JSON, and at the auth boundary it has three shapes that mean different things. Discriminate them by the error field, because no two of them share a remedy. Two further 403s carry a named dual-control payload and are documented under named JSON payloads rather than here.

The capability gate

When a caller authenticated but their role does not hold the capability a route requires, the engine returns the Forbidden body: { "error": "forbidden", "required": "<capability>", "have": "<role>" } (gate in engine/src/admin/router-core.ts, and the Forbidden interface in engine/src/admin/identity.ts). The required field is the capability the route gates on, and have is the caller’s resolved role, so a client can say “your operator role cannot run a key ceremony” rather than showing a mystery failure.

StatusCodeWhenFix
403forbidden (with required and have)The caller’s role lacks the required capability for this routeGrant the caller a role that holds the named capability, or have someone who already holds it perform the action. The owner-exclusive capabilities (keys.ceremony, posture.riskaccept) cannot be conferred to any other role

For a custom-role caller the have field reports the viewer floor while the real authority is the caller’s resolved capability set, so read the failure as “this caller does not hold required”, not as a literal statement about the floor role (gate, engine/src/admin/router-core.ts).

The cross origin write rejection (CSRF)

A separate 403 guards cookie-borne sessions against a forged cross origin write. A mutating request (anything that is not GET or OPTIONS) authenticated by an ambient session cookie must also carry an Origin that exactly matches the configured console origin; if it does not, the engine returns { "error": "csrf origin check failed" } before reading the body or touching state (engine/src/admin/router.ts).

StatusCodeWhenFix
403csrf origin check failedA cookie-session mutating request arrived with a missing or mismatched OriginSend the request from the configured console origin. The Access method and the bare ADMIN_TOKEN bearer present an explicit header instead of an ambient cookie and are not subject to this check

The double-submit token rejection

The session-termination routes and the sign-in-factor revoke route carry a second, independent CSRF check on top of the Origin check above: a double-submit token. A cookie-borne request to one of those routes must present the x-downpipes-csrf header carrying the same value as its __Host-downpipes_csrf cookie; one that does not is refused with { "error": "csrf token check failed" } before any rate-limit or storage work (csrfBlock, engine/src/admin/router-account-session.ts). It is a different string from the origin rejection deliberately, because it has a different cause and a different fix.

StatusCodeWhenFix
403csrf token check failedA cookie-borne request to a session-termination route or to POST /admin/signin-factors/revoke presented no x-downpipes-csrf header, or one that did not match the paired cookieCall GET /admin/whoami on the cookie session first. It returns csrfToken in the body and sets the paired __Host-downpipes_csrf cookie in the same response, so echo the body value in the header on the request. A client that never called whoami, or a browser that blocked or partitioned the cookie, produces this without any cross-site attempt being involved

The way to tell the three apart is the error value, and a client must match all three rather than the first two. A forbidden body means change the caller’s role. A csrf origin check failed body means fix where the request is coming from. A csrf token check failed body means the request is coming from the right place but is not carrying the token that proves it.

Only a cookie-borne session can hit a CSRF 403. A passkey, native OIDC or native SAML session can hit both. A token session minted from the ADMIN_TOKEN bearer can hit the origin check only. An Access caller or a bare-bearer caller never hits either, because their explicit header is not a cross-site-ridable ambient credential.

Named JSON payloads worth knowing

Beyond the bare channels, a few routes return a named JSON body that an integrator should match on. Each entry below gives the body shape.

restore not approved on an apply

A restore apply (POST /admin/restore with confirm: true) needs the restore.apply capability. When an owner turns on restore dual control, the apply also needs a second identity’s approval. That approval binds to the same plan hash, and the maker cannot be the checker. The policy is off by default. If the capability passes but no usable approval exists, the engine returns a 403 with { "error": "restore not approved", "planHash": "<hash>" } and records a denied apply (engine/src/admin/router-restore.ts).

StatusCodeWhenFix
403restore not approved (with planHash)An apply was attempted with the right capability but without a second identity’s approval for that exact planHave a different authorised identity approve the request at POST /admin/restore/approve with this planHash, then re-submit the apply. A caller cannot approve their own request

The planHash in the body is the engine’s server-recomputed binding for the plan, so it is the value to pass to the approve route. This is a capability gate the console routes to “awaiting approval”, not an error to retry blindly. A changed plan produces a different hash and voids any prior approval, which is why dual control is covered in full on dual control.

A request raised against a plan that has aged out

POST /admin/restore/request binds the approval to a dry run, and the approval’s twenty-four hours are counted from that dry run rather than from the request. A request raised more than twenty-four hours and thirty minutes after the plan it names is refused with a 400 carrying { "error": "the plan this request is raised against is older than this engine will vouch for; run the dry run again and request from the plan it returns" } (engine/src/sched/scheduler-do-restore-approval.ts).

StatusCodeWhenFix
400the plan this request is raised against is older than this engine will vouch forA request was raised against a dry run whose window has closed, typically a plan left open long enough that what it predicted may no longer be what an apply would doRe-run the dry run for the same selection and raise the request against the plan it returns. Read the returned plan rather than assuming it matches the old one, because the counts and the fidelity warnings are recomputed

Match this separately from the 403 refusals above: it is not an authority problem and no amount of approving will clear it. The remedy is a fresh preview, which is also the disclosure the requester needs. The window and where it starts are covered on dual control.

prune not approved on a retention prune

Retention pruning has the same request, approve, apply shape as a restore, and a dual-control gate on every real apply. An apply with no usable approval bound to that exact plan is a 403 with { "error": "prune not approved", "downpipeId", "planHash", "mode": "not-approved", "retainedRuns", "supersededRuns" } (engine/src/admin/router-retention-prune.ts).

StatusCodeWhenFix
403prune not approved (with planHash and mode: "not-approved")A prune apply reached the dual-control gate with no approval armed for that planHave a different authorised identity approve the plan, then re-submit. Match on this string separately from restore not approved: they are two different routes and a client that matches only the restore one falls through to a generic denial, which reads as a role problem and is not one

The two run counts in the body are the reason to parse this rather than treat it as a bare refusal. They tell the caller what the refused plan would have done, so an approver can be shown the consequence before they approve rather than after.

unknown report kind, unknown framework and unknown channel

A report request for a kind the engine does not recognise is a 404 with { "error": "unknown report kind" }. An evidence pack asked for a ?framework= the engine does not map is a 404 with { "error": "unknown framework" } (both handleReport, engine/src/admin/router-posture.ts). The notification test-send route returns { "error": "unknown channel" } with the same 404 status for a channelId the engine does not hold (engine/src/admin/router-ops.ts).

StatusCodeWhenFix
404unknown report kindGET /admin/reports/:kind named a report kind the engine does not produceRequest one of the supported kinds. The kind is validated before any data is gathered, so a typo never produces a partial report
404unknown frameworkGET /admin/reports/evidence-pack?framework= named a framework the engine does not mapRequest a framework the engine maps, or omit the parameter for the all default. Like the kind above this is checked before anything is gathered, so a typo never silently yields an empty pack
404unknown channelPOST /admin/notify/test named a channelId the engine does not holdUse the id of a configured channel

report could not be rendered

One 500 on this surface is route-specific rather than the generic dispatch fault below. This 500 is worth knowing because of what it does not mean. A report request with ?format=pdf whose data was gathered and signed successfully but whose PDF render threw returns { "error": "report could not be rendered" } (handleReport, engine/src/admin/router-posture.ts).

StatusCodeWhenFix
500report could not be renderedThe PDF renderer threw on a report that was otherwise assembled and signedRequest the same report without ?format=pdf. The JSON body is unaffected, because the failure is in the rendering step alone, so an auditor waiting on the report is not blocked while the render fault is investigated

Do not read this as an unavailable report. The distinction is the whole value of the separate string: the generic internal error says the engine could not answer, and this one says it answered and could not draw the answer.

The 202 owner-action-queued response

Several owner operations are dual-control gated. When the gate is on and no armed approval exists yet, the route records a pending approval and returns 202 with { "ownerActionQueued": true, "id": "<id>", "status": "pending" } (ownerActionQueuedResponse, engine/src/admin/router-core.ts). This is the single most important non-failure to handle correctly.

StatusCodeWhenFix
202ownerActionQueued (with id)A gated owner action was submitted with dual control on and no second-owner approval armedTreat it as queued, not failed. A second owner approves the recorded action, then the original owner re-submits the same request, and the gate consumes the approval and proceeds

Do not retry a 202 as though it errored. The action has been recorded under id and is waiting for a second owner; the privileged step (for example minting a one-shot token) runs only on the approved re-submit, never on this first call (ownerActionGate, engine/src/admin/router-core.ts).

Rate limiting and the 429

Every rate-limit 429 the engine returns carries the same body: an RFC 9457 problem+json document (application/problem+json) with type, title, status and an error field that a client matches on (error === "rate limited", or "rate limit unavailable" from a fail-closed limiter), plus a Retry-After header in whole seconds, rounded up so a client never retries a hair early into a still-saturated window, and the advisory IETF RateLimit and RateLimit-Policy headers beside it. One builder produces all of them (rateLimitedResponse, engine/src/admin/router-core.ts), so the shape does not vary by which limiter refused. A dual-control approve that arrives within five seconds of its request gets a different 429. Its application/json body is { error, refusal: "too-soon", retryAfterSeconds }, with a Retry-After header (engine/src/sched/scheduler-do.ts).

There is more than one limiter. They differ in what they are keyed on and how they behave when their own backing store is unavailable.

StatusCodeWhenFix
429rate limited (with Retry-After)The caller exceeded the per-caller cap for the fixed window. This is the limiter on the mutating routes: every admin write is a POST, and the GET reads are exempt from this oneBack off for the Retry-After seconds, then retry. The bucket is per verified caller, not per source IP, so a shared egress does not pool everyone into one limit
429rate limited (with Retry-After)The source address exceeded the per-IP cap on the bare ADMIN_TOKEN break-glass compare: at most ten attempts a minute from one address. This limiter sits inside the auth gate, so unlike the per-caller one it applies to GET as well, and it fails closed, denying the compare outright if its own backing store cannot be reachedBack off for the Retry-After seconds. A correct token does not exempt you, because the throttle is checked before the token is compared at all
429rate limit unavailable (with Retry-After)The per-IP limiter on the unauthenticated sign-in ceremony could not reach its own backing store, and that surface fails closedRetry after the advertised window. This is an engine-side availability fault, not a statement about your credential

Two properties matter for an integrator. First, the per-caller limiter is keyed on the verified caller (the caller’s stable subject, or one shared bucket for the bare token) rather than the source IP, so a shared NAT does not lock out legitimate callers; the two per-IP limiters above are keyed the other way, in their own separate namespaces, because they guard credential-guessing surfaces where the caller is not yet established. Second, the per-caller limiter fails open: if its backing store is unavailable the request is admitted rather than denied, because this is an authenticated admin and recovery surface where availability beats strict limiting (rateLimited, engine/src/admin/router-core.ts). The two per-IP limiters fail closed for the opposite reason: a brute-force gate that opens when its counter breaks is not a gate. The windows and caps are constants in the scheduler Durable Object (RATE_LIMIT_WINDOW_MS, RATE_LIMIT_MAX_PER_WINDOW and the per-IP caps beside them, engine/src/sched/scheduler-do-limits.ts); read them there rather than hard-coding a number, since they are an operational tuning, not a wire contract.

A 429 is never a session loss. The break-glass throttle in particular refuses before any credential is compared, so it establishes nothing about the token presented. A client that treats a rate limit as an expired session and signs the operator out is reading it wrongly, which is why the engine answers 429 for it rather than the 401 every genuine authentication failure returns.

Ordering of the gate, the limiter and dual control on an apply

On a restore apply the engine runs the checks in a deliberate order so a security-relevant denial is never masked by a rate limit (engine/src/admin/router.ts). The capability gate runs first, so an unauthorised apply always returns the forbidden 403 and a denied-apply audit entry, even if the caller is also over their rate limit. The per-caller limiter runs next, so an authorised caller hammering the apply surface still meets a 429. The dual-control approval check runs last, before any byte is written, and a missing approval is the restore not approved 403. The approval is consumed only after a successful apply, so a failed apply leaves the approval usable for a retry without a fresh round of dual control.

The 500 channel: an unhandled fault

A 500 is the engine’s own last-resort handler, not a route-specific refusal. Both the admin dispatch and the outer Worker entry run under a single top-level catch, so a route handler that throws unexpectedly lands here rather than escaping the response unlogged (engine/src/index.ts). The body is the same RFC 9457 problem+json shape the 429s use, with error: "internal error" and a requestId that also rides on the response as the x-downpipe-request-id header, so a client and the engine’s own log line can be matched by the one value.

StatusCodeWhenFix
500internal error (with requestId)An unhandled error inside the engine’s dispatch, not caused by anything identifiable in the request itselfRetry once. If it persists, quote the requestId (also the x-downpipe-request-id response header) when raising it, since it is the join key to the engine’s own log line

What this reference does not cover

Two things are deliberately out of scope here. The first is the reader’s process exit codes: the offline downpipe tool returns numeric exit statuses (0 verified, 2 unverified, and so on) that are a different contract from these HTTP codes, documented on the CLI exit codes. The second is the DELETE, PUT and PATCH verbs: the admin API mutates exclusively through POST and reads through GET, so those verbs are not part of the surface and are not documented as error rows.

Where this fits

These are the wire-level outcomes; the pages below give the surrounding model.

Last updated .