Skip to content
downpipes docs

Selecting what a downpipe covers: prefix selectors, source-native scope, and one D1 source per database

A selector is the rule a downpipe uses to decide which records it captures from a source. It is a pair of prefix lists, one to include and one to exclude, matched against the source-native name of each record. This page sets out how that match works, which source types it applies to, and the one rule that surprises people: the selector chooses which databases a downpipe covers, never which rows inside a single D1 database. It is written for a self-hoster who wants to predict what a backup will and will not contain.

The selector is deliberately small, so you can reason about it. For six of the eight source types the matching rule is a prefix match, and exclude always beats include. There are two exceptions, not one.

Cloudflare config matches by exact surface id rather than by prefix. D1 does not run the prefix rule at all: its crawl takes the selector and never consults it, so the selector picks which databases a downpipe covers and nothing inside one. Both are covered below. Everything else on this page is how that one rule meets each source type and how the console expresses it.

The one matching rule

A record is in scope when its name matches at least one include prefix and matches no exclude prefix. The engine implements this in a single pure function, inScope, and KV, R2, Secrets, Stream, Images and Workers all crawl through that one function. That is the six. The engine also has an Artifacts source, which the console does not offer and no configuration can select (ARTIFACTS_GA, console/src/lib/token-source.ts); it crawls through inScope as well.

Cloudflare config is the first exception: a surface is matched by exact surface id, via a separate function called surfaceSelected, not by the inScope prefix rule the other source types share. D1 is the second: its crawl signature takes the selector and discards it, so no prefix test runs anywhere in a D1 capture. The shape is otherwise the same, exclude wins and an empty include means everything, it is just comparing whole ids rather than testing a prefix.

Two edge cases fall straight out of that rule. An empty include list means everything is included, because the rule treats “no include prefixes” as “do not narrow by include at all”. And exclude always wins, because a name that matches an exclude prefix is rejected even when it also matched an include prefix. So the way to read a selector is: start from the include set (or from everything, if include is empty), then carve out anything the exclude set matches.

The match is a prefix test, name.startsWith(prefix), not a glob and not a regular expression. A prefix of uploads/ matches uploads/a.jpg and uploads/2026/x but not avatars/uploads/a.jpg, because the name must start with the prefix. There are no wildcards to learn.

The comparison is case-sensitive

The prefix test is a byte comparison, so case matters and nothing is normalised. An include of Users/ does not match users/profile.json, and an exclude of TMP/ does not exclude tmp/scratch.

This is the selector mistake with the quietest failure. A downpipe whose include prefixes match nothing is not an error: it is a downpipe with an empty scope, and it runs, succeeds and writes an archive containing nothing from that source. Nothing on the screen is red, because nothing went wrong; the selector simply admitted no records. The same typo in an EXCLUDE prefix fails the other way and just as quietly, admitting the records you meant to leave out.

So check the case against a real key rather than against what you remember the naming convention to be. The run’s own record count is the check that catches it: a source you expected to contribute thousands of records contributing none is the symptom, and it is visible on the first run rather than at recovery.

Whitespace, by contrast, is not significant. The console splits your list on commas, trims each entry and drops the empties, so uploads/, avatars/ and uploads/,avatars/ are the same two prefixes. Only the characters inside a prefix count, and their case.

An empty prefix is refused

An entry that is the empty string is not a prefix, and the engine rejects any request that carries one. Every name starts with the empty string, so a single "" in exclude matches every record and puts the whole source out of scope. A downpipe saved that way would capture nothing and report a successful run; a restore scoped that way would write nothing and report a successful restore. That is the worst shape a selector mistake can take, because the failure you would act on never appears, so the empty entry is refused outright rather than quietly dropped from the list (selectorPrefixFault, engine/src/sources/selector.ts). A caller who sent one meant something by it, and what they are short of is the knowledge that it covers everything.

The refusal sits at every boundary that takes a selector from a caller:

WhereWhat happens
Saving a downpipe (POST /admin/downpipes)rejected with 400, and the downpipe is not stored
A restore dry run or apply (POST /admin/restore)rejected with 400 before any plan is built
Raising a restore request (POST /admin/restore/request)rejected with 400 before the request reaches an approver
A restorability proof (POST /admin/restore/verify)rejected with 400 before the proof runs

Each returns the reason as { "error": "..." }. On the POST /restore and POST /restore/request routes the engine runs the check before it consults your role, because the check is a statement about the request rather than about you, so the answer is the same whichever role you hold. On the POST /restore/verify route the engine checks your role first, then the selector (engine/src/admin/router-restore.ts). The engine grants restore.verify to viewer and up (engine/src/admin/identity-rbac.ts), so in practice the selector answer still reaches a signed-in caller. The exclude message names the consequence rather than only the fault:

exclude contains an empty prefix (""), which matches every record name and would put the
whole source out of scope, so nothing would be backed up or restored; remove it, or give
the prefix you meant

An empty entry in include is refused by the same rule, with its own message, even though an empty include list already means everything: one rule that covers both fields is easier to hold than two that differ on which field forgives it. An entry that is not a string is refused as well, in either field, because startsWith coerces its argument and a number in the list would silently become a prefix nobody wrote.

You will not meet this refusal from the console, because the empties are dropped from your typed list before the request is built, as above. It is reachable over the API, which is the path your own automation and any machine-to-machine integration take. The API is where the rule has to hold.

What the names are, per source type

The prefixes are matched against the source-native name of each record, and that name is different for each source type.

Source typeThe name a prefix matchesExample include prefix
KVthe KV key namesession:
R2the R2 object keyuploads/2026/
Secretsthe secret namedb-
Streamthe video uid(the console has no field for Stream; every video is captured)
Imagesthe image id (and the _variants record)(the console has no field for Images; every image is captured)
Workersthe script-id-prefixed record name(the console has no field for Workers; every script is captured)
Cloudflare configthe surface id, matched exactly, not by prefix(picked from a catalogue)

For KV and R2 the prefix is the obvious thing: the key. A KV downpipe with an include of session: and no exclude captures every key whose name begins with session:. An R2 downpipe with an include of uploads/ and an exclude of uploads/tmp/ captures everything under uploads/ except the temporary tree, because exclude wins over the broader include.

Stream and Images are account-scoped inventories, and the engine’s inScope selector can in principle narrow each by the video uid or the image id respectively. But the console hides the include and exclude prefix fields for both, the same way it does for Cloudflare config and Workers, and always sends an empty include and exclude for a Stream or Images source regardless of what is typed: a Stream or Images downpipe always captures every video or image in the account. Narrowing either source is an engine capability with no console control.

Workers records are named by script id with a per-aspect suffix (<id>, <id>/settings, <id>/versions, <id>/schedules). The engine’s inScope selector can in principle scope the Workers crawl by that name too. But the console shows no include or exclude control for Workers at all either: the standard prefix fields are hidden for a Workers source, the same as they are for Cloudflare config, Stream and Images, so a Workers downpipe always captures every script in the account. Cloudflare config does not have key prefixes at all; you choose which configuration surfaces a downpipe covers from a catalogue in the console, and that selection is the scope. The Cloudflare config surface reference lists the surfaces.

For a Secrets source the prefix matches on the secret name, and unlike D1 it actually takes effect: a Secrets downpipe with an include of db- captures only the secrets whose name begins with db-, so the selector subsets which secrets are captured. The secret name is wiring, not the value; nothing about the prefix reads or reveals a secret value. This is the opposite of D1, where the console shows the prefix fields but they are inert (see below).

The D1 rule: one source, one whole database

This is the rule worth committing to memory. The selector does not apply inside a single D1 database. It does not subset rows, and it does not subset tables. A D1 source captures the entire database, every table the export reads.

It captures it as a sequence of records, not as one. The export emits one header record (every table’s CREATE plus its ordered column names), then many row-page records, then one schema record carrying the indexes, triggers and views (crawlFrom, engine/src/sources/d1.ts). That sequence is what makes a multi-gigabyte database backable at all, because no single record ever holds more than one keyset page of rows, and it is what lets a large export resume between records. Size a D1 backup from that record sequence; D1 in depth has the per-record byte guards.

From engine 0.3.6, a WITHOUT ROWID table pages by its PRIMARY KEY, and its resume point holds the last key of the page. When that key is more than 4 KiB encoded, the page gets no resume point, unless it is the last page of the table.

So what does the selector do for D1? It chooses which databases a downpipe covers, expressed by configuring one D1 source per database. If you want to back up three databases, you add three D1 sources, one each. There is no way to write a selector that captures “only these rows” or “only these tables” of one database, because a D1 backup is a single sequentially-consistent dump of the whole database, and a partial dump would not be a coherent snapshot. The D1 in depth reference covers how that whole-database dump is produced and replayed.

A practical consequence for D1

To narrow what you protect in D1, split the data across databases and protect the ones you want, rather than reaching for an include or exclude prefix. The downpipe editor does show the prefix fields when D1 is selected, but they have no effect inside a D1 database; for D1 they are inert. Treat the database, not the row, as the unit of scope.

Two worked examples

A concrete pair makes the include-and-exclude interaction clear. Read each row as: given these prefixes, is the named record captured?

SelectorRecord nameIn scope?Why
include [], exclude []uploads/a.jpgYesempty include means everything, nothing excluded
include ["uploads/"], exclude ["uploads/tmp/"]uploads/2026/a.jpgYesmatches the include, matches no exclude
include ["uploads/"], exclude ["uploads/tmp/"]uploads/tmp/scratchNoexclude wins over the broader include
include ["uploads/"], exclude []avatars/x.pngNomatches no include prefix

The second and third rows are the important pair: the same selector keeps a permanent object and drops a scratch object, purely because the exclude prefix carves the scratch tree out of the broader include.

How the console expresses it

The selector is set in the downpipe editor, under the “Advanced: prefixes, cost, restore test, overrides” disclosure rather than on the main form, so the calm default is “back up everything in this source”. The disclosure offers a comma-separated include field (“Include prefixes”, with the hint that empty means all keys) and a comma-separated exclude field (“Exclude prefixes”, with the hint that exclude wins over include). Those two fields are exactly the include and exclude lists the engine matches with. For Cloudflare config the standard prefix fields are hidden; you pick which surfaces a downpipe covers from a catalogue, and that ticked selection becomes the engine’s include list. For Workers, Stream and Images the prefix fields are hidden too, and the console shows only an informational note about what the source backs up. So a Workers, Stream or Images downpipe always captures every script, video or image in the account; treat the account, not a script, video or image name, as the unit of scope for these three, because there is no per-item scoping control in the console.

The guided create wizard has no include or exclude prefix fields. Its “Use the advanced editor” link opens the downpipe editor.

Bulk protect is the fast path for many sources at once. From the Sources screen you tick the attached-but-not-yet-protected sources and choose one shared schedule. The console then creates one downpipe per KV, R2 or D1 binding on that shared cadence. The ticked secrets go into one Secrets downpipe. Each of those downpipes is created with an empty include and an empty exclude, so a bulk-protected source captures everything by default; you can narrow any one of them afterwards by editing its selector. So bulk protect is “protect all of these wholesale on one schedule”, and per-source narrowing is a deliberate follow-up.

friendlyName, a presentation detail that does not change scope

A small thing that often comes up: the console strips the SRC_ prefixes from a binding name when it shows a downpipe, so a binding named SRC_D1_app_db reads as “app db” in the interface. This is friendlyName, and it is presentation only. The binding name on the engine is unchanged, the selector still matches against the real source-native record names, and nothing about scope shifts. It exists so a fleet of scoped downpipes reads as human names rather than as code.

How scope interacts with the cost estimate

Before you commit a downpipe, the console can project its cost, and the estimate respects the selector. The estimate counts only the in-scope records, so a tight selector previews a smaller backup. The estimate differs between source types in two ways.

A KV estimate counts the in-scope keys from list metadata only and never reads a value, so it reports a record count but reports the byte total as unknown. KV list results do not carry value sizes, so the console cannot know how many bytes a KV backup will move without reading every value, which the estimate deliberately does not do. An R2 estimate is different: R2 list results do carry object sizes, so an R2 estimate returns a real byte total alongside the record count. So when you read a cost preview, expect a real size for R2 and “unknown” for the KV byte total. Both are scoped to exactly what the selector admits.

Why the estimate never reads values to size them

The headline cost risk for a large KV namespace is a high-frequency full re-read, where every value read is a billable operation. If the estimate itself read each value to measure it, the act of previewing the cost would incur the very cost it is trying to predict. So the estimate is a projection from list metadata: it lists keys (cheap) and counts the in-scope ones, and reports bytes as unknown for KV because the list does not carry sizes. R2 list does carry sizes, so its estimate can be exact without reading a single object body. This is the same reason a backup of a large namespace is a careful, scoped, scheduled thing rather than a default-everything-often habit.

Where this fits

  • Connect a source is the task page for attaching and protecting a source from the console.
  • Sources overview sets out what each source type captures and how each restores.
  • D1 in depth covers the whole-database dump that is the reason the selector cannot subset a D1 database.
  • Snapshot consistency explains what a backup of a live KV or R2 store actually represents.
  • Cost prediction goes deeper on the estimate and the KV read-cost projection the wizard shows.

Last updated .