Skip to content
downpipes docs

Snapshot consistency: live crawls versus a point-in-time view, and what a downpipes backup actually represents

There is no whole-account point-in-time snapshot. A downpipes backup of KV or R2 is a live crawl of a store that may be changing while the crawl runs, so a single backup can mix keys that were written at slightly different instants. This page explains that, and names the one source that does give a coherent cross-table snapshot and under what condition. It defines what “skipped” means so you do not read it as a failure. It is written for a developer who needs to know what guarantee a backup carries before relying on it.

A backup is consistent per record, not consistent across the whole account at a single instant. For most KV and R2 workloads that is exactly what you want and exactly what a periodic backup can promise. Where you need cross-record consistency, D1 gives it, and the practical reading at the end of this page tells you how to lean on that.

KV is a live crawl

A KV backup lists the namespace and reads each in-scope key’s value. KV has no point-in-time snapshot to read against, so this is a crawl over a live namespace. A key deleted between appearing in a list page and the crawl reaching it reads back as gone. The crawl seals a small _vanished marker record in its place, and a restore skips that marker. That is the designed behaviour, not an error.

The crawl is ordered. KV list results are lexicographic, which is what lets a large namespace be backed up across more than one engine invocation: each fully-yielded page records a resume watermark (its last key), and a resumed crawl continues after that key without re-reading any value. So while a KV backup is not a point-in-time view, it is a single ordered pass in which each key is read once.

R2 is a live ordered crawl too

R2 behaves the same way, with one extra safeguard on its large objects. The bucket is listed in lexicographic order and each in-scope object is read. A streamed object over 8 MiB is read twice (once to address it, once to seal it), so each pass of that streamed read is pinned to the etag observed at list time, because without the pin an object replaced between those two reads could seal bytes that do not match the recorded address. With the pin, a replaced object reads as gone, so the writer skips and counts the record rather than archiving something that can never verify. A smaller object is read with a single plain get, so it is one read that cannot tear and needs no pin.

Object size decides how an object is read. An object over 8 MiB is streamed: it is read in passes and fed into the seal in windows, never held whole in memory. A smaller object is read whole, which is simpler and deduplicates the same way. Neither path buffers a large object.

Skipped is expected, not data loss

When a key or object vanishes or changes mid-crawl, the run does not fail. A deleted key or small object leaves a _vanished marker record, which a restore skips. A large object replaced or deleted mid-read is excluded from the archive entirely. This is the correct outcome for a live store: the alternative would be archiving a value that no longer exists, or one whose bytes do not match its recorded address. A skip is a normal outcome of the run, not a failure.

D1 is the one consistent point-in-time view

D1 is the exception, and it is a real one.

Record structure

A D1 database is backed up as a resumable sequence of records, not one whole-database record: one header record carrying every table’s CREATE statement and column names with no rows, then many per-table keyset row-page records (bounded by a per-page byte target of 8 MiB, a hard cap of 96 MiB, and up to 2,000 rows per page), then one schema record carrying the CREATE INDEX, TRIGGER and VIEW statements, applied last.

Session-bookmark mechanism

The whole sequence reads through one D1 Session anchored with withSession("first-primary"). The first query goes to the primary, which holds the newest committed state, and establishes a bookmark; every later read on that session, across every record in the sequence and across every resumed slice, is constrained to a replica at or after that bookmark. The header, every row page and the schema record therefore see one sequentially-consistent point-in-time snapshot. A database under concurrent write load is exported torn-free, rather than mixing pre-write and post-write rows across tables.

Each record is produced incrementally, not buffered as one dump. The plan (every table’s DDL and columns, plus the index/trigger/view text) is read once up front, on the first query, which is also what establishes the session bookmark described above. The header record is built from that plan and emitted first; each table’s rows are then read in keyset-ordered pages (paging by rowid where the table has one, or by PRIMARY KEY for a WITHOUT ROWID table from engine 0.3.6) and emitted one page at a time; the schema record is built from the same plan but emitted last, after every row page, so a restore can create the tables, load the rows, then add the indexes and triggers. Each record is one buffered value of bounded size, so peak memory is about one row page and a multi-gigabyte database is backed up rather than rejected.

The D1 snapshot is conditional

The D1 snapshot’s cross-table consistency is not unconditional in every environment: it holds on the production D1 binding, which supports withSession. On an older or local binding that does not expose withSession, the export degrades gracefully to a best-effort read against the bare database: the reads are still ordered keyset pages, but the single cross-table snapshot guarantee is then best-effort rather than sequentially consistent. So the guarantee is “sequentially consistent on the production binding, best-effort without withSession”, never “always a perfect snapshot everywhere”.

What is buffered, and what is not

Streaming is the rule, and the exception is narrow.

CaptureHeld whole in memory?
A large R2 object (over 8 MiB)No, streamed in windows
A small R2 objectYes, read whole (it is small)
A D1 database (header, row-page and schema records)No, produced and sealed one record at a time; each record is buffered, and a row page is capped at 96 MiB
A Worker script bodyYes, buffered, capped at 128 MiB (from engine 0.3.6 the engine refuses a larger body before it holds all of it)

The Worker script body is the deliberate buffered exception. The engine marks a script past the cap unavailable and does not archive it. From engine 0.3.6 the engine refuses such a body before it holds all of the body in memory. If the declared length is over the cap, the engine reads none of the body. If the response declares no length, the engine stops as soon as the bytes pass the cap. It then holds no more than the cap and the chunk it is reading.

Where you see skipped counts

The Runs screen does not treat a mid-crawl skip as a clean success. A run with vanished or skipped records shows a warning badge that reads “N not fully captured”, and its detail names each kind. The run itself does not fail. Run now on the downpipes screen uses “skipped” in another sense: a trigger that arrives while a run is in progress is skipped, and an information toast says so.

How this connects to recovery

Because there is no within-run point-in-time view for KV and R2, recovery is at run granularity, not at an arbitrary instant inside a run. When you choose a moment to restore to, downpipes resolves it to the latest successful run that completed at or before that moment. That run is read from a bounded per-downpipe history ring. If the moment picked is older than the oldest run the ring still holds, the engine reports a miss rather than returning a run. There is no arbitrary-time picker, because there is no continuous timeline to pick from: the recovery points are your retained runs. See the recovery overview and restore flow for how a run is chosen and replayed.

A short practical reading

Two habits follow from the difference between a live crawl and a sequentially-consistent snapshot. Where you need several records to be consistent with each other, put that data in D1 and lean on its sequentially-consistent snapshot. That is the strongest cross-record guarantee downpipes offers. For KV and R2, treat each backup as a consistent-per-key ordered crawl rather than a global instant. Lean on cadence and on failover copies across destinations. A key caught mid-change in one run is then captured cleanly in the next.

Where this fits

Last updated .