# ADR 022: The trash never expires; history is the audit log; "restore this version" is a walk of undos

> item 10a's asset GC).

Source: https://nooklet.danielalder.cz/decisions/022-trash-history-retention

Date: 2026-09-12. Status: accepted. Implements research/13 §4.2 item 8 (and the retention half of
item 10a's asset GC).

## Context

Every write is already an op in the log, every deletion is already a tombstone (`deleted_at`,
ADR 003), and every write already records a full before/after image per entity in `changes`
(ADR 013). So "trash" and "page history" were, in the data, already there — what was missing was
a way to read them and a way to act on them, and an honest statement of what they promise.

Three of the v1 tool descriptions promised deletions were "restorable for 30 days". Nothing
implemented that: `nooklet gc` (M6) trims the `op` table and never touches a `page`/`block` row,
so tombstones were kept forever and nothing listed them. The review of 2026-09-12 struck the
wording (B-56). This ADR decides what the promise actually is.

## Decisions

### 1. Retention: none. The trash is kept indefinitely.

`page`/`block` rows with a tombstone are never hard-deleted. `trash.list` shows all of them;
`trash.restore` brings any of them back. The asset GC (below) treats a reference from a
tombstoned block as a live reference for the same reason.

**Rejected: purge tombstones after N days.** It costs more than it saves, in four places:

- `nooklet verify` replays the op log and diffs it against live state (ADR 003's core guarantee).
  A purged row reappears on replay, so every purge is a permanent, growing "expected divergence"
  — unless the purge is gated behind the op-log GC floor, at which point it is a second GC with
  its own floor logic for a table that is a rounding error in size (tens of KB of tombstones next
  to an 85 MB vector table, sql-schema.md rule 28).
- A device that was offline through the purge can still push a `block.text` for the purged block.
  Core's `applyOps` treats a missing row as "no such block" (noop), so the edit is silently lost,
  and a `block.place` of a live child under a purged parent is a rejected move. Both are new
  failure modes for a feature whose whole point is trust.
- `((block-ref))` targets, `changes.entity_id`, `path_ref` and the FTS shadow tables all key on
  the id. A purge is a cascade, and each table has its own trigger semantics.
- The saving is disk that nobody will miss.

**Not done, recorded:** a "purge now" button on the trash view, gated behind the GC floor, is the
shape a future purge would take if someone actually wants one. Nothing in this design prevents it.

### 2. What one trash entry is, and what a restore brings back

The grouping key is the **tombstone instant**. Every deleter stamps one `now` across the whole
action — the editor's `deleteSelectedBlocks`, the API's `block.delete` (the subtree) and
`page.delete` (the page and its blocks). So:

- A **page** entry is a tombstoned page. Restoring it un-deletes the page and every block on it
  whose `deleted_at` equals the page's; blocks that were never tombstoned (hidden only because
  the page was) simply reappear. A block deleted separately, earlier, keeps its own tombstone —
  it is not offered while its page is in the trash (it could not be restored on its own), and it
  becomes its own entry once the page is back.
- A **block** entry is a tombstoned block on a live page whose parent is live or was tombstoned
  at a different instant — the root of a delete action. Restoring it un-deletes the descendants
  that share its instant, plus any tombstoned ancestor, so the block is actually visible rather
  than live-but-hidden. Only the ancestor chain is revived, not the ancestors' other deleted
  children; those stay as their own entries.
- A page whose name has since been taken by a live page is refused with `conflict` unless
  `new_name` is given. Core's `applyPageDelete(null)` does not re-check the live-name unique
  index (B-90), so without the guard the write would fail on a SQL constraint mid-transaction.

**Known limitation:** two separate delete actions in the same millisecond share an instant and
merge into one entry; restoring one restores both. Real actions are never that close; in-process
tests space theirs out.

**Rejected: the deleting batch as the grouping key.** More precise on paper, but a sync push is
one batch per HTTP request (200 ops), so a large subtree delete from a device spans batches and
would split into several entries; the timestamp does not.

**Rejected: a raw `UPDATE … SET deleted_at = NULL`.** It would be one line and wrong twice, for
the reasons ADR 018 gives for its migration: other devices never hear of it, and `verify` flags
it. The restore mints `page.delete`/`block.delete {deletedAt: null}` ops through
`serverApplyOps`, exactly as `batch.undo` does — so it replicates, it is audited, and it is
itself a batch that `batch.undo` can reverse.

### 3. History is the audit log, grouped by batch

`page.history` reads the `changes` rows for the page entity and for every block currently on the
page, grouped by `batch_id` (one batch = one write call, one sync push, one undo), newest first,
each with a summary in words and the per-entity before/after images the client diffs. A batch
whose ops were all LWW-stale or identical re-sends changed nothing visible and is hidden.

"This page's" rows are attributed by the block's *current* page: a block moved here brings its
earlier history along, one moved away takes it. Same choice as `changes.since`'s `page` filter;
the alternative (parsing every row's `place.pageId`) is a full scan for a distinction nobody
asked for.

### 4. "Restore this version" is `batch.undo` of every newer batch, newest first — from the client

The op log has no page snapshots to jump to. What it has is each batch's before-image, and
`batch.undo` already turns one batch into a complete, audited, LWW-winning compensating write.
Restoring the page as it was after batch K is therefore: undo every batch newer than K on this
page, newest first. The History view does exactly that after a confirm that says how many.

What this honestly is and is not:

- **Not atomic across batches.** If step 3 of 5 fails, the page is left after step 2, and the
  view says so. Each step is atomic on its own.
- **Not page-scoped.** A batch that also touched another page (a cross-page move, a `batch` op,
  a graph-wide replace) is undone there too. The confirm says so.
- **Recorded.** Each undo is its own batch, so the walk shows up in the timeline and can be
  undone in turn — there is no separate redo.

**Rejected: a server-side `page.restore_version` op.** It would call `batch.undo`'s logic in a
loop inside one savepoint, which buys all-or-nothing across the walk — but the cross-page
semantics are identical, the audit trail is identical (still one batch per undo), and it means
either re-implementing or re-exporting `batch.undo`'s compensation. Worth doing when someone
actually hits the partial-failure case; recorded here so the shape is known.

**Rejected: page snapshots.** A per-write full-page image would make "jump to version K" one
op, at the cost of storing every page's whole text on every keystroke batch. The per-entity
images already stored are the right granularity for a tool whose unit of change is the block.

### 5. Asset orphans and the trash

An asset is an orphan only when no block text or property value — **live or tombstoned** —
mentions `assets/<id>.<ext>`, and the upload is older than a grace period (7 days by default,
`--asset-grace`). The trash-only rule follows from decision 1: a page restored from the trash
must not come back with broken images. The grace period exists because the one legitimate
unreferenced window is between an upload and the block write that embeds it, which can sit in an
offline device's push queue for as long as a laptop stays closed; a week covers a holiday, and a
monthly `nooklet gc` still collects. A `changes` row for the asset newer than the cutoff extends
the grace, so a future "touched on deduplicated re-upload" audit row (B-91) needs no GC change.

**Amended 2026-09-13 (M7 server/sync review, F10): page history counts as a reference too.** An
asset mentioned in a page or block pre/post-image in `changes` is kept, reported as
`keptByHistoryOnly`. Decision 4 makes "restore this version" a walk of `batch.undo`, which
rewrites block text from `changes.before_json`; an image removed by an *edit* (nothing in the
trash) and collected a week later came back from a history restore as a broken image, recoverable
only by digging the file out of the pre-GC backup. The reasoning is decision 1's: the saving is
disk nobody will miss, the loss is a restore people trust.

What it costs, stated plainly: `changes` is never trimmed, so an asset any recorded write ever
embedded is never collected. The GC still collects uploads no write ever pointed at — an agent's
`asset_upload` that was never used, a paste whose block write never reached the server — which is
the window the grace period was built for; it no longer collects "a picture pasted and later
edited out of its block".

Rejected: *extend the grace from the last history mention* (an image edited out within the last 7
days is kept, older ones collected) — keeps GC useful for old edits, but a restore of a version
from last month still shows a broken image, the failure this amendment exists to remove.
*Document the limitation and keep collecting* — honest, and the smallest change, but it leaves a
first-class feature (History's restore) silently lossy. If disk ever matters, the shape of a real
answer is trimming history itself behind an explicit horizon, at which point the assets it alone
kept become collectable with no GC change.

## Consequences

- `docs/spec/mcp-tools.md` gains `trash.list`, `trash.restore`, `page.history` (rows 25–27); the
  earlier note that `trash.*` would be HTTP-only is superseded — an agent that deletes the wrong
  thing needs the trash as much as a person does.
- Tool descriptions say "the trash has no expiry", never a number of days.
- `/trash` and `/history/<name>` are routes; the history route is not nested under `/page/…`
  because that route is a splat and would read `/history` as part of the page name.
- The web client's `data/history.ts` carries its own copy of `store.ts`'s `stamped` invalidation
  idiom; `store.ts` keeps its signals private and is owned by another workstream this milestone.
- Open follow-ups, logged: B-90 (core should reject a colliding un-delete the way it rejects a
  colliding rename), B-91 (record a deduplicated re-upload so the asset GC sees it as recent).

## Amendment (2026-09-13): the walk keeps later edits (B-251)

Decision 4 said the walk is "not page-scoped" and left it there. Exploratory QA on the real graph
showed what that costs: restoring one journal day walked back through a graph-wide replace and its
undo, and `batch.undo`'s last-writer-wins before-images overwrote every later edit to the 835
blocks the replace had touched — a page merge's 19 link rewrites, a Turn-into-page link — with
nothing on screen saying so. The confirm's one sentence about other pages did not describe that.

`batch.undo` now takes `keep_later_edits` and `ignore_batches`. With the flag, a field another
batch changed after the one being undone is left as it is and reported in `kept`; the walk passes
its own batches and the undo batches it has written so far as `ignore_batches`, so its earlier
steps are not mistaken for later edits. The History view passes both for Undo and for Restore and
names the pages where something was kept. The walk is still not page-scoped — a replace is still
undone on the other pages it touched, where nothing has changed since — but it no longer
destroys work there.

**Default stays last-writer-wins** for the API and MCP (ADR 013's contract: an agent undoing its
own last write gets exactly the before-state). Whether agents should default to keeping later
edits too is left to the owner; the tool description now tells them when to pass the flag.

**Rejected: a whole-entity version check (`if_version`).** A later collapse toggle or an unrelated
property would then block restoring a block's text. The `changes` rows already carry each later
batch's before/after image, so the check is per field at no extra storage.

**Rejected: page-scoping the walk** (undo only this page's entities). It would leave a cross-page
batch half-undone — a merge's moved blocks back but its link rewrites not, a replace undone on one
page of the 300 it touched — which is a state no write ever produced.
