# ADR 015: Live UI control — a dedicated `/ui/live` socket, the existing command registry, consent by default-asymmetry

> Follows from ADR 013's forward-looking item. Full survey and rationale: docs/research/09-live-ui-control.md.

Source: https://nooklet.danielalder.cz/decisions/015-live-ui-control-channel

Date: 2026-09-10. Status: accepted (design); implementation lands with M2's client.

Follows from ADR 013's forward-looking item. Full survey and rationale:
`docs/research/09-live-ui-control.md`.

## Decision

An agent can observe and drive a *currently-open* client window, in addition to (never instead
of) the headless 18-tool data API, which keeps working with no client open at all.

1. **Transport: a second, dedicated WebSocket (`/ui/live`)**, not a multiplexed extension of the
   sync "poke" socket (ADR 003). The sync channel is mandatory, always-on, per-device, and
   protected by property tests; this one is optional, per-window, opt-in, and needs genuine
   request/response semantics. Keeping them separate buys one clean invariant: **the socket being
   open *is* the feature being active**, which is exactly the fact the consent badge displays.
2. **Identity is a `window_id`, not a `device_id`.** One device routinely has several tabs open
   (the client already elects a writer tab via `navigator.locks`). Each window mints a random
   `window_id` in `sessionStorage`; the server keeps an in-memory-only table of live sessions —
   ephemeral by nature, never persisted, never something `rebuild()` reproduces.
   Resolution policy differs by risk: reads (`ui_get_state`, `ui_list_windows`) default to the
   most-recently-active window and return the other candidates alongside; `ui_run_command` with
   several windows open and no `window_id` returns `ambiguous_window` rather than guessing, since
   acting on the wrong window is materially worse than reading the wrong one.
3. **Observed state is the client's existing `WhenContext`/`CommandContext`, serialized** — page,
   zoom root, focused/selected blocks, cursor offset, viewport, open panel or dialog. nooklet does
   not build a second "what is the user doing" model for agents; it exposes the one the command
   dispatcher already computes before every keystroke.
4. **Actions are the existing `Command` registry (ADR 009), invoked by id + args**, with each
   command's own `when` clause and scope check enforced exactly as for a human. Only two new
   commands are needed (`nav.openPage`, `nav.revealBlock`) for the one thing no keybinding-driven
   command does today: jump straight to a known page or block with no picker.
5. **New MCP surface**: `ui_list_windows`, `ui_get_state`, `ui_run_command`, plus thin
   `ui_navigate`/`ui_highlight` wrappers for the two constant cases. Tools, not resources, are the
   day-one surface — MCP 2026-07-28 *does* natively support subscribable resources, and a
   `nooklet://ui/window/{id}` resource should be added in parallel later, but Claude Code currently
   surfaces resources only as manual `@mentions`, not something an autonomous loop polls.
6. **Consent is asymmetric by default, and visible rather than hidden.** Two independent,
   device-local, unsynced toggles: *"let agents view this window"* defaults **on** (read-only, no
   write risk, and it is what makes the feature discoverable); *"let agents control this window"*
   defaults **off**, one deliberate opt-in per window. A persistent status-bar badge shows
   off/observed/controlled and opens a recent-activity log. Anything a remote command touches gets
   a distinct flash attributed to the agent, in a different colour from the human's own accent.
   Since 2026-09-13 (B-540) the badge is a muted icon button rather than a text pill, its sentence
   moved to the tooltip and accessible name — except that *controlled* stays drawn in the agent
   accent on a tinted ground, because a consent signal for "an agent can act here" must not blend
   into the chrome.
7. **A new `ui:control` capability flag**, orthogonal to `read`/`write`/`admin`. A broad `write`
   token for headless data cleanup should not thereby be able to drive someone's screen, and a
   `read` + `ui:control` token is a coherent "can watch and point at things, cannot edit" grant.
8. **No blocking confirmation on remote commands.** The human is by construction looking at the
   screen when this channel is live, and everything destructive here is soft-deleted and
   reversible, so destructive ops surface a toast with one-tap Undo via `batch_undo` (ADR 013)
   instead of a dialog per call. `page_delete`'s existing hard `requiresUserInteraction` gate is
   unchanged — that one is page-level and harder to casually undo.

## Why

Most competitors' AI integrations are headless: they read and write a backend the human's UI also
happens to read. Letting an agent see and drive the actual screen someone is looking at, live and
attributably, is a genuinely different capability, and the architecture already contains both
halves it needs — a live server-client connection and a declarative command registry — so this is
cheap to build and expensive to retrofit.

The closest real precedent is **VS Code's Language Model Tools API plus Copilot Chat's editor
context**: the only surveyed system that is also "an app with its own command registry and live
focus/selection model, exposed to an LLM over a semantic channel," including a `when`-gated
availability model nearly identical to nooklet's own. Figma's agent canvas access is the second
template (structured operations against the document model, gated by an explicit capability).
Home Assistant's WebSocket API — which the user already depends on daily — is the transport
pattern: subscribe for state, call services for actions, confirm effects by watching the stream.

Browser automation and "computer use" were rejected as the wrong shape: they exist to break into
applications from the outside, paying in screenshots, vision tokens, pixel ambiguity, and a
documented 2–5 seconds per action. nooklet owns both ends of this connection, so it can have a
typed JSON snapshot and a command id in tens of milliseconds instead.

## Consequences

- `docs/spec/commands-and-keymap.md` needs `nav.openPage`, `nav.revealBlock`, and a
  `Command.remoteInvocable` flag added before implementation.
- Client-side op emission needs an origin/actor override so remotely-triggered local edits are
  attributed to the agent, not the human, in the audit trail.
- `ui_run_command` needs its own rate-limit bucket, separate from `write`'s — a live collaboration
  burst has a different natural rate than a batch import.
- Whether `ui:control` needs a third, narrower tier is left open; per-command scope checks may
  already give a fine enough boundary.
- A desktop shell could later add OS-level extras a browser tab cannot have (raising the window
  when a remote command fires), through ADR 005's existing `platform` adapter. Explicitly optional
  and not required by this design.

## Amendment (2026-10-04, B-708): off and hidden by default on a phone

Owner: "agents control should probably be off for mobile? that doesn't make much sense."

**Change.** On the Capacitor app and on any touch-only device (what the command system already
calls `mobile`: an iOS/Android user agent, or a coarse pointer with no hover), "let agents view
this window" now defaults **off**, and the top-bar badge is not drawn while both switches are off.
Both switches are also in Settings → Agent access, on every device; on a phone that is where the
channel is turned on, and turning it on brings the badge back. A choice the device already stored
wins either way. Desktop behaviour is unchanged (view on, control off, badge always shown).
Code: `apps/web/src/live/device-default.ts`, `consent.ts#liveBadgeShown`.

**Why.** This channel exists for an agent working beside a window someone is looking at — the same
computer, the next terminal (§"Why" above). A phone is rarely that: nobody runs an agent next to
it, it is one more socket to hold open on a radio that sleeps, and the badge took a slot in a top
bar with little room. "Discoverable by default" (§6) was the argument for view-on, and it is
weakest exactly where the feature is least useful.

**What stays true.** "The socket being open *is* the feature being active" (§1): with view off,
`app/CommandLayer.tsx#LiveConnection` never calls `connectLiveSocket`, so the device opens no
`/ui/live` at all — checked by `e2e/tests/phone-ui.spec.ts` "B-708…" (Chromium and WebKit, iPhone
viewport with touch), which watches the page's WebSockets before and after the Settings opt-in. And
the badge is shown whenever an agent can see the window (§6): hiding it is tied to the switches
being off, never to the device alone.

**Cost.** Someone who does want an agent on their phone or tablet must find the switch in Settings.
A touch-screen laptop with a mouse attached reports hover, so it keeps the desktop default; an
iPad with a keyboard but no pointer gets the phone default.
