ADR 015: Live UI control — a dedicated `/ui/live` socket, the existing command registry, consent by default-asymmetry
Date: 2026-09-10. Status: accepted (design); implementation lands with M2's client.
Follows from ADR 013's forward-looking item. Full survey and rationale:
docs/research/09-live-ui-control.md.
Decision#
An agent can observe and drive a currently-open client window, in addition to (never instead of) the headless 18-tool data API, which keeps working with no client open at all.
- Transport: a second, dedicated WebSocket (
/ui/live), not a multiplexed extension of the sync "poke" socket (ADR 003). The sync channel is mandatory, always-on, per-device, and protected by property tests; this one is optional, per-window, opt-in, and needs genuine request/response semantics. Keeping them separate buys one clean invariant: the socket being open is the feature being active, which is exactly the fact the consent badge displays. - Identity is a
window_id, not adevice_id. One device routinely has several tabs open (the client already elects a writer tab vianavigator.locks). Each window mints a randomwindow_idinsessionStorage; the server keeps an in-memory-only table of live sessions — ephemeral by nature, never persisted, never somethingrebuild()reproduces. Resolution policy differs by risk: reads (ui_get_state,ui_list_windows) default to the most-recently-active window and return the other candidates alongside;ui_run_commandwith several windows open and nowindow_idreturnsambiguous_windowrather than guessing, since acting on the wrong window is materially worse than reading the wrong one. - Observed state is the client's existing
WhenContext/CommandContext, serialized — page, zoom root, focused/selected blocks, cursor offset, viewport, open panel or dialog. nooklet does not build a second "what is the user doing" model for agents; it exposes the one the command dispatcher already computes before every keystroke. - Actions are the existing
Commandregistry (ADR 009), invoked by id + args, with each command's ownwhenclause and scope check enforced exactly as for a human. Only two new commands are needed (nav.openPage,nav.revealBlock) for the one thing no keybinding-driven command does today: jump straight to a known page or block with no picker. - New MCP surface:
ui_list_windows,ui_get_state,ui_run_command, plus thinui_navigate/ui_highlightwrappers for the two constant cases. Tools, not resources, are the day-one surface — MCP 2026-07-28 does natively support subscribable resources, and anooklet://ui/window/{id}resource should be added in parallel later, but Claude Code currently surfaces resources only as manual@mentions, not something an autonomous loop polls. - Consent is asymmetric by default, and visible rather than hidden. Two independent, device-local, unsynced toggles: "let agents view this window" defaults on (read-only, no write risk, and it is what makes the feature discoverable); "let agents control this window" defaults off, one deliberate opt-in per window. A persistent status-bar badge shows off/observed/controlled and opens a recent-activity log. Anything a remote command touches gets a distinct flash attributed to the agent, in a different colour from the human's own accent. Since 2026-09-13 (B-540) the badge is a muted icon button rather than a text pill, its sentence moved to the tooltip and accessible name — except that controlled stays drawn in the agent accent on a tinted ground, because a consent signal for "an agent can act here" must not blend into the chrome.
- A new
ui:controlcapability flag, orthogonal toread/write/admin. A broadwritetoken for headless data cleanup should not thereby be able to drive someone's screen, and aread+ui:controltoken is a coherent "can watch and point at things, cannot edit" grant. - No blocking confirmation on remote commands. The human is by construction looking at the
screen when this channel is live, and everything destructive here is soft-deleted and
reversible, so destructive ops surface a toast with one-tap Undo via
batch_undo(ADR 013) instead of a dialog per call.page_delete's existing hardrequiresUserInteractiongate is unchanged — that one is page-level and harder to casually undo.
Why#
Most competitors' AI integrations are headless: they read and write a backend the human's UI also happens to read. Letting an agent see and drive the actual screen someone is looking at, live and attributably, is a genuinely different capability, and the architecture already contains both halves it needs — a live server-client connection and a declarative command registry — so this is cheap to build and expensive to retrofit.
The closest real precedent is VS Code's Language Model Tools API plus Copilot Chat's editor
context: the only surveyed system that is also "an app with its own command registry and live
focus/selection model, exposed to an LLM over a semantic channel," including a when-gated
availability model nearly identical to nooklet's own. Figma's agent canvas access is the second
template (structured operations against the document model, gated by an explicit capability).
Home Assistant's WebSocket API — which the user already depends on daily — is the transport
pattern: subscribe for state, call services for actions, confirm effects by watching the stream.
Browser automation and "computer use" were rejected as the wrong shape: they exist to break into applications from the outside, paying in screenshots, vision tokens, pixel ambiguity, and a documented 2–5 seconds per action. nooklet owns both ends of this connection, so it can have a typed JSON snapshot and a command id in tens of milliseconds instead.
Consequences#
docs/spec/commands-and-keymap.mdneedsnav.openPage,nav.revealBlock, and aCommand.remoteInvocableflag added before implementation.- Client-side op emission needs an origin/actor override so remotely-triggered local edits are attributed to the agent, not the human, in the audit trail.
ui_run_commandneeds its own rate-limit bucket, separate fromwrite's — a live collaboration burst has a different natural rate than a batch import.- Whether
ui:controlneeds a third, narrower tier is left open; per-command scope checks may already give a fine enough boundary. - A desktop shell could later add OS-level extras a browser tab cannot have (raising the window
when a remote command fires), through ADR 005's existing
platformadapter. Explicitly optional and not required by this design.
Amendment (2026-10-04, B-708): off and hidden by default on a phone#
Owner: "agents control should probably be off for mobile? that doesn't make much sense."
Change. On the Capacitor app and on any touch-only device (what the command system already
calls mobile: an iOS/Android user agent, or a coarse pointer with no hover), "let agents view
this window" now defaults off, and the top-bar badge is not drawn while both switches are off.
Both switches are also in Settings → Agent access, on every device; on a phone that is where the
channel is turned on, and turning it on brings the badge back. A choice the device already stored
wins either way. Desktop behaviour is unchanged (view on, control off, badge always shown).
Code: apps/web/src/live/device-default.ts, consent.ts#liveBadgeShown.
Why. This channel exists for an agent working beside a window someone is looking at — the same computer, the next terminal (§"Why" above). A phone is rarely that: nobody runs an agent next to it, it is one more socket to hold open on a radio that sleeps, and the badge took a slot in a top bar with little room. "Discoverable by default" (§6) was the argument for view-on, and it is weakest exactly where the feature is least useful.
What stays true. "The socket being open is the feature being active" (§1): with view off,
app/CommandLayer.tsx#LiveConnection never calls connectLiveSocket, so the device opens no
/ui/live at all — checked by e2e/tests/phone-ui.spec.ts "B-708…" (Chromium and WebKit, iPhone
viewport with touch), which watches the page's WebSockets before and after the Settings opt-in. And
the badge is shown whenever an agent can see the window (§6): hiding it is tied to the switches
being off, never to the device alone.
Cost. Someone who does want an agent on their phone or tablet must find the switch in Settings. A touch-screen laptop with a mouse attached reports hover, so it keeps the desktop default; an iPad with a keyboard but no pointer gets the phone default.