Skip to content

~8 min readgrounded in apps/web/app/api/devices/** · apps/worker/src/{devices,relay}.ts · packages/contracts/src/devices.ts · examples/echo-device/

Devices·

A device is a row in the worker's D1 devices table that belongs to one user. The agent's use_device tool, the /devices page and the phone apps all read that table. What makes the design work across a robot arm, a laptop daemon, an e-ink display and a 3D printer is that there are only two traffic patterns, chosen by the device's kind:

Dial-in (daemon, browser, cli) Dial-out (endpoint)
Who holds the credential the device: a tind_… token, returned once at enrollment; the worker keeps only its SHA-256 the worker: the device's url and bearer secret
Traffic the device heartbeats and polls a relay mailbox; the app never connects to it the worker makes outbound HTTPS calls to the device's own API
Presence online = heartbeat within the last 60 s online: null — reachability is unknown until invoked
Typical hardware a laptop running tiny-vercel, a phone gatewaying a BLE board, an ESP32 display a printer, an arm, anything already serving an authenticated HTTPS API
Reference npx tiny-vercel (CLI) examples/echo-device/

Every device route bounds its worker round-trip to 10 s (the endpoint chat proxy excepted — an agent turn on a robot gets 90 s and runs on Node, since Edge 504s a response whose first byte takes over 25 s).

Enrollment·

The owner's session is the enrollment authority. POST /api/devices {name, platform?, kind?, capabilities?} → {ok, device_id, device_token}. The token appears exactly once. For an endpoint device the body carries url and secret instead and no token is minted. DELETE /api/devices {deviceId} revokes instantly. GET /api/devices lists the fleet with id, name, kind, online, last_seen, capabilities, lan_url, url.

Lost the token — a phone looking at a board that was paired from a laptop, a reinstall that emptied the Keychain? Do not enroll the hardware again: that mints a second row and the first sits in the fleet forever with a frozen last_seen. POST /api/devices/adopt {deviceId} rotates the token for a device you already own and returns the new one once; the old token stops working immediately, the row keeps its id, events and transcripts. Endpoint devices cannot be adopted — nothing inbound may speak as them.

The consent flow that gives a laptop the session it needs to enroll is in Identity.

Dial-in devices·

No session on any of these — the device token authenticates and resolves the owner in one SQL lookup (id = ? AND token_hash = ? AND revoked = 0), so a device can only act inside its own user's world and a spoofed userId in a body is never read. They are also deliberately outside the per-IP free-tier limiter: a daemon heartbeats continuously and a wearable's wakes are not a quota.

Route Body What it does
POST /api/devices/heartbeat {deviceId, token, capabilities?, lanUrl?, wantUnread?} presence; refreshes capabilities and the LAN address the phone app can dial directly. A wrong token 401s without revealing whether the id exists
PUT /api/devices/relay {deviceId, token, max?} poll undelivered envelopes
PATCH /api/devices/relay {deviceId, token, inReplyTo, payload} answer one
POST /api/devices/event {deviceId, token, kind, detail?} push something the device noticed onto the owner's event ring — the half the pull model cannot cover (a wake word fires when it fires). kind is allowlisted by the worker
POST /api/devices/task-result {deviceId, token, taskId, summary?, result?} a background task the daemon offloaded has finished: deposited under a task_* ticket, one ring event, one push
POST /api/devices/transcript device token + transcript the paired phone stores words it transcribed on-device; outlives the envelope
POST /api/devices/ask device token + text or a media id a device asks its owner's tiny and gets {text, card?} back — prose plus an optional card an e-ink can render natively
POST /api/devices/messages {deviceId, token, op: send\|inbox\|thread\|unread, …} the DM rail for a device with a keyboard

The relay mailbox·

The owner side is session-gated: POST /api/devices/relay {toDevice, payload} queues an envelope (JSON, ≤ 8 KB) and returns {id}; GET /api/devices/relay?inReplyTo=<id> fetches the reply. Envelopes are typed — {type:'invoke', prompt} asks the device's local agent to do something with its own tools, {type:'notify', title, body, tag, url} shows a notification — and anything else is passed through for firmware to interpret.

Two sweep tiers in apps/worker/src/relay.ts: undelivered envelopes are dead-lettered after 1 hour; delivered requests and their replies stay 24 hours, which is the window in which a task_* or batch_* ticket can be redeemed. RelaySendResult tells the caller exactly what happened when a send does not queue: no_such_device, too_big, bad_request, server_key, relay_fault, unreachable, no_envelope — each with delivered: 'no' | 'unknown' and whether a retry makes sense.

Endpoint devices·

The worker holds the bearer and dials the device; the device's credential never passes through the app. Actions are an allowlist, not a free path, because the caller is an LLM tool argument:

Action The worker calls Budget Returns
telemetry GET {url}/api/telemetry 20 s JSON
chat POST {url}/api/chat {prompt} 90 s {reply} (or {result}, {text}, plain text)
snapshot GET {url}/api/camera/snapshot 10 s image/png, image/jpeg or image/webp bytes — the content type is allowlisted because the frame is served back from your origin

Rules enforced in apps/worker/src/devices.ts: https:// and a public hostname — any IP literal in any encoding, localhost, .local, .internal and dotless hosts are refused, because the worker fetches this URL server-side and a private address would make the registry an SSRF pivot into Cloudflare's network; a non-empty secret; redirects are never followed (a 3xx could bounce the bearer to another origin); a 401/403 is reported as "device rejected our credential". A snapshot is a still frame on purpose — a multipart stream would pin a worker invocation open with no timeout able to fire.

App-side doors: GET /api/devices/endpoint?deviceId=&action=telemetry|snapshot (session-gated read proxy; GET so a snapshot works as an <img src>) and POST /api/devices/endpoint/chat {deviceId, prompt} (one agent turn, Node runtime). The web chat's use_device dials the worker directly with the internal key; a laptop has no such key, so its local use_device uses these two.

Try it without hardware·

examples/echo-device/ is a 90-line Node server that answers the three calls. Run it, expose it with any HTTPS tunnel, and enroll:

export ECHO_TOKEN=$(openssl rand -hex 24)
node examples/echo-device/server.mjs                 # http://127.0.0.1:8080
cloudflared tunnel --url http://127.0.0.1:8080       # prints https://<words>.trycloudflare.com
node examples/echo-device/enroll.mjs --app https://<your-app> --url https://<words>.trycloudflare.com --name echo

enroll.mjs walks the same consent flow as the CLI, enrolls with kind: "endpoint", then calls ?action=telemetry — one line proves app → worker → tunnel → your process → back. Then in chat: "ask my device named echo for its status."

What use_device does·

The agent tool has three actions. list reads the fleet with each device's declared capabilities — the description tells the model to match the task to a device that declares the power it needs and never to claim a device lacks something without reading that list. invoke sends an invoke envelope to a dial-in device (waiting about 45 s; a slower task returns pending: true with an envelope id and the user gets a push when it finishes) or a chat call to an endpoint device. result redeems an envelope, task_* or batch_* ticket within the 24-hour window. Only the owner's devices are reachable, and a machine is refused when asked to invoke itself.