~8 min readgrounded in apps/web/app/api/devices/** · apps/worker/src/{devices,relay}.ts · packages/contracts/src/devices.ts · examples/echo-device/
Devices·
A device is a row in the worker's D1 devices table that belongs to one user. The agent's use_device tool, the /devices page and the phone apps all read that table. What makes the design work across a robot arm, a laptop daemon, an e-ink display and a 3D printer is that there are only two traffic patterns, chosen by the device's kind:
Dial-in (daemon, browser, cli) |
Dial-out (endpoint) |
|
|---|---|---|
| Who holds the credential | the device: a tind_… token, returned once at enrollment; the worker keeps only its SHA-256 |
the worker: the device's url and bearer secret |
| Traffic | the device heartbeats and polls a relay mailbox; the app never connects to it | the worker makes outbound HTTPS calls to the device's own API |
| Presence | online = heartbeat within the last 60 s |
online: null — reachability is unknown until invoked |
| Typical hardware | a laptop running tiny-vercel, a phone gatewaying a BLE board, an ESP32 display |
a printer, an arm, anything already serving an authenticated HTTPS API |
| Reference | npx tiny-vercel (CLI) |
examples/echo-device/ |
Every device route bounds its worker round-trip to 10 s (the endpoint chat proxy excepted — an agent turn on a robot gets 90 s and runs on Node, since Edge 504s a response whose first byte takes over 25 s).
Enrollment·
The owner's session is the enrollment authority. POST /api/devices {name, platform?, kind?, capabilities?} → {ok, device_id, device_token}. The token appears exactly once. For an endpoint device the body carries url and secret instead and no token is minted. DELETE /api/devices {deviceId} revokes instantly. GET /api/devices lists the fleet with id, name, kind, online, last_seen, capabilities, lan_url, url.
Lost the token — a phone looking at a board that was paired from a laptop, a reinstall that emptied the Keychain? Do not enroll the hardware again: that mints a second row and the first sits in the fleet forever with a frozen last_seen. POST /api/devices/adopt {deviceId} rotates the token for a device you already own and returns the new one once; the old token stops working immediately, the row keeps its id, events and transcripts. Endpoint devices cannot be adopted — nothing inbound may speak as them.
The consent flow that gives a laptop the session it needs to enroll is in Identity.
Dial-in devices·
No session on any of these — the device token authenticates and resolves the owner in one SQL lookup (id = ? AND token_hash = ? AND revoked = 0), so a device can only act inside its own user's world and a spoofed userId in a body is never read. They are also deliberately outside the per-IP free-tier limiter: a daemon heartbeats continuously and a wearable's wakes are not a quota.
| Route | Body | What it does |
|---|---|---|
POST /api/devices/heartbeat |
{deviceId, token, capabilities?, lanUrl?, wantUnread?} |
presence; refreshes capabilities and the LAN address the phone app can dial directly. A wrong token 401s without revealing whether the id exists |
PUT /api/devices/relay |
{deviceId, token, max?} |
poll undelivered envelopes |
PATCH /api/devices/relay |
{deviceId, token, inReplyTo, payload} |
answer one |
POST /api/devices/event |
{deviceId, token, kind, detail?} |
push something the device noticed onto the owner's event ring — the half the pull model cannot cover (a wake word fires when it fires). kind is allowlisted by the worker |
POST /api/devices/task-result |
{deviceId, token, taskId, summary?, result?} |
a background task the daemon offloaded has finished: deposited under a task_* ticket, one ring event, one push |
POST /api/devices/transcript |
device token + transcript | the paired phone stores words it transcribed on-device; outlives the envelope |
POST /api/devices/ask |
device token + text or a media id | a device asks its owner's tiny and gets {text, card?} back — prose plus an optional card an e-ink can render natively |
POST /api/devices/messages |
{deviceId, token, op: send\|inbox\|thread\|unread, …} |
the DM rail for a device with a keyboard |
The relay mailbox·
The owner side is session-gated: POST /api/devices/relay {toDevice, payload} queues an envelope (JSON, ≤ 8 KB) and returns {id}; GET /api/devices/relay?inReplyTo=<id> fetches the reply. Envelopes are typed — {type:'invoke', prompt} asks the device's local agent to do something with its own tools, {type:'notify', title, body, tag, url} shows a notification — and anything else is passed through for firmware to interpret.
Two sweep tiers in apps/worker/src/relay.ts: undelivered envelopes are dead-lettered after 1 hour; delivered requests and their replies stay 24 hours, which is the window in which a task_* or batch_* ticket can be redeemed. RelaySendResult tells the caller exactly what happened when a send does not queue: no_such_device, too_big, bad_request, server_key, relay_fault, unreachable, no_envelope — each with delivered: 'no' | 'unknown' and whether a retry makes sense.
Endpoint devices·
The worker holds the bearer and dials the device; the device's credential never passes through the app. Actions are an allowlist, not a free path, because the caller is an LLM tool argument:
| Action | The worker calls | Budget | Returns |
|---|---|---|---|
telemetry |
GET {url}/api/telemetry |
20 s | JSON |
chat |
POST {url}/api/chat {prompt} |
90 s | {reply} (or {result}, {text}, plain text) |
snapshot |
GET {url}/api/camera/snapshot |
10 s | image/png, image/jpeg or image/webp bytes — the content type is allowlisted because the frame is served back from your origin |
Rules enforced in apps/worker/src/devices.ts: https:// and a public hostname — any IP literal in any encoding, localhost, .local, .internal and dotless hosts are refused, because the worker fetches this URL server-side and a private address would make the registry an SSRF pivot into Cloudflare's network; a non-empty secret; redirects are never followed (a 3xx could bounce the bearer to another origin); a 401/403 is reported as "device rejected our credential". A snapshot is a still frame on purpose — a multipart stream would pin a worker invocation open with no timeout able to fire.
App-side doors: GET /api/devices/endpoint?deviceId=&action=telemetry|snapshot (session-gated read proxy; GET so a snapshot works as an <img src>) and POST /api/devices/endpoint/chat {deviceId, prompt} (one agent turn, Node runtime). The web chat's use_device dials the worker directly with the internal key; a laptop has no such key, so its local use_device uses these two.
Try it without hardware·
examples/echo-device/ is a 90-line Node server that answers the three calls. Run it, expose it with any HTTPS tunnel, and enroll:
export ECHO_TOKEN=$(openssl rand -hex 24)
node examples/echo-device/server.mjs # http://127.0.0.1:8080
cloudflared tunnel --url http://127.0.0.1:8080 # prints https://<words>.trycloudflare.com
node examples/echo-device/enroll.mjs --app https://<your-app> --url https://<words>.trycloudflare.com --name echo
enroll.mjs walks the same consent flow as the CLI, enrolls with kind: "endpoint", then calls ?action=telemetry — one line proves app → worker → tunnel → your process → back. Then in chat: "ask my device named echo for its status."
What use_device does·
The agent tool has three actions. list reads the fleet with each device's declared capabilities — the description tells the model to match the task to a device that declares the power it needs and never to claim a device lacks something without reading that list. invoke sends an invoke envelope to a dial-in device (waiting about 45 s; a slower task returns pending: true with an envelope id and the user gets a push when it finishes) or a chat call to an endpoint device. result redeems an envelope, task_* or batch_* ticket within the 24-hour window. Only the owner's devices are reachable, and a machine is refused when asked to invoke itself.