~4 min readgrounded in apps/worker/src/scheduler.ts · apps/worker/wrangler.toml · apps/web/app/api/{jobs,job-run}/route.ts · apps/web/app/api/chat/route.ts (schedule tool)
Scheduling jobs·
"Every morning at 09:00, check my repos and message me on Telegram" is one sentence to the agent and one row in D1. Here is what happens to it.
Creating a job·
The agent's schedule tool (in the chat route) takes:
| Argument | Meaning |
|---|---|
action |
create · list · delete |
name |
the job's name |
prompt |
what the job should do each run |
schedule |
recurring — */30m, */2h (every N minutes/hours) or daily@09:00 (UTC) |
run_in_minutes |
one-shot — run once N minutes from now |
id |
for delete |
It calls POST /api/jobs in the app, which forwards to the worker's POST /jobs on the internal-key channel with the session's user id. Validation is strict on purpose: */0m is refused (it would pass a digit check but never fire, holding a quota slot forever); daily@25:70 is refused rather than silently rolled over into an unintended time; a one-shot must have a finite run_at. Quota: MAX_JOBS_PER_USER = 10 — the eleventh returns 429 job limit reached (10). GET /jobs?userId= and DELETE /jobs complete the set (the app exposes them as GET / DELETE /api/jobs).
The tick·
wrangler.toml declares crons = ["* * * * *"]. Every minute the worker's scheduled handler runs runDueJobs (alongside Telegram polling, tool-update sweeps and — when payments are on — the reconcilers). For each enabled job it decides, purely and testably, one of:
- fire — due now, or missed within the last 24 hours (
CATCH_UP_SECONDS): a job missed during a short outage still runs, once; - skip — not due;
- skip-stale — missed by more than 24 hours: advance
last_fired_atwithout running, so a long backlog does not fire all at once. A stale one-shot is also disabled — see below.
Exactly once. Two cron invocations can overlap. The double-fire guard is a compare-and-swap UPDATE … WHERE last_fired_at IS ? — IS, not =, because a never-fired job has last_fired_at NULL and = NULL is never true in SQL, which would make such rows unclaimable forever. Only one runner's UPDATE reports changes = 1; the other sees 0 and skips.
The run·
A due job is executed by POST /api/job-run in the app (X-Internal-Key guarded, Node runtime with maxDuration = 120 — the Edge runtime returns a platform 504 for any turn slower than ~25 s, measured). It runs one non-streaming agent turn as the job's tiny with the owner's capability set: forged my_* tools, the tiny's OpenAPI skills, its MCP servers, server memory (learn / recall / unlearn) and use_telegram. The agent gets JOB_DEADLINE_S = 50 seconds before it is cancelled, so the worker's own wait never outlives the run.
Each run writes a job_runs row and one event on the owner's ring: job_result with a ✅ push on success, job_error with a ❌ push on failure.
The one silent case, made loud·
A one-shot whose fire time is more than a day in the past when the tick sees it — a long outage, or simply an agent that mis-parsed a date and scheduled it far in the past — is abandoned: enabled = 0, and it will never run. That used to be the only outcome nobody was told about, and the only one where the person has to act. Now it pushes a sentence that names the due time (not "now" — last_fired_at is about to be overwritten with the moment of abandonment, a time the job provably did not run at) and says the job is off and theirs to re-schedule. Recurring jobs are deliberately not announced when they skip a stale slot: an */5m job fires again in five minutes and nothing was lost; a push per missed slot after an outage would be a flood.
Operating notes·
- Jobs run on the worker's clock in UTC;
daily@HH:MMis UTC. - The cron is part of the worker deploy — there is nothing to enable in the dashboard;
wrangler deployregisters the trigger. - A job's tiny must still exist and be owned by the job's user at run time; the run executes as that user, so their memory and tools are what the job sees.