~5 min readgrounded in apps/web/lib/chat/model.ts · apps/web/lib/chat/model-config.ts · apps/web/app/api/model-providers/route.ts · docs/ENV.md
Configure models·
createModel() in apps/web/lib/chat/model.ts resolves the model for a chat request from three layers, most specific first:
- Per-request headers —
x-tiny-model-provider,x-tiny-model-id,x-tiny-model-api-key,x-tiny-model-base-url,x-tiny-model-max-tokens,x-tiny-model-region,x-tiny-model-additional-fields. Native apps and thetiny-vercelCLI send these from local settings (tiny-vercel onboardpicks the provider;TINY_MODEL_*overrides). - The user's synced provider settings — rows in
/api/model-providers, one per provider, with the active one chosen. Keys are encrypted at rest in D1 (MODEL_CONFIG_ENC_KEY, falling back toINTERNAL_API_KEY) and can be pulled by the user's own devices withGET /api/model-providers?full=1behind their session. - The deployment default —
TINY_MODEL_PROVIDERand the matching key on the Vercel app. This is what anonymous visitors and users without their own key get, so it is also what the free-tier rate limit protects (NEXT_PUBLIC_FREE_TIER_REQUESTS_PER_DAY, default 50 per day per IP).
If no key exists at any layer, preflightModelCheck refuses before the agent is built: No API key configured for provider '<name>'. Bring your own key via model settings or configure the server.
Providers the server routes to·
TINY_MODEL_PROVIDER |
Variables | Default model id | Notes |
|---|---|---|---|
openai (default) |
OPENAI_API_KEY, OPENAI_MODEL_ID |
gpt-5.6-luna |
also the embedding provider (text-embedding-3-small) — this key is required regardless of chat provider |
bedrock |
AWS_BEARER_TOKEN_BEDROCK, BEDROCK_MODEL_ID, BEDROCK_REGION or AWS_REGION (default us-west-2) |
global.anthropic.claude-sonnet-4-6 |
Edge-safe direct fetch to ConverseStream with a bearer token — no AWS SDK, because node:http is not available on the Edge runtime |
gemini (normalised to google) |
GEMINI_API_KEY or GOOGLE_API_KEY, GEMINI_MODEL_ID |
gemini-2.5-flash |
|
gateway / vercel |
AI_GATEWAY_API_KEY, AI_GATEWAY_MODEL_ID |
openai/gpt-5-mini |
Vercel AI Gateway; any model the gateway serves |
STRANDS_ADDITIONAL_REQUEST_FIELDS (JSON) is passed through to the provider as extra request fields — for example {"anthropic_beta":["context-1m-2025-08-07"]} on Bedrock. Temperature is 1 everywhere.
OpenAI-compatible providers·
Anything with a /v1/chat/completions endpoint works through the openai case with a base URL: when baseUrl is set the client switches from OpenAI's Responses API (which only api.openai.com serves) to the chat API. The Settings screen ships these presets (apps/web/lib/chat/model-config.ts):
| Preset | Base URL |
|---|---|
| Anthropic | https://api.anthropic.com/v1/ |
| OpenRouter | https://openrouter.ai/api/v1 |
| Groq | https://api.groq.com/openai/v1 |
| DeepSeek | https://api.deepseek.com/v1 |
| Mistral | https://api.mistral.ai/v1 |
| xAI (Grok) | https://api.x.ai/v1 |
| Perplexity | https://api.perplexity.ai |
| Google Gemini (OpenAI-compatible) | https://generativelanguage.googleapis.com/v1beta/openai/ |
| Custom | whatever you type |
Two more presets are not servers at all: Tiny (free, rate-limited) uses the deployment default, and On-device (WebLLM, offline) runs a model in the browser and never calls /api/chat for generation.
Set the deployment default·
add() { printf '%s' "$2" | npx vercel env add "$1" production --force; }
add TINY_MODEL_PROVIDER bedrock
add AWS_BEARER_TOKEN_BEDROCK <token>
add BEDROCK_MODEL_ID global.anthropic.claude-sonnet-4-6
add BEDROCK_REGION us-west-2
npx vercel redeploy <your-project>.vercel.app
Keep OPENAI_API_KEY set on both the app and the worker even when chat runs elsewhere — memory search embeds with it.
Let users bring their own key·
Nothing to enable. Settings → Model lists the presets above; a saved provider is upserted with POST /api/model-providers {provider, modelId?, baseUrl?, region?, maxTokens?, additionalFields?, apiKey?, isActive?} and the chat route reads the active one. Set a dedicated MODEL_CONFIG_ENC_KEY on the worker (openssl rand -hex 32) before the first user stores a key, so rotating INTERNAL_API_KEY later does not orphan every stored credential.