Skip to content

~5 min readgrounded in apps/web/lib/chat/model.ts · apps/web/lib/chat/model-config.ts · apps/web/app/api/model-providers/route.ts · docs/ENV.md

Configure models·

createModel() in apps/web/lib/chat/model.ts resolves the model for a chat request from three layers, most specific first:

  1. Per-request headers — x-tiny-model-provider, x-tiny-model-id, x-tiny-model-api-key, x-tiny-model-base-url, x-tiny-model-max-tokens, x-tiny-model-region, x-tiny-model-additional-fields. Native apps and the tiny-vercel CLI send these from local settings (tiny-vercel onboard picks the provider; TINY_MODEL_* overrides).
  2. The user's synced provider settings — rows in /api/model-providers, one per provider, with the active one chosen. Keys are encrypted at rest in D1 (MODEL_CONFIG_ENC_KEY, falling back to INTERNAL_API_KEY) and can be pulled by the user's own devices with GET /api/model-providers?full=1 behind their session.
  3. The deployment default — TINY_MODEL_PROVIDER and the matching key on the Vercel app. This is what anonymous visitors and users without their own key get, so it is also what the free-tier rate limit protects (NEXT_PUBLIC_FREE_TIER_REQUESTS_PER_DAY, default 50 per day per IP).

If no key exists at any layer, preflightModelCheck refuses before the agent is built: No API key configured for provider '<name>'. Bring your own key via model settings or configure the server.

Providers the server routes to·

TINY_MODEL_PROVIDER Variables Default model id Notes
openai (default) OPENAI_API_KEY, OPENAI_MODEL_ID gpt-5.6-luna also the embedding provider (text-embedding-3-small) — this key is required regardless of chat provider
bedrock AWS_BEARER_TOKEN_BEDROCK, BEDROCK_MODEL_ID, BEDROCK_REGION or AWS_REGION (default us-west-2) global.anthropic.claude-sonnet-4-6 Edge-safe direct fetch to ConverseStream with a bearer token — no AWS SDK, because node:http is not available on the Edge runtime
gemini (normalised to google) GEMINI_API_KEY or GOOGLE_API_KEY, GEMINI_MODEL_ID gemini-2.5-flash
gateway / vercel AI_GATEWAY_API_KEY, AI_GATEWAY_MODEL_ID openai/gpt-5-mini Vercel AI Gateway; any model the gateway serves

STRANDS_ADDITIONAL_REQUEST_FIELDS (JSON) is passed through to the provider as extra request fields — for example {"anthropic_beta":["context-1m-2025-08-07"]} on Bedrock. Temperature is 1 everywhere.

OpenAI-compatible providers·

Anything with a /v1/chat/completions endpoint works through the openai case with a base URL: when baseUrl is set the client switches from OpenAI's Responses API (which only api.openai.com serves) to the chat API. The Settings screen ships these presets (apps/web/lib/chat/model-config.ts):

Preset Base URL
Anthropic https://api.anthropic.com/v1/
OpenRouter https://openrouter.ai/api/v1
Groq https://api.groq.com/openai/v1
DeepSeek https://api.deepseek.com/v1
Mistral https://api.mistral.ai/v1
xAI (Grok) https://api.x.ai/v1
Perplexity https://api.perplexity.ai
Google Gemini (OpenAI-compatible) https://generativelanguage.googleapis.com/v1beta/openai/
Custom whatever you type

Two more presets are not servers at all: Tiny (free, rate-limited) uses the deployment default, and On-device (WebLLM, offline) runs a model in the browser and never calls /api/chat for generation.

Set the deployment default·

add() { printf '%s' "$2" | npx vercel env add "$1" production --force; }
add TINY_MODEL_PROVIDER bedrock
add AWS_BEARER_TOKEN_BEDROCK <token>
add BEDROCK_MODEL_ID global.anthropic.claude-sonnet-4-6
add BEDROCK_REGION us-west-2
npx vercel redeploy <your-project>.vercel.app

Keep OPENAI_API_KEY set on both the app and the worker even when chat runs elsewhere — memory search embeds with it.

Let users bring their own key·

Nothing to enable. Settings → Model lists the presets above; a saved provider is upserted with POST /api/model-providers {provider, modelId?, baseUrl?, region?, maxTokens?, additionalFields?, apiKey?, isActive?} and the chat route reads the active one. Set a dedicated MODEL_CONFIG_ENC_KEY on the worker (openssl rand -hex 32) before the first user stores a key, so rotating INTERNAL_API_KEY later does not orphan every stored credential.