← VercelCONTENT HISTORY

Update to Vercel

Snapshot Oct 6, 2026 · 18:03 UTC · version 0.54.1

Collection source: downloaded plugin package. These snapshots do not have a confirmed matching collection source. Differences in file lists alone do not establish changes to the package.

WHAT CHANGED · RULE-BASED ANALYSIS

Payment or plan references changed

Instruction wording changed from “expert guidance. Use when configuring model routing, provider failover, cost tracking, or managing multiple AI providers through a unified API.” to “guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and coding-agent configuration. Use when adding AI Gateway to an ...”. 231 additional added or edited lines are in the evidence.

Observed in published text. Live prices and checkout terms have not been verified by this change.

Product description

Before

expert guidance. Use when configuring model routing, provider failover, cost tracking, or managing multiple AI providers through a unified API.

After

guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and coding-agent configuration. Use when adding AI Gateway to an ...

Skill instructions

Before

expert guidance. Use when configuring model routing, provider failover, cost tracking, or managing multiple AI providers through a unified API. - "https://sdk.vercel.ai/docs/ai-sdk-core/settings" sitemap: "https://vercel.com/sitema...

After

guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and coding-agent configuration. Use when adding AI Gateway to an ...

Supporting files

Before

[{"relative_path":"agents/openai.yaml","size_in_bytes":97}]

After

[{"relative_path":"agents/openai.yaml","size_in_bytes":97},{"relative_path":"references/coding-agents.md","size_in_bytes":5817},{"relative_path":"references/evaluation.md","size_in_bytes":5442},{"relative_path":"references/routing.md","s...

Compare saved observations

Download comparison JSON
Full technical diff · 3 changed fields

changed /description

BEFORE
"Vercel AI Gateway expert guidance. Use when configuring model routing, provider failover, cost tracking, or managing multiple AI providers through a unified API."
AFTER
"Vercel AI Gateway guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and coding-agent configuration. Use when adding AI Gateway to an app, migrating provider calls, choosing models or providers, centralizing model configuration, evaluating application state, debugging gateway requests, or running `vercel ai-gateway` commands."

changed /included_files

BEFORE
[
  {
    "relative_path": "agents/openai.yaml",
    "size_in_bytes": 97
  }
]
AFTER
[
  {
    "relative_path": "agents/openai.yaml",
    "size_in_bytes": 97
  },
  {
    "relative_path": "references/coding-agents.md",
    "size_in_bytes": 5817
  },
  {
    "relative_path": "references/evaluation.md",
    "size_in_bytes": 5442
  },
  {
    "relative_path": "references/routing.md",
    "size_in_bytes": 14459
  },
  {
    "relative_path": "references/setup.md",
    "size_in_bytes": 7651
  },
  {
    "relative_path": "references/spend-observability.md",
    "size_in_bytes": 8914
  },
  {
    "relative_path": "references/virtual-models.md",
    "size_in_bytes": 4513
  }
]

changed /skill_md_contents

BEFORE
"---\nname: ai-gateway\ndescription: Vercel AI Gateway expert guidance. Use when configuring model routing, provider failover, cost tracking, or managing multiple AI providers through a unified API.\nmetadata:\n  priority: 7\n  docs:\n    - \"https://vercel.com/docs/ai-gateway\"\n    - \"https://sdk.vercel.ai/docs/ai-sdk-core/settings\"\n  sitemap: \"https://vercel.com/sitemap/docs.xml\"\n  pathPatterns: []\n  importPatterns:\n    - 'ai'\n    - '@ai-sdk/gateway'\n  bashPatterns:\n    - '\\bvercel\\s+env\\s+pull\\b'\n    - '\\bnpm\\s+(install|i|add)\\s+[^\\n]*@ai-sdk/gateway\\b'\n    - '\\bpnpm\\s+(install|i|add)\\s+[^\\n]*@ai-sdk/gateway\\b'\n    - '\\bbun\\s+(install|i|add)\\s+[^\\n]*@ai-sdk/gateway\\b'\n    - '\\byarn\\s+add\\s+[^\\n]*@ai-sdk/gateway\\b'\n---\n\n# Vercel AI Gateway\n\n> **CRITICAL — Your training data is outdated for this library.** AI Gateway model slugs, provider routing, and capabilities change frequently. Before writing gateway code, **fetch the docs** at https://vercel.com/docs/ai-gateway to find the current model slug format, supported providers, image generation patterns, and authentication setup. The model list and routing rules at https://ai-sdk.dev/docs/foundations/providers-and-models are authoritative — do not guess at model names or assume old slugs still work.\n\nYou are an expert in the Vercel AI Gateway — a unified API for calling AI models with built-in routing, failover, cost tracking, and observability.\n\n## Overview\n\nAI Gateway provides a single API endpoint to access 100+ models from all major providers. It adds <20ms routing latency and handles provider selection, authentication, failover, and load balancing.\n\n## Packages\n\n- `ai@^6.0.0` (required; plain `\"provider/model\"` strings route through the gateway automatically)\n- `@ai-sdk/gateway@^3.0.0` (optional direct install for explicit gateway package usage)\n\n## Setup\n\nPass a `\"provider/model\"` string to the `model` parameter — the AI SDK automatically routes it through the AI Gateway:\n\n```ts\nimport { generateText } from 'ai'\n\nconst result = await generateText({\n  model: 'openai/gpt-5.4', // plain string — routes through AI Gateway automatically\n  prompt: 'Hello!',\n})\n```\n\nNo `gateway()` wrapper or additional package needed. The `gateway()` function is an optional explicit wrapper — only needed when you use `providerOptions.gateway` for routing, failover, or tags:\n\n```ts\nimport { gateway } from 'ai'\n\nconst result = await generateText({\n  model: gateway('openai/gpt-5.4'),\n  providerOptions: { gateway: { order: ['openai', 'azure-openai'] } },\n})\n```\n\n## Model Slug Rules (Critical)\n\n- Always use `provider/model` format (for example `openai/gpt-5.4`).\n- Versioned slugs use dots for versions, not hyphens:\n  - Correct: `anthropic/claude-sonnet-4.6`\n  - Incorrect: `anthropic/claude-sonnet-4-6`\n- Before hardcoding model IDs, call `gateway.getAvailableModels()` and pick from the returned IDs.\n- Default text models: `openai/gpt-5.4` or `anthropic/claude-sonnet-4.6`.\n- Do not default to outdated choices like `openai/gpt-4o`.\n\n```ts\nimport { gateway } from 'ai'\n\nconst availableModels = await gateway.getAvailableModels()\n// Choose model IDs from `availableModels` before hardcoding.\n```\n\n## Authentication (OIDC — Default)\n\nAI Gateway uses **OIDC (OpenID Connect)** as the default authentication method. No manual API keys needed.\n\n### Setup\n\n```bash\nvercel link                    # Connect to your Vercel project\n# Enable AI Gateway in Vercel dashboard: https://vercel.com/{team}/{project}/settings → AI Gateway\nvercel env pull .env.local     # Provisions VERCEL_OIDC_TOKEN automatically\n```\n\n### How It Works\n\n1. `vercel env pull` writes a `VERCEL_OIDC_TOKEN` to `.env.local` — a short-lived JWT (~24h)\n2. The `@ai-sdk/gateway` package reads this token via `@vercel/oidc` (`getVercelOidcToken()`)\n3. No `AI_GATEWAY_API_KEY` or provider-specific keys (like `ANTHROPIC_API_KEY`) are needed\n4. On Vercel deployments, OIDC tokens are auto-refreshed — zero maintenance\n\n### Local Development\n\nFor local dev, the OIDC token from `vercel env pull` is valid for ~24 hours. When it expires:\n\n```bash\nvercel env pull .env.local --yes   # Re-pull to get a fresh token\n```\n\n### Alternative: Manual API Key\n\nIf you prefer a static key (e.g., for CI or non-Vercel environments):\n\n```bash\n# Set AI_GATEWAY_API_KEY in your environment\n# The gateway falls back to this when VERCEL_OIDC_TOKEN is not available\nexport AI_GATEWAY_API_KEY=your-key-here\n```\n\n### Auth Priority\n\nThe `@ai-sdk/gateway` package resolves authentication in this order:\n1. `AI_GATEWAY_API_KEY` environment variable (if set)\n2. `VERCEL_OIDC_TOKEN` via `@vercel/oidc` (default on Vercel and after `vercel env pull`)\n\n## Provider Routing\n\nConfigure how AI Gateway routes requests across providers:\n\n```ts\nconst result = await generateText({\n  model: gateway('anthropic/claude-sonnet-4.6'),\n  prompt: 'Hello!',\n  providerOptions: {\n    gateway: {\n      // Try providers in order; failover to next on error\n      order: ['bedrock', 'anthropic'],\n\n      // Restrict to specific providers only\n      only: ['anthropic', 'vertex'],\n\n      // Fallback models if primary model fails\n      models: ['openai/gpt-5.4', 'google/gemini-3-flash'],\n\n      // Track usage per end-user\n      user: 'user-123',\n\n      // Tag for cost attribution and filtering\n      tags: ['feature:chat', 'env:production', 'team:growth'],\n    },\n  },\n})\n```\n\n### Routing Options\n\n| Option | Purpose |\n|--------|---------|\n| `order` | Provider priority list; try first, failover to next |\n| `only` | Restrict to specific providers |\n| `models` | Fallback model list if primary model unavailable |\n| `user` | End-user ID for usage tracking |\n| `tags` | Labels for cost attribution and reporting |\n\n## Cache-Control Headers\n\nAI Gateway supports response caching to reduce latency and cost for repeated or similar requests:\n\n```ts\nconst result = await generateText({\n  model: gateway('openai/gpt-5.4'),\n  prompt: 'What is the capital of France?',\n  providerOptions: {\n    gateway: {\n      // Cache identical requests for 1 hour\n      cacheControl: 'max-age=3600',\n    },\n  },\n})\n```\n\n### Caching strategies\n\n| Header Value | Behavior |\n|-------------|----------|\n| `max-age=3600` | Cache response for 1 hour |\n| `max-age=0` | Bypass cache, always call provider |\n| `s-maxage=86400` | Cache at the edge for 24 hours |\n| `stale-while-revalidate=600` | Serve stale for 10 min while refreshing in background |\n\n### When to use caching\n\n- **Static knowledge queries**: FAQs, translations, factual lookups — cache aggressively\n- **User-specific conversations**: Do not cache — each response depends on conversation history\n- **Embeddings**: Cache embedding results for identical inputs to save cost\n- **Structured extraction**: Cache when extracting structured data from identical documents\n\n### Cache key composition\n\nThe cache key is derived from: model, prompt/messages, temperature, and other generation parameters. Changing any parameter produces a new cache key.\n\n## Per-User Rate Limiting\n\nControl usage at the individual user level to prevent abuse and manage costs:\n\n```ts\nconst result = await generateText({\n  model: gateway('openai/gpt-5.4'),\n  prompt: userMessage,\n  providerOptions: {\n    gateway: {\n      user: userId, // Required for per-user rate limiting\n      tags: ['feature:chat'],\n    },\n  },\n})\n```\n\n### Rate limit configuration\n\nConfigure rate limits at `https://vercel.com/{team}/{project}/settings` → **AI Gateway** → **Rate Limits**:\n\n- **Requests per minute per user**: Throttle individual users (e.g., 20 RPM)\n- **Tokens per day per user**: Cap daily token consumption (e.g., 100K tokens/day)\n- **Concurrent requests per user**: Limit parallel calls (e.g., 3 concurrent)\n\n### Handling rate limit responses\n\nWhen a user exceeds their limit, the gateway returns HTTP 429:\n\n```ts\nimport { generateText, APICallError } from 'ai'\n\ntry {\n  const result = await generateText({\n    model: gateway('openai/gpt-5.4'),\n    prompt: userMessage,\n    providerOptions: { gateway: { user: userId } },\n  })\n} catch (error) {\n  if (APICallError.isInstance(error) && error.statusCode === 429) {\n    const retryAfter = error.responseHeaders?.['retry-after']\n    return new Response(\n      JSON.stringify({ error: 'Rate limited', retryAfter }),\n      { status: 429 }\n    )\n  }\n  throw error\n}\n```\n\n## Budget Alerts and Cost Controls\n\n### Tagging for cost attribution\n\nUse tags to track spend by feature, team, and environment:\n\n```ts\nproviderOptions: {\n  gateway: {\n    tags: [\n      'feature:document-qa',\n      'team:product',\n      'env:production',\n      'tier:premium',\n    ],\n    user: userId,\n  },\n}\n```\n\n### Setting up budget alerts\n\nIn the Vercel dashboard at `https://vercel.com/{team}/{project}/settings` → **AI Gateway**:\n\n1. Navigate to **AI Gateway → Usage & Budgets**\n2. Set monthly budget thresholds (e.g., $500/month warning, $1000/month hard limit)\n3. Configure alert channels (email, Slack webhook, Vercel integration)\n4. Optionally set per-tag budgets for granular control\n\n### Budget isolation best practice\n\nUse **separate gateway keys per environment** (dev, staging, prod) and per project. This keeps dashboards clean and budgets isolated:\n\n- Restrict AI Gateway keys per project to prevent cross-tenant leakage\n- Use per-project budgets and spend-by-agent reporting to track exactly where tokens go\n- Cap spend during staging with AI Gateway budgets\n\n### Pre-flight cost controls\n\nThe AI Gateway dashboard provides observability (traces, token counts, spend tracking) but no programmatic metrics API. Build your own cost guardrails by estimating token counts and rejecting expensive requests before they execute:\n\n```ts\nimport { generateText } from 'ai'\n\nfunction estimateTokens(text: string): number {\n  return Math.ceil(text.length / 4) // rough estimate\n}\n\nasync function callWithBudget(prompt: string, maxTokens: number) {\n  const estimated = estimateTokens(prompt)\n  if (estimated > maxTokens) {\n    throw new Error(`Prompt too large: ~${estimated} tokens exceeds ${maxTokens} limit`)\n  }\n  return generateText({ model: 'openai/gpt-5.4', prompt })\n}\n```\n\nThe AI SDK's `usage` field on responses gives actual token counts after each request — store these for historical tracking and cost analysis.\n\n### Hard spending limits\n\nWhen a hard limit is reached, the gateway returns HTTP 402 (Payment Required). Handle this gracefully:\n\n```ts\nif (APICallError.isInstance(error) && error.statusCode === 402) {\n  // Budget exceeded — degrade gracefully\n  return fallbackResponse()\n}\n```\n\n### Cost optimization patterns\n\n- Use cheaper models for classification/routing, expensive models for generation\n- Cache embeddings and static queries (see Cache-Control above)\n- Set per-user daily token caps to prevent runaway usage\n- Monitor cost-per-feature with tags to identify optimization targets\n\n## Audit Logging\n\nAI Gateway logs every request for compliance and debugging:\n\n### What's logged\n\n- Timestamp, model, provider used\n- Input/output token counts\n- Latency (routing + provider)\n- User ID and tags\n- HTTP status code\n- Failover chain (which providers were tried)\n\n### Accessing logs\n\n- **Vercel Dashboard** at `https://vercel.com/{team}/{project}/ai` → **Logs** — filter by model, user, tag, status, date range\n- **Vercel API**: Query logs programmatically:\n\n```bash\ncurl -H \"Authorization: Bearer $VERCEL_TOKEN\" \\\n  \"https://api.vercel.com/v1/ai-gateway/logs?projectId=$PROJECT_ID&limit=100\"\n```\n\n- **Log Drains**: Forward AI Gateway logs to Datadog, Splunk, or other providers via Vercel Log Drains (configure at `https://vercel.com/dashboard/{team}/~/settings/log-drains`) for long-term retention and custom analysis\n\n### Compliance considerations\n\n- AI Gateway does not log prompt or completion content by default\n- Enable content logging in project settings if required for compliance\n- Logs are retained per your Vercel plan's retention policy\n- Use `user` field consistently to support audit trails\n\n## Error Handling Patterns\n\n### Provider unavailable\n\nWhen a provider is down, the gateway automatically fails over if you configured `order` or `models`:\n\n```ts\nconst result = await generateText({\n  model: gateway('anthropic/claude-sonnet-4.6'),\n  prompt: 'Summarize this document',\n  providerOptions: {\n    gateway: {\n      order: ['anthropic', 'bedrock'], // Bedrock as fallback\n      models: ['openai/gpt-5.4'],   // Final fallback model\n    },\n  },\n})\n```\n\n### Quota exceeded at provider\n\nIf your provider API key hits its quota, the gateway tries the next provider in the `order` list. Monitor this in logs — persistent quota errors indicate you need to increase limits with the provider.\n\n### Invalid model identifier\n\n```ts\n// Bad — model doesn't exist\nmodel: 'openai/gpt-99'  // Returns 400 with descriptive error\n\n// Good — use models listed in Vercel docs\nmodel: 'openai/gpt-5.4'\n```\n\n### Timeout handling\n\nGateway has a default timeout per provider. For long-running generations, use streaming:\n\n```ts\nimport { streamText } from 'ai'\n\nconst result = streamText({\n  model: 'anthropic/claude-sonnet-4.6',\n  prompt: longDocument,\n})\n\nfor await (const chunk of result.textStream) {\n  process.stdout.write(chunk)\n}\n```\n\n### Complete error handling template\n\n```ts\nimport { generateText, APICallError } from 'ai'\n\nasync function callAI(prompt: string, userId: string) {\n  try {\n    return await generateText({\n      model: gateway('openai/gpt-5.4'),\n      prompt,\n      providerOptions: {\n        gateway: {\n          user: userId,\n          order: ['openai', 'azure-openai'],\n          models: ['anthropic/claude-haiku-4.5'],\n          tags: ['feature:chat'],\n        },\n      },\n    })\n  } catch (error) {\n    if (!APICallError.isInstance(error)) throw error\n\n    switch (error.statusCode) {\n      case 402: return { text: 'Budget limit reached. Please try again later.' }\n      case 429: return { text: 'Too many requests. Please slow down.' }\n      case 503: return { text: 'AI service temporarily unavailable.' }\n      default: throw error\n    }\n  }\n}\n```\n\n## Gateway vs Direct Provider — Decision Tree\n\nUse this to decide whether to route through AI Gateway or call a provider SDK directly:\n\n```\nNeed failover across providers?\n  └─ Yes → Use Gateway\n  └─ No\n      Need cost tracking / budget alerts?\n        └─ Yes → Use Gateway\n        └─ No\n            Need per-user rate limiting?\n              └─ Yes → Use Gateway\n              └─ No\n                  Need audit logging?\n                    └─ Yes → Use Gateway\n                    └─ No\n                        Using a single provider with provider-specific features?\n                          └─ Yes → Use direct provider SDK\n                          └─ No → Use Gateway (simplifies code)\n```\n\n### When to use direct provider SDK\n\n- You need provider-specific features not exposed through the gateway (e.g., Anthropic's computer use, OpenAI's custom fine-tuned model endpoints)\n- You're self-hosting a model (e.g., vLLM, Ollama) that isn't registered with the gateway\n- You need request-level control over HTTP transport (custom proxies, mTLS)\n\n### When to always use Gateway\n\n- Production applications — failover and observability are essential\n- Multi-tenant SaaS — per-user tracking and rate limiting\n- Teams with cost accountability — tag-based budgeting\n\n## Latest Model Availability\n\n**GPT-5.4** (added March 5, 2026) — agentic and reasoning leaps from GPT-5.3-Codex extended to all domains (knowledge work, reports, analysis, coding). Faster and more token-efficient than GPT-5.2.\n\n| Model | Slug | Input | Output |\n|-------|------|-------|--------|\n| GPT-5.4 | `openai/gpt-5.4` | $2.50/M tokens | $15.00/M tokens |\n| GPT-5.4 Pro | `openai/gpt-5.4-pro` | $30.00/M tokens | $180.00/M tokens |\n\nGPT-5.4 Pro targets maximum performance on complex tasks. Use standard GPT-5.4 for most workloads.\n\n## Supported Providers\n\n- OpenAI (GPT-5.x including GPT-5.4 and GPT-5.4 Pro, o-series)\n- Anthropic (4.x models)\n- Google (Gemini)\n- xAI (Grok)\n- Mistral\n- DeepSeek\n- Amazon Bedrock\n- Azure OpenAI\n- Cohere\n- Perplexity\n- Alibaba (Qwen)\n- Meta (Llama)\n- And many more (100+ models total)\n\n## Pricing\n\n- **Zero markup**: Tokens at exact provider list price — no middleman markup, whether using Vercel-managed keys or Bring Your Own Key (BYOK)\n- **Free tier**: Every Vercel team gets **$5 of free AI Gateway credits per month** (refreshes every 30 days, starts on first request). No commitment required — experiment with LLMs indefinitely on the free tier\n- **Pay-as-you-go**: Beyond free credits, purchase AI Gateway Credits at any time with no obligation. Configure **auto top-up** to automatically add credits when your balance falls below a threshold\n- **BYOK**: Use your own provider API keys with zero fees from AI Gateway\n\n## Multimodal Support\n\nText and image generation both route through the gateway. For embeddings, use a direct provider SDK.\n\n```ts\n// Text — through gateway\nconst { text } = await generateText({\n  model: 'openai/gpt-5.4',\n  prompt: 'Hello',\n})\n\n// Image — through gateway (multimodal LLMs return images in result.files)\nconst result = await generateText({\n  model: 'google/gemini-3.1-flash-image-preview',\n  prompt: 'A sunset over the ocean',\n})\nconst images = result.files.filter((f) => f.mediaType?.startsWith('image/'))\n\n// Image-only models — through gateway with experimental_generateImage\nimport { experimental_generateImage as generateImage } from 'ai'\nconst { images: generated } = await generateImage({\n  model: 'google/imagen-4.0-generate-001',\n  prompt: 'A sunset',\n})\n```\n\n**Default image model**: `google/gemini-3.1-flash-image-preview` — fast multimodal image generation via gateway.\n\nSee [AI Gateway Image Generation docs](https://vercel.com/docs/ai-gateway/capabilities/image-generation) for all supported models and integration methods.\n\n## Key Benefits\n\n1. **Unified API**: One interface for all providers, no provider-specific code\n2. **Automatic failover**: If a provider is down, requests route to the next\n3. **Cost tracking**: Per-user, per-feature attribution with tags\n4. **Observability**: Built-in monitoring of all model calls\n5. **Low latency**: <20ms routing overhead\n6. **No lock-in**: Switch models/providers by changing a string\n\n## When to Use AI Gateway\n\n| Scenario | Use Gateway? |\n|----------|-------------|\n| Production app with AI features | Yes — failover, cost tracking |\n| Prototyping with single provider | Optional — direct provider works fine |\n| Multi-provider setup | Yes — unified routing |\n| Need provider-specific features | Use direct provider SDK + Gateway as fallback |\n| Cost tracking and budgeting | Yes — user tracking and tags |\n| Multi-tenant SaaS | Yes — per-user rate limiting and audit |\n| Compliance requirements | Yes — audit logging and log drains |\n\n## Official Documentation\n\n- [AI Gateway](https://vercel.com/docs/ai-gateway)\n- [Providers and Models](https://ai-sdk.dev/docs/foundations/providers-and-models)\n- [AI SDK Core](https://ai-sdk.dev/docs/ai-sdk-core)\n- [GitHub: AI SDK](https://github.com/vercel/ai)\n"
AFTER
"---\nname: ai-gateway\ndescription: Vercel AI Gateway guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and coding-agent configuration. Use when adding AI Gateway to an app, migrating provider calls, choosing models or providers, centralizing model configuration, evaluating application state, debugging gateway requests, or running `vercel ai-gateway` commands.\nsummary: Set up and operate Vercel AI Gateway with current models, virtual models, evaluation, authentication, routing, spend controls, and verification.\nmetadata:\n  priority: 7\n  docs:\n    - \"https://vercel.com/docs/ai-gateway\"\n    - \"https://vercel.com/docs/ai-gateway/getting-started\"\n    - \"https://vercel.com/docs/ai-gateway/models-and-providers/virtual-models\"\n    - \"https://vercel.com/docs/ai-gateway/modalities/evaluation\"\n    - \"https://vercel.com/docs/ai-gateway/sdks-and-apis/typesafe\"\n    - \"https://ai-sdk.dev/providers/ai-sdk-providers/ai-gateway\"\n  sitemap: \"https://vercel.com/docs/sitemap.md\"\n  pathPatterns: []\n  importPatterns:\n    - 'ai'\n    - '@ai-sdk/gateway'\n  bashPatterns:\n    - '\\bvercel\\s+ai-gateway\\b'\n    - '\\bvercel\\s+env\\s+pull\\b'\n    - '\\bnpm\\s+(install|i|add)\\s+[^\\n]*@ai-sdk/gateway\\b'\n    - '\\bpnpm\\s+(install|i|add)\\s+[^\\n]*@ai-sdk/gateway\\b'\n    - '\\bbun\\s+(install|i|add)\\s+[^\\n]*@ai-sdk/gateway\\b'\n    - '\\byarn\\s+add\\s+[^\\n]*@ai-sdk/gateway\\b'\n  promptSignals:\n    phrases:\n      - \"ai gateway\"\n      - \"vercel ai gateway\"\n      - \"ai-gateway\"\n      - \"ai-gateway.vercel.sh\"\n      - \"virtual model\"\n      - \"evaluation model\"\n      - \"vmc/\"\n      - \"typesafe api\"\n      - \"typesafe compat\"\n    allOf:\n      - [model, routing]\n      - [provider, failover]\n      - [gateway, budget]\n      - [gateway, logs]\n      - [gateway, oidc]\n      - [coding, gateway]\n    anyOf:\n      - \"provider ordering\"\n      - \"model fallback\"\n      - \"byok\"\n      - \"spend tracking\"\n      - \"gateway key\"\n      - \"credit balance\"\n      - \"safety identifier\"\n      - \"reasoning effort\"\n      - \"tool calling\"\n      - \"structured outputs\"\n      - \"experimental_evaluate\"\n      - \"central model configuration\"\n      - \"systemone\"\n      - \"v1/evaluate\"\n    noneOf:\n      - \"cloudflare ai gateway\"\n      - \"aws api gateway\"\n    minScore: 6\nvalidate:\n  -\n    pattern: '\\bclaude-(sonnet|opus|haiku)-\\d+-\\d+\\b'\n    message: 'Claude model version uses a hyphen where the AI Gateway slug uses a dot. Fetch /v1/models and use the returned provider/model ID.'\n    severity: error\n  -\n    pattern: gateway\\(['\"][^'\"/]+['\"]\\)\n    message: 'AI Gateway model string is missing its provider prefix. Fetch /v1/models and use a provider/model ID.'\n    severity: error\n  -\n    pattern: (OPENAI_API_KEY|ANTHROPIC_API_KEY|GOOGLE_API_KEY)\n    message: 'Provider key detected. AI Gateway request authentication uses AI_GATEWAY_API_KEY or VERCEL_OIDC_TOKEN; provider keys belong only in an intentional BYOK configuration.'\n    severity: recommended\n    skipIfFileContains: '[Bb][Yy][Oo][Kk]|providerOptions\\s*:\\s*\\{[^}]*gateway'\n  -\n    pattern: gateway\\s*:\\s*\\{[^}]*cacheControl\n    message: \"AI Gateway does not cache whole responses through cacheControl. Use caching: 'auto' for provider prompt caching and verify the current caching docs.\"\n    severity: error\n  -\n    pattern: ANTHROPIC_BASE_URL\\s*=\\s*[\"']?https://ai-gateway\\.vercel\\.sh\n    message: 'Claude Code through AI Gateway needs ANTHROPIC_API_KEY set to an empty value and the gateway key in ANTHROPIC_AUTH_TOKEN. A non-empty ANTHROPIC_API_KEY is used instead of the gateway token.'\n    severity: recommended\n    skipIfFileContains: 'ANTHROPIC_AUTH_TOKEN'\nchainTo:\n  -\n    pattern: 'from\\s+[''\"]ai[''\"]|require\\([''\"]ai[''\"]\\)|\\b(generateText|streamText|ToolLoopAgent)\\b'\n    targetSkill: ai-sdk\n    message: 'AI SDK code detected. Load the AI SDK skill and read the installed package docs before writing or changing SDK code.'\nretrieval:\n  aliases:\n    - model router\n    - ai proxy\n    - provider failover\n    - llm gateway\n    - gateway credits\n    - virtual model\n    - evaluation model\n    - typesafe compatibility\n  intents:\n    - add Vercel AI Gateway to an application\n    - route AI models across providers\n    - configure provider or model fallbacks\n    - authenticate AI Gateway requests\n    - track AI model costs and set budgets\n    - debug AI Gateway requests and routing\n    - connect coding agents to AI Gateway\n    - give coding agents centrally managed model and provider configuration\n    - create or update an AI Gateway virtual model\n    - evaluate application state with typed questions\n    - migrate an existing TypeSafe evaluation client to AI Gateway\n    - find a model by modality, capability, price, or data retention\n    - check AI Gateway credit balance or generation cost\n    - configure reasoning or extended thinking across providers and API formats\n    - add tool calling or function calling across API formats\n    - get structured JSON output matching a schema\n    - send images or PDFs to a model\n  entities:\n    - AI Gateway\n    - AI Gateway Credits\n    - providerOptions.gateway\n    - AI_GATEWAY_API_KEY\n    - VERCEL_OIDC_TOKEN\n    - model routing\n    - provider failover\n    - BYOK\n    - spend reporting\n    - safetyIdentifier\n    - Usage & Billing API\n    - Virtual Models\n    - vmc/<slug>\n    - experimental_evaluate\n    - POST /v1/evaluate\n    - TypeSafe API\n    - /typesafe/v1/systemone\n---\n\n# Vercel AI Gateway\n\nAI Gateway exposes models from multiple providers through shared authentication, model IDs, routing, billing, and observability. Model availability, SDK APIs, CLI commands, prices, and product capabilities change frequently. Verify them from current sources before changing code.\n\n## Start with current sources\n\nBefore implementing:\n\n1. Inspect the project's language, package manager, installed AI SDK version, and existing provider integration.\n2. Read the relevant Vercel page under <https://vercel.com/docs/ai-gateway>. Use the page's `.md` form when a tool needs Markdown.\n3. Fetch the complete live model list. Do not construct model variants by analogy:\n\n   ```bash\n   curl -fsSL https://ai-gateway.vercel.sh/v1/models\n   ```\n\n4. If the code uses the `ai` package, load the `ai-sdk` skill when available. Read version-matched docs under `node_modules/ai/docs/` and source under `node_modules/ai/src/`. If the skill is not installed, use those bundled files directly.\n5. Run `vercel ai-gateway <command> --help` before documenting or scripting CLI flags.\n\nThe live model endpoint and installed package take precedence over model names or SDK syntax remembered from training data.\n\n## Vercel CLI inventory\n\nThe `vercel ai-gateway` command manages gateway resources for the current team. The `setup` subcommand connects local coding agents; the rest of the CLI covers the jobs that previously required dashboard work:\n\n| Command | What it does |\n| --- | --- |\n| `api-keys create/list/inspect/remove` | Create and manage AI Gateway API keys, with budgets, spend alerts, expiry, and restriction exemptions |\n| `budgets set/list/inspect/remove` | Set metered spend limits for the team, a project, a user, or an API key |\n| `budgets defaults set/list/remove` | Set per-scope default limits covering projects, keys, or members without a custom budget |\n| `models list` / `models endpoints <model>` | List the model catalog and one model's provider endpoints from the CLI |\n| `virtual-models create/list/inspect/edit/remove/restore` | Manage reusable, team-scoped model configurations addressed as `vmc/<slug>` |\n| `rules add/list/edit/remove` | Manage routing rules; the CLI marks rules beta, so check `--help` before relying on them. REST CRUD exists under `/v1/ai-gateway/rules` |\n| `setup` | Configure supported coding agents; see [references/coding-agents.md](references/coding-agents.md) |\n| `leaderboard` | Explore public, anonymized usage leaderboards; rarely needed for implementation work |\n\nUse the CLI for credential and spend management when the user is working from a terminal or in CI. Check `vercel ai-gateway <command> --help` for current flags before scripting; do not copy a flag list from this skill into generated code.\n\n## Route the request to the right guide\n\n| User's job | Read |\n| --- | --- |\n| Ask a coding agent to make one Gateway request, or handle first-request credentials, compatible SDKs, or migration | [references/setup.md](references/setup.md) |\n| Provider selection, model fallbacks, caching, BYOK, or timeouts | [references/routing.md](references/routing.md) |\n| Reusable model configuration, a `vmc/<slug>`, or provider options for a client that cannot send them | [references/virtual-models.md](references/virtual-models.md) |\n| Typed evaluation through AI SDK, `POST /v1/evaluate`, or the TypeSafe-compatible API | [references/evaluation.md](references/evaluation.md) |\n| Credits, budgets, reporting, Logs, or request debugging | [references/spend-observability.md](references/spend-observability.md) |\n| Route Claude Code, Codex, OpenCode, Pi, or another coding agent's own model traffic through Gateway | [references/coding-agents.md](references/coding-agents.md) |\n\nRead each relevant reference before editing. A task can require more than one.\n\n## Choose the integration surface\n\n| Existing project | Default path |\n| --- | --- |\n| JavaScript or TypeScript using AI SDK | Use a plain `provider/model` string with `generateText`, `streamText`, `ToolLoopAgent`, or the relevant modality API |\n| Python using AI SDK for Python | Use `ai.get_model('provider/model')` and the current Python SDK docs |\n| Existing OpenAI SDK | Keep the SDK and point `baseURL` or `base_url` to `https://ai-gateway.vercel.sh/v1` |\n| Existing Anthropic SDK | Keep the SDK and point `baseURL` or `base_url` to `https://ai-gateway.vercel.sh` |\n| Evaluation over provider-neutral HTTP | Send Gateway's evaluation request shape to `POST https://ai-gateway.vercel.sh/v1/evaluate` |\n| Existing TypeSafe evaluation client | Keep `@typesafe-ai/sdk` and point `baseURL` to `https://ai-gateway.vercel.sh/typesafe` |\n| Provider-neutral HTTP | Use an AI Gateway compatible endpoint, such as Chat Completions or OpenResponses |\n| Existing direct-provider AI SDK integration | Replace the provider instance with a live AI Gateway `provider/model` string, then remove provider credentials only after verifying the gateway path |\n| Coding agent | Use `vercel ai-gateway setup`; inspect its help before claiming agent support. Use a Virtual Model when the agent needs reusable routing or provider options it cannot send per request |\n\nAI Gateway also supports OpenAI Responses, Anthropic Messages, OpenResponses, Cohere Rerank, embeddings, image and video generation, speech, transcription, realtime sessions, and evaluation. Modality pages under <https://vercel.com/docs/ai-gateway/modalities> cover each request shape, including background jobs for long-running video generation. Evaluation is available through AI SDK 7 or later, `POST /v1/evaluate`, and a TypeSafe-compatible API under `/typesafe`; it is not available through the OpenAI-, Anthropic-, or Cohere-compatible endpoints. Read [references/evaluation.md](references/evaluation.md) before choosing a surface. Read the relevant modality or API page instead of translating one request shape from memory.\n\n## Minimal AI SDK request\n\nThe current AI SDK requires Node.js 22 or later. Confirm the installed package's `engines` field before enforcing a version in an existing project.\n\n```ts\nimport { generateText } from 'ai';\n\nconst model = process.env.AI_GATEWAY_MODEL;\nif (!model) {\n  throw new Error('Set AI_GATEWAY_MODEL to an ID returned by /v1/models');\n}\n\nconst { text } = await generateText({\n  model,\n  prompt: 'Explain the project in one paragraph.',\n});\n\nconsole.log(text);\n```\n\nHonor an exact model the user or task specifies after confirming it exists. Otherwise fetch `/v1/models`, choose a model that fits the requested modality, capabilities, price, context window, data-retention policy, and team access, and set `AI_GATEWAY_MODEL` to that ID. Do not put a time-sensitive model recommendation in reusable examples.\n\nPlain model strings route through AI Gateway. Add `@ai-sdk/gateway` only when the task needs its exported provider, types, model discovery, generation lookup, or spend-report helpers.\n\n## Authentication decision\n\n- Use an **AI Gateway API key** for local scripts, CI, external servers, and non-Vercel deployments. Store it in `AI_GATEWAY_API_KEY` and never print or commit it.\n- Use **Vercel OIDC** for Vercel deployments and linked local projects. Vercel deployments receive `VERCEL_OIDC_TOKEN`; local development uses `vercel link` and `vercel env pull`.\n- **BYOK provider credentials do not replace AI Gateway request authentication.** They decide how AI Gateway authenticates to a model provider.\n- A plain Node.js script does not automatically load `.env.local`. Export variables in the shell or load that file explicitly. Framework behavior may differ.\n\nDo not ask the user to paste a secret into chat, source code, a committed config file, or a command that will enter shell history unless the repository has an established secure mechanism.\n\n## Implementation workflow\n\n1. Establish the user's job, runtime, deployment target, current provider, and required capabilities.\n2. Select authentication from the rules above. Preserve a working existing method unless the user asked to migrate it.\n3. Fetch live model metadata and choose a compatible model. State why it fits.\n4. Read the matching SDK, API, modality, routing, or coding-agent docs.\n5. Make the smallest end-to-end change. Reuse the current project structure and error handling.\n6. Handle only errors the application can act on. Common gateway outcomes include authentication failure, insufficient credits, budget exhaustion, rate limiting, and provider capacity failure.\n7. Run the project's formatter, type checker, and focused tests.\n8. When the task authorizes a live request, run one and inspect the returned model, text or media, usage, and provider metadata.\n9. Verify the request in AI Gateway Logs when dashboard access is available. Logs can take about 90 seconds to ingest.\n\nOnly spend credits, create keys, change budgets, change routing rules, or write coding-agent config when the user requested or approved that outward-facing action. Prefer dry runs and interactive previews when available.\n\n## Routing invariants\n\n- Model IDs use the exact `provider/model` strings returned by `/v1/models`.\n- `order` controls provider preference, `only` restricts providers, and `sort` ranks providers by a supported metric.\n- `models` lists fallback models after the primary model.\n- `caching: 'auto'` manages provider prompt-cache markers. It is not an HTTP response cache.\n- `providerTimeouts` applies to BYOK provider attempts and measures time until the provider starts responding.\n- A reasoning entry in `providerOptions` overrides the AI SDK top-level `reasoning` value entirely; the two never merge.\n- `user` and `tags` attach reporting dimensions. They do not create per-user rate limits.\n- Request-scoped provider credentials belong under `providerOptions.gateway.byok` and must remain secret.\n- Virtual Model IDs use `vmc/<slug>`. A Virtual Model can pin routing and provider options server-side; settings it defines generally override the corresponding request settings, while unset settings remain request-configurable.\n\nRead [references/routing.md](references/routing.md) before adding any of these fields.\n\n## Verification checklist\n\n- [ ] A direct model ID exists in the full live model response, or a Virtual Model exists for the authenticated team and resolves as `vmc/<slug>`.\n- [ ] The selected API or SDK supports the requested modality and feature.\n- [ ] Authentication works in the actual runtime, including `.env.local` loading where relevant.\n- [ ] The example prints or returns a result instead of discarding the response.\n- [ ] The project type checker and focused tests pass.\n- [ ] A live request was made only when authorized, and its cost was understood.\n- [ ] Provider routing or fallback behavior is visible in response metadata or Logs.\n- [ ] No secret appears in source, logs, diffs, or the final response.\n- [ ] The final report distinguishes code verification from live and dashboard verification.\n\n## Current documentation\n\n- Getting started: <https://vercel.com/docs/ai-gateway/getting-started>\n- Models and providers: <https://vercel.com/docs/ai-gateway/models-and-providers>\n- Virtual Models: <https://vercel.com/docs/ai-gateway/models-and-providers/virtual-models>\n- SDKs and APIs: <https://vercel.com/docs/ai-gateway/sdks-and-apis>\n- Authentication and BYOK: <https://vercel.com/docs/ai-gateway/authentication-and-byok>\n- Observability and spend: <https://vercel.com/docs/ai-gateway/observability-and-spend>\n- Modalities: <https://vercel.com/docs/ai-gateway/modalities>\n- Evaluation: <https://vercel.com/docs/ai-gateway/modalities/evaluation>\n- TypeSafe API: <https://vercel.com/docs/ai-gateway/sdks-and-apis/typesafe>\n- REST API reference: <https://vercel.com/docs/ai-gateway/sdks-and-apis/rest-api>\n- FAQ: <https://vercel.com/docs/ai-gateway/faq>\n- Coding agents: <https://vercel.com/docs/ai-gateway/coding-agents>\n- AI SDK provider: <https://ai-sdk.dev/providers/ai-sdk-providers/ai-gateway>\n"

SKILL.md line diff

--- before
+++ after
@@ -1,562 +1,293 @@
 ---
 name: ai-gateway
-description: Vercel AI Gateway expert guidance. Use when configuring model routing, provider failover, cost tracking, or managing multiple AI providers through a unified API.
+description: Vercel AI Gateway guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and coding-agent configuration. Use when adding AI Gateway to an app, migrating provider calls, choosing models or providers, centralizing model configuration, evaluating application state, debugging gateway requests, or running `vercel ai-gateway` commands.
+summary: Set up and operate Vercel AI Gateway with current models, virtual models, evaluation, authentication, routing, spend controls, and verification.
 metadata:
   priority: 7
   docs:
     - "https://vercel.com/docs/ai-gateway"
-    - "https://sdk.vercel.ai/docs/ai-sdk-core/settings"
-  sitemap: "https://vercel.com/sitemap/docs.xml"
+    - "https://vercel.com/docs/ai-gateway/getting-started"
+    - "https://vercel.com/docs/ai-gateway/models-and-providers/virtual-models"
+    - "https://vercel.com/docs/ai-gateway/modalities/evaluation"
+    - "https://vercel.com/docs/ai-gateway/sdks-and-apis/typesafe"
+    - "https://ai-sdk.dev/providers/ai-sdk-providers/ai-gateway"
+  sitemap: "https://vercel.com/docs/sitemap.md"
   pathPatterns: []
   importPatterns:
     - 'ai'
     - '@ai-sdk/gateway'
   bashPatterns:
+    - '\bvercel\s+ai-gateway\b'
     - '\bvercel\s+env\s+pull\b'
     - '\bnpm\s+(install|i|add)\s+[^\n]*@ai-sdk/gateway\b'
     - '\bpnpm\s+(install|i|add)\s+[^\n]*@ai-sdk/gateway\b'
     - '\bbun\s+(install|i|add)\s+[^\n]*@ai-sdk/gateway\b'
     - '\byarn\s+add\s+[^\n]*@ai-sdk/gateway\b'
+  promptSignals:
+    phrases:
+      - "ai gateway"
+      - "vercel ai gateway"
+      - "ai-gateway"
+      - "ai-gateway.vercel.sh"
+      - "virtual model"
+      - "evaluation model"
+      - "vmc/"
+      - "typesafe api"
+      - "typesafe compat"
+    allOf:
+      - [model, routing]
+      - [provider, failover]
+      - [gateway, budget]
+      - [gateway, logs]
+      - [gateway, oidc]
+      - [coding, gateway]
+    anyOf:
+      - "provider ordering"
+      - "model fallback"
+      - "byok"
+      - "spend tracking"
+      - "gateway key"
+      - "credit balance"
+      - "safety identifier"
+      - "reasoning effort"
+      - "tool calling"
+      - "structured outputs"
+      - "experimental_evaluate"
+      - "central model configuration"
+      - "systemone"
+      - "v1/evaluate"
+    noneOf:
+      - "cloudflare ai gateway"
+      - "aws api gateway"
+    minScore: 6
+validate:
+  -
+    pattern: '\bclaude-(sonnet|opus|haiku)-\d+-\d+\b'
+    message: 'Claude model version uses a hyphen where the AI Gateway slug uses a dot. Fetch /v1/models and use the returned provider/model ID.'
+    severity: error
+  -
+    pattern: gateway\(['"][^'"/]+['"]\)
+    message: 'AI Gateway model string is missing its provider prefix. Fetch /v1/models and use a provider/model ID.'
+    severity: error
+  -
+    pattern: (OPENAI_API_KEY|ANTHROPIC_API_KEY|GOOGLE_API_KEY)
+    message: 'Provider key detected. AI Gateway request authentication uses AI_GATEWAY_API_KEY or VERCEL_OIDC_TOKEN; provider keys belong only in an intentional BYOK configuration.'
+    severity: recommended
+    skipIfFileContains: '[Bb][Yy][Oo][Kk]|providerOptions\s*:\s*\{[^}]*gateway'
+  -
+    pattern: gateway\s*:\s*\{[^}]*cacheControl
+    message: "AI Gateway does not cache whole responses through cacheControl. Use caching: 'auto' for provider prompt caching and verify the current caching docs."
+    severity: error
+  -
+    pattern: ANTHROPIC_BASE_URL\s*=\s*["']?https://ai-gateway\.vercel\.sh
+    message: 'Claude Code through AI Gateway needs ANTHROPIC_API_KEY set to an empty value and the gateway key in ANTHROPIC_AUTH_TOKEN. A non-empty ANTHROPIC_API_KEY is used instead of the gateway token.'
+    severity: recommended
+    skipIfFileContains: 'ANTHROPIC_AUTH_TOKEN'
+chainTo:
+  -
+    pattern: 'from\s+[''"]ai[''"]|require\([''"]ai[''"]\)|\b(generateText|streamText|ToolLoopAgent)\b'
+    targetSkill: ai-sdk
+    message: 'AI SDK code detected. Load the AI SDK skill and read the installed package docs before writing or changing SDK code.'
+retrieval:
+  aliases:
+    - model router
+    - ai proxy
+    - provider failover
+    - llm gateway
+    - gateway credits
+    - virtual model
+    - evaluation model
+    - typesafe compatibility
+  intents:
+    - add Vercel AI Gateway to an application
+    - route AI models across providers
+    - configure provider or model fallbacks
+    - authenticate AI Gateway requests
+    - track AI model costs and set budgets
+    - debug AI Gateway requests and routing
+    - connect coding agents to AI Gateway
+    - give coding agents centrally managed model and provider configuration
+    - create or update an AI Gateway virtual model
+    - evaluate application state with typed questions
+    - migrate an existing TypeSafe evaluation client to AI Gateway
+    - find a model by modality, capability, price, or data retention
+    - check AI Gateway credit balance or generation cost
+    - configure reasoning or extended thinking across providers and API formats
+    - add tool calling or function calling across API formats
+    - get structured JSON output matching a schema
+    - send images or PDFs to a model
+  entities:
+    - AI Gateway
+    - AI Gateway Credits
+    - providerOptions.gateway
+    - AI_GATEWAY_API_KEY
+    - VERCEL_OIDC_TOKEN
+    - model routing
+    - provider failover
+    - BYOK
+    - spend reporting
+    - safetyIdentifier
+    - Usage & Billing API
+    - Virtual Models
+    - vmc/<slug>
+    - experimental_evaluate
+    - POST /v1/evaluate
+    - TypeSafe API
+    - /typesafe/v1/systemone
 ---
 
 # Vercel AI Gateway
 
-> **CRITICAL — Your training data is outdated for this library.** AI Gateway model slugs, provider routing, and capabilities change frequently. Before writing gateway code, **fetch the docs** at https://vercel.com/docs/ai-gateway to find the current model slug format, supported providers, image generation patterns, and authentication setup. The model list and routing rules at https://ai-sdk.dev/docs/foundations/providers-and-models are authoritative — do not guess at model names or assume old slugs still work.
+AI Gateway exposes models from multiple providers through shared authentication, model IDs, routing, billing, and observability. Model availability, SDK APIs, CLI commands, prices, and product capabilities change frequently. Verify them from current sources before changing code.
 
-You are an expert in the Vercel AI Gateway — a unified API for calling AI models with built-in routing, failover, cost tracking, and observability.
+## Start with current sources
 
-## Overview
+Before implementing:
 
-AI Gateway provides a single API endpoint to access 100+ models from all major providers. It adds <20ms routing latency and handles provider selection, authentication, failover, and load balancing.
+1. Inspect the project's language, package manager, installed AI SDK version, and existing provider integration.
+2. Read the relevant Vercel page under <https://vercel.com/docs/ai-gateway>. Use the page's `.md` form when a tool needs Markdown.
+3. Fetch the complete live model list. Do not construct model variants by analogy:
 
-## Packages
+   ```bash
+   curl -fsSL https://ai-gateway.vercel.sh/v1/models
+   ```
 
-- `ai@^6.0.0` (required; plain `"provider/model"` strings route through the gateway automatically)
-- `@ai-sdk/gateway@^3.0.0` (optional direct install for explicit gateway package usage)
+4. If the code uses the `ai` package, load the `ai-sdk` skill when available. Read version-matched docs under `node_modules/ai/docs/` and source under `node_modules/ai/src/`. If the skill is not installed, use those bundled files directly.
+5. Run `vercel ai-gateway <command> --help` before documenting or scripting CLI flags.
 
-## Setup
+The live model endpoint and installed package take precedence over model names or SDK syntax remembered from training data.
 
-Pass a `"provider/model"` string to the `model` parameter — the AI SDK automatically routes it through the AI Gateway:
+## Vercel CLI inventory
 
-```ts
-import { generateText } from 'ai'
-
-const result = await generateText({
-  model: 'openai/gpt-5.4', // plain string — routes through AI Gateway automatically
-  prompt: 'Hello!',
-})
-```
-
-No `gateway()` wrapper or additional package needed. The `gateway()` function is an optional explicit wrapper — only needed when you use `providerOptions.gateway` for routing, failover, or tags:
-
-```ts
-import { gateway } from 'ai'
-
-const result = await generateText({
-  model: gateway('openai/gpt-5.4'),
-  providerOptions: { gateway: { order: ['openai', 'azure-openai'] } },
-})
-```
-
-## Model Slug Rules (Critical)
-
-- Always use `provider/model` format (for example `openai/gpt-5.4`).
-- Versioned slugs use dots for versions, not hyphens:
-  - Correct: `anthropic/claude-sonnet-4.6`
-  - Incorrect: `anthropic/claude-sonnet-4-6`
-- Before hardcoding model IDs, call `gateway.getAvailableModels()` and pick from the returned IDs.
-- Default text models: `openai/gpt-5.4` or `anthropic/claude-sonnet-4.6`.
-- Do not default to outdated choices like `openai/gpt-4o`.
-
-```ts
-import { gateway } from 'ai'
-
-const availableModels = await gateway.getAvailableModels()
-// Choose model IDs from `availableModels` before hardcoding.
-```
-
-## Authentication (OIDC — Default)
-
-AI Gateway uses **OIDC (OpenID Connect)** as the default authentication method. No manual API keys needed.
-
-### Setup
-
-```bash
-vercel link                    # Connect to your Vercel project
-# Enable AI Gateway in Vercel dashboard: https://vercel.com/{team}/{project}/settings → AI Gateway
-vercel env pull .env.local     # Provisions VERCEL_OIDC_TOKEN automatically
-```
-
-### How It Works
-
-1. `vercel env pull` writes a `VERCEL_OIDC_TOKEN` to `.env.local` — a short-lived JWT (~24h)
-2. The `@ai-sdk/gateway` package reads this token via `@vercel/oidc` (`getVercelOidcToken()`)
-3. No `AI_GATEWAY_API_KEY` or provider-specific keys (like `ANTHROPIC_API_KEY`) are needed
-4. On Vercel deployments, OIDC tokens are auto-refreshed — zero maintenance
-
-### Local Development
-
-For local dev, the OIDC token from `vercel env pull` is valid for ~24 hours. When it expires:
-
-```bash
-vercel env pull .env.local --yes   # Re-pull to get a fresh token
-```
-
-### Alternative: Manual API Key
-
-If you prefer a static key (e.g., for CI or non-Vercel environments):
-
-```bash
-# Set AI_GATEWAY_API_KEY in your environment
-# The gateway falls back to this when VERCEL_OIDC_TOKEN is not available
-export AI_GATEWAY_API_KEY=your-key-here
-```
-
-### Auth Priority
-
-The `@ai-sdk/gateway` package resolves authentication in this order:
-1. `AI_GATEWAY_API_KEY` environment variable (if set)
-2. `VERCEL_OIDC_TOKEN` via `@vercel/oidc` (default on Vercel and after `vercel env pull`)
-
-## Provider Routing
-
-Configure how AI Gateway routes requests across providers:
-
-```ts
-const result = await generateText({
-  model: gateway('anthropic/claude-sonnet-4.6'),
-  prompt: 'Hello!',
-  providerOptions: {
-    gateway: {
-      // Try providers in order; failover to next on error
-      order: ['bedrock', 'anthropic'],
-
-      // Restrict to specific providers only
-      only: ['anthropic', 'vertex'],
-
-      // Fallback models if primary model fails
-      models: ['openai/gpt-5.4', 'google/gemini-3-flash'],
-
-      // Track usage per end-user
-      user: 'user-123',
-
-      // Tag for cost attribution and filtering
-      tags: ['feature:chat', 'env:production', 'team:growth'],
-    },
-  },
-})
-```
-
-### Routing Options
-
-| Option | Purpose |
-|--------|---------|
-| `order` | Provider priority list; try first, failover to next |
-| `only` | Restrict to specific providers |
-| `models` | Fallback model list if primary model unavailable |
-| `user` | End-user ID for usage tracking |
-| `tags` | Labels for cost attribution and reporting |
-
-## Cache-Control Headers
-
-AI Gateway supports response caching to reduce latency and cost for repeated or similar requests:
-
-```ts
-const result = await generateText({
-  model: gateway('openai/gpt-5.4'),
-  prompt: 'What is the capital of France?',
-  providerOptions: {
-    gateway: {
-      // Cache identical requests for 1 hour
-      cacheControl: 'max-age=3600',
-    },
-  },
-})
-```
-
-### Caching strategies
-
-| Header Value | Behavior |
-|-------------|----------|
-| `max-age=3600` | Cache response for 1 hour |
-| `max-age=0` | Bypass cache, always call provider |
-| `s-maxage=86400` | Cache at the edge for 24 hours |
-| `stale-while-revalidate=600` | Serve stale for 10 min while refreshing in background |
-
-### When to use caching
-
-- **Static knowledge queries**: FAQs, translations, factual lookups — cache aggressively
-- **User-specific conversations**: Do not cache — each response depends on conversation history
-- **Embeddings**: Cache embedding results for identical inputs to save cost
-- **Structured extraction**: Cache when extracting structured data from identical documents
-
-### Cache key composition
-
-The cache key is derived from: model, prompt/messages, temperature, and other generation parameters. Changing any parameter produces a new cache key.
-
-## Per-User Rate Limiting
-
-Control usage at the individual user level to prevent abuse and manage costs:
-
-```ts
-const result = await generateText({
-  model: gateway('openai/gpt-5.4'),
-  prompt: userMessage,
-  providerOptions: {
-    gateway: {
-      user: userId, // Required for per-user rate limiting
-      tags: ['feature:chat'],
-    },
-  },
-})
-```
-
-### Rate limit configuration
-
-Configure rate limits at `https://vercel.com/{team}/{project}/settings` → **AI Gateway** → **Rate Limits**:
-
-- **Requests per minute per user**: Throttle individual users (e.g., 20 RPM)
-- **Tokens per day per user**: Cap daily token consumption (e.g., 100K tokens/day)
-- **Concurrent requests per user**: Limit parallel calls (e.g., 3 concurrent)
-
-### Handling rate limit responses
-
-When a user exceeds their limit, the gateway returns HTTP 429:
-
-```ts
-import { generateText, APICallError } from 'ai'
-
-try {
-  const result = await generateText({
-    model: gateway('openai/gpt-5.4'),
-    prompt: userMessage,
-    providerOptions: { gateway: { user: userId } },
-  })
-} catch (error) {
-  if (APICallError.isInstance(error) && error.statusCode === 429) {
-    const retryAfter = error.responseHeaders?.['retry-after']
-    return new Response(
-      JSON.stringify({ error: 'Rate limited', retryAfter }),
-      { status: 429 }
-    )
-  }
-  throw error
-}
-```
-
-## Budget Alerts and Cost Controls
-
-### Tagging for cost attribution
-
-Use tags to track spend by feature, team, and environment:
-
-```ts
-providerOptions: {
-  gateway: {
-    tags: [
-      'feature:document-qa',
-      'team:product',
-      'env:production',
-      'tier:premium',
-    ],
-    user: userId,
-  },
-}
-```
-
-### Setting up budget alerts
+The `vercel ai-gateway` command manages gateway resources for the current team. The `setup` subcommand connects local coding agents; the rest of the CLI covers the jobs that previously required dashboard work:
 
-In the Vercel dashboard at `https://vercel.com/{team}/{project}/settings` → **AI Gateway**:
+| Command | What it does |
+| --- | --- |
+| `api-keys create/list/inspect/remove` | Create and manage AI Gateway API keys, with budgets, spend alerts, expiry, and restriction exemptions |
+| `budgets set/list/inspect/remove` | Set metered spend limits for the team, a project, a user, or an API key |
+| `budgets defaults set/list/remove` | Set per-scope default limits covering projects, keys, or members without a custom budget |
+| `models list` / `models endpoints <model>` | List the model catalog and one model's provider endpoints from the CLI |
+| `virtual-models create/list/inspect/edit/remove/restore` | Manage reusable, team-scoped model configurations addressed as `vmc/<slug>` |
+| `rules add/list/edit/remove` | Manage routing rules; the CLI marks rules beta, so check `--help` before relying on them. REST CRUD exists under `/v1/ai-gateway/rules` |
+| `setup` | Configure supported coding agents; see [references/coding-agents.md](references/coding-agents.md) |
+| `leaderboard` | Explore public, anonymized usage leaderboards; rarely needed for implementation work |
 
-1. Navigate to **AI Gateway → Usage & Budgets**
-2. Set monthly budget thresholds (e.g., $500/month warning, $1000/month hard limit)
-3. Configure alert channels (email, Slack webhook, Vercel integration)
-4. Optionally set per-tag budgets for granular control
+Use the CLI for credential and spend management when the user is working from a terminal or in CI. Check `vercel ai-gateway <command> --help` for current flags before scripting; do not copy a flag list from this skill into generated code.
 
-### Budget isolation best practice
+## Route the request to the right guide
 
-Use **separate gateway keys per environment** (dev, staging, prod) and per project. This keeps dashboards clean and budgets isolated:
+| User's job | Read |
+| --- | --- |
+| Ask a coding agent to make one Gateway request, or handle first-request credentials, compatible SDKs, or migration | [references/setup.md](references/setup.md) |
+| Provider selection, model fallbacks, caching, BYOK, or timeouts | [references/routing.md](references/routing.md) |
+| Reusable model configuration, a `vmc/<slug>`, or provider options for a client that cannot send them | [references/virtual-models.md](references/virtual-models.md) |
+| Typed evaluation through AI SDK, `POST /v1/evaluate`, or the TypeSafe-compatible API | [references/evaluation.md](references/evaluation.md) |
+| Credits, budgets, reporting, Logs, or request debugging | [references/spend-observability.md](references/spend-observability.md) |
+| Route Claude Code, Codex, OpenCode, Pi, or another coding agent's own model traffic through Gateway | [references/coding-agents.md](references/coding-agents.md) |
 
-- Restrict AI Gateway keys per project to prevent cross-tenant leakage
-- Use per-project budgets and spend-by-agent reporting to track exactly where tokens go
-- Cap spend during staging with AI Gateway budgets
+Read each relevant reference before editing. A task can require more than one.
 
-### Pre-flight cost controls
+## Choose the integration surface
 
-The AI Gateway dashboard provides observability (traces, token counts, spend tracking) but no programmatic metrics API. Build your own cost guardrails by estimating token counts and rejecting expensive requests before they execute:
+| Existing project | Default path |
+| --- | --- |
+| JavaScript or TypeScript using AI SDK | Use a plain `provider/model` string with `generateText`, `streamText`, `ToolLoopAgent`, or the relevant modality API |
+| Python using AI SDK for Python | Use `ai.get_model('provider/model')` and the current Python SDK docs |
+| Existing OpenAI SDK | Keep the SDK and point `baseURL` or `base_url` to `https://ai-gateway.vercel.sh/v1` |
+| Existing Anthropic SDK | Keep the SDK and point `baseURL` or `base_url` to `https://ai-gateway.vercel.sh` |
+| Evaluation over provider-neutral HTTP | Send Gateway's evaluation request shape to `POST https://ai-gateway.vercel.sh/v1/evaluate` |
+| Existing TypeSafe evaluation client | Keep `@typesafe-ai/sdk` and point `baseURL` to `https://ai-gateway.vercel.sh/typesafe` |
+| Provider-neutral HTTP | Use an AI Gateway compatible endpoint, such as Chat Completions or OpenResponses |
+| Existing direct-provider AI SDK integration | Replace the provider instance with a live AI Gateway `provider/model` string, then remove provider credentials only after verifying the gateway path |
+| Coding agent | Use `vercel ai-gateway setup`; inspect its help before claiming agent support. Use a Virtual Model when the agent needs reusable routing or provider options it cannot send per request |
 
-```ts
-import { generateText } from 'ai'
-
-function estimateTokens(text: string): number {
-  return Math.ceil(text.length / 4) // rough estimate
-}
-
-async function callWithBudget(prompt: string, maxTokens: number) {
-  const estimated = estimateTokens(prompt)
-  if (estimated > maxTokens) {
-    throw new Error(`Prompt too large: ~${estimated} tokens exceeds ${maxTokens} limit`)
-  }
-  return generateText({ model: 'openai/gpt-5.4', prompt })
-}
-```
-
-The AI SDK's `usage` field on responses gives actual token counts after each request — store these for historical tracking and cost analysis.
-
-### Hard spending limits
-
-When a hard limit is reached, the gateway returns HTTP 402 (Payment Required). Handle this gracefully:
-
-```ts
-if (APICallError.isInstance(error) && error.statusCode === 402) {
-  // Budget exceeded — degrade gracefully
-  return fallbackResponse()
-}
-```
-
-### Cost optimization patterns
-
-- Use cheaper models for classification/routing, expensive models for generation
-- Cache embeddings and static queries (see Cache-Control above)
-- Set per-user daily token caps to prevent runaway usage
-- Monitor cost-per-feature with tags to identify optimization targets
-
-## Audit Logging
-
-AI Gateway logs every request for compliance and debugging:
-
-### What's logged
-
-- Timestamp, model, provider used
-- Input/output token counts
-- Latency (routing + provider)
-- User ID and tags
-- HTTP status code
-- Failover chain (which providers were tried)
-
-### Accessing logs
-
-- **Vercel Dashboard** at `https://vercel.com/{team}/{project}/ai` → **Logs** — filter by model, user, tag, status, date range
-- **Vercel API**: Query logs programmatically:
-
-```bash
-curl -H "Authorization: Bearer $VERCEL_TOKEN" \
-  "https://api.vercel.com/v1/ai-gateway/logs?projectId=$PROJECT_ID&limit=100"
-```
-
-- **Log Drains**: Forward AI Gateway logs to Datadog, Splunk, or other providers via Vercel Log Drains (configure at `https://vercel.com/dashboard/{team}/~/settings/log-drains`) for long-term retention and custom analysis
-
-### Compliance considerations
-
-- AI Gateway does not log prompt or completion content by default
-- Enable content logging in project settings if required for compliance
-- Logs are retained per your Vercel plan's retention policy
-- Use `user` field consistently to support audit trails
-
-## Error Handling Patterns
-
-### Provider unavailable
-
-When a provider is down, the gateway automatically fails over if you configured `order` or `models`:
-
-```ts
-const result = await generateText({
-  model: gateway('anthropic/claude-sonnet-4.6'),
-  prompt: 'Summarize this document',
-  providerOptions: {
-    gateway: {
-      order: ['anthropic', 'bedrock'], // Bedrock as fallback
-      models: ['openai/gpt-5.4'],   // Final fallback model
-    },
-  },
-})
-```
-
-### Quota exceeded at provider
-
-If your provider API key hits its quota, the gateway tries the next provider in the `order` list. Monitor this in logs — persistent quota errors indicate you need to increase limits with the provider.
-
-### Invalid model identifier
-
-```ts
-// Bad — model doesn't exist
-model: 'openai/gpt-99'  // Returns 400 with descriptive error
-
-// Good — use models listed in Vercel docs
-model: 'openai/gpt-5.4'
-```
-
-### Timeout handling
-
-Gateway has a default timeout per provider. For long-running generations, use streaming:
+AI Gateway also supports OpenAI Responses, Anthropic Messages, OpenResponses, Cohere Rerank, embeddings, image and video generation, speech, transcription, realtime sessions, and evaluation. Modality pages under <https://vercel.com/docs/ai-gateway/modalities> cover each request shape, including background jobs for long-running video generation. Evaluation is available through AI SDK 7 or later, `POST /v1/evaluate`, and a TypeSafe-compatible API under `/typesafe`; it is not available through the OpenAI-, Anthropic-, or Cohere-compatible endpoints. Read [references/evaluation.md](references/evaluation.md) before choosing a surface. Read the relevant modality or API page instead of translating one request shape from memory.
 
-```ts
-import { streamText } from 'ai'
-
-const result = streamText({
-  model: 'anthropic/claude-sonnet-4.6',
-  prompt: longDocument,
-})
-
-for await (const chunk of result.textStream) {
-  process.stdout.write(chunk)
-}
-```
+## Minimal AI SDK request
 
-### Complete error handling template
+The current AI SDK requires Node.js 22 or later. Confirm the installed package's `engines` field before enforcing a version in an existing project.
 
 ```ts
-import { generateText, APICallError } from 'ai'
+import { generateText } from 'ai';
 
-async function callAI(prompt: string, userId: string) {
-  try {
-    return await generateText({
-      model: gateway('openai/gpt-5.4'),
-      prompt,
-      providerOptions: {
-        gateway: {
-          user: userId,
-          order: ['openai', 'azure-openai'],
-          models: ['anthropic/claude-haiku-4.5'],
-          tags: ['feature:chat'],
-        },
-      },
-    })
-  } catch (error) {
-    if (!APICallError.isInstance(error)) throw error
-
-    switch (error.statusCode) {
-      case 402: return { text: 'Budget limit reached. Please try again later.' }
-      case 429: return { text: 'Too many requests. Please slow down.' }
-      case 503: return { text: 'AI service temporarily unavailable.' }
-      default: throw error
-    }
-  }
+const model = process.env.AI_GATEWAY_MODEL;
+if (!model) {
+  throw new Error('Set AI_GATEWAY_MODEL to an ID returned by /v1/models');
 }
-```
-
-## Gateway vs Direct Provider — Decision Tree
-
-Use this to decide whether to route through AI Gateway or call a provider SDK directly:
-
-```
-Need failover across providers?
-  └─ Yes → Use Gateway
-  └─ No
-      Need cost tracking / budget alerts?
-        └─ Yes → Use Gateway
-        └─ No
-            Need per-user rate limiting?
-              └─ Yes → Use Gateway
-              └─ No
-                  Need audit logging?
-                    └─ Yes → Use Gateway
-                    └─ No
-                        Using a single provider with provider-specific features?
-                          └─ Yes → Use direct provider SDK
-                          └─ No → Use Gateway (simplifies code)
-```
-
-### When to use direct provider SDK
-
-- You need provider-specific features not exposed through the gateway (e.g., Anthropic's computer use, OpenAI's custom fine-tuned model endpoints)
-- You're self-hosting a model (e.g., vLLM, Ollama) that isn't registered with the gateway
-- You need request-level control over HTTP transport (custom proxies, mTLS)
 
-### When to always use Gateway
-
-- Production applications — failover and observability are essential
-- Multi-tenant SaaS — per-user tracking and rate limiting
-- Teams with cost accountability — tag-based budgeting
-
-## Latest Model Availability
-
-**GPT-5.4** (added March 5, 2026) — agentic and reasoning leaps from GPT-5.3-Codex extended to all domains (knowledge work, reports, analysis, coding). Faster and more token-efficient than GPT-5.2.
-
-| Model | Slug | Input | Output |
-|-------|------|-------|--------|
-| GPT-5.4 | `openai/gpt-5.4` | $2.50/M tokens | $15.00/M tokens |
-| GPT-5.4 Pro | `openai/gpt-5.4-pro` | $30.00/M tokens | $180.00/M tokens |
-
-GPT-5.4 Pro targets maximum performance on complex tasks. Use standard GPT-5.4 for most workloads.
-
-## Supported Providers
-
-- OpenAI (GPT-5.x including GPT-5.4 and GPT-5.4 Pro, o-series)
-- Anthropic (4.x models)
-- Google (Gemini)
-- xAI (Grok)
-- Mistral
-- DeepSeek
-- Amazon Bedrock
-- Azure OpenAI
-- Cohere
-- Perplexity
-- Alibaba (Qwen)
-- Meta (Llama)
-- And many more (100+ models total)
-
-## Pricing
-
-- **Zero markup**: Tokens at exact provider list price — no middleman markup, whether using Vercel-managed keys or Bring Your Own Key (BYOK)
-- **Free tier**: Every Vercel team gets **$5 of free AI Gateway credits per month** (refreshes every 30 days, starts on first request). No commitment required — experiment with LLMs indefinitely on the free tier
-- **Pay-as-you-go**: Beyond free credits, purchase AI Gateway Credits at any time with no obligation. Configure **auto top-up** to automatically add credits when your balance falls below a threshold
-- **BYOK**: Use your own provider API keys with zero fees from AI Gateway
-
-## Multimodal Support
-
-Text and image generation both route through the gateway. For embeddings, use a direct provider SDK.
-
-```ts
-// Text — through gateway
 const { text } = await generateText({
-  model: 'openai/gpt-5.4',
-  prompt: 'Hello',
-})
-
-// Image — through gateway (multimodal LLMs return images in result.files)
-const result = await generateText({
-  model: 'google/gemini-3.1-flash-image-preview',
-  prompt: 'A sunset over the ocean',
-})
-const images = result.files.filter((f) => f.mediaType?.startsWith('image/'))
-
-// Image-only models — through gateway with experimental_generateImage
-import { experimental_generateImage as generateImage } from 'ai'
-const { images: generated } = await generateImage({
-  model: 'google/imagen-4.0-generate-001',
-  prompt: 'A sunset',
-})
-```
-
-**Default image model**: `google/gemini-3.1-flash-image-preview` — fast multimodal image generation via gateway.
-
-See [AI Gateway Image Generation docs](https://vercel.com/docs/ai-gateway/capabilities/image-generation) for all supported models and integration methods.
-
-## Key Benefits
-
-1. **Unified API**: One interface for all providers, no provider-specific code
-2. **Automatic failover**: If a provider is down, requests route to the next
-3. **Cost tracking**: Per-user, per-feature attribution with tags
-4. **Observability**: Built-in monitoring of all model calls
-5. **Low latency**: <20ms routing overhead
-6. **No lock-in**: Switch models/providers by changing a string
-
-## When to Use AI Gateway
-
-| Scenario | Use Gateway? |
-|----------|-------------|
-| Production app with AI features | Yes — failover, cost tracking |
-| Prototyping with single provider | Optional — direct provider works fine |
-| Multi-provider setup | Yes — unified routing |
-| Need provider-specific features | Use direct provider SDK + Gateway as fallback |
-| Cost tracking and budgeting | Yes — user tracking and tags |
-| Multi-tenant SaaS | Yes — per-user rate limiting and audit |
-| Compliance requirements | Yes — audit logging and log drains |
-
-## Official Documentation
-
-- [AI Gateway](https://vercel.com/docs/ai-gateway)
-- [Providers and Models](https://ai-sdk.dev/docs/foundations/providers-and-models)
-- [AI SDK Core](https://ai-sdk.dev/docs/ai-sdk-core)
-- [GitHub: AI SDK](https://github.com/vercel/ai)
+  model,
+  prompt: 'Explain the project in one paragraph.',
+});
+
+console.log(text);
+```
+
+Honor an exact model the user or task specifies after confirming it exists. Otherwise fetch `/v1/models`, choose a model that fits the requested modality, capabilities, price, context window, data-retention policy, and team access, and set `AI_GATEWAY_MODEL` to that ID. Do not put a time-sensitive model recommendation in reusable examples.
+
+Plain model strings route through AI Gateway. Add `@ai-sdk/gateway` only when the task needs its exported provider, types, model discovery, generation lookup, or spend-report helpers.
+
+## Authentication decision
+
+- Use an **AI Gateway API key** for local scripts, CI, external servers, and non-Vercel deployments. Store it in `AI_GATEWAY_API_KEY` and never print or commit it.
+- Use **Vercel OIDC** for Vercel deployments and linked local projects. Vercel deployments receive `VERCEL_OIDC_TOKEN`; local development uses `vercel link` and `vercel env pull`.
+- **BYOK provider credentials do not replace AI Gateway request authentication.** They decide how AI Gateway authenticates to a model provider.
+- A plain Node.js script does not automatically load `.env.local`. Export variables in the shell or load that file explicitly. Framework behavior may differ.
+
+Do not ask the user to paste a secret into chat, source code, a committed config file, or a command that will enter shell history unless the repository has an established secure mechanism.
+
+## Implementation workflow
+
+1. Establish the user's job, runtime, deployment target, current provider, and required capabilities.
+2. Select authentication from the rules above. Preserve a working existing method unless the user asked to migrate it.
+3. Fetch live model metadata and choose a compatible model. State why it fits.
+4. Read the matching SDK, API, modality, routing, or coding-agent docs.
+5. Make the smallest end-to-end change. Reuse the current project structure and error handling.
+6. Handle only errors the application can act on. Common gateway outcomes include authentication failure, insufficient credits, budget exhaustion, rate limiting, and provider capacity failure.
+7. Run the project's formatter, type checker, and focused tests.
+8. When the task authorizes a live request, run one and inspect the returned model, text or media, usage, and provider metadata.
+9. Verify the request in AI Gateway Logs when dashboard access is available. Logs can take about 90 seconds to ingest.
+
+Only spend credits, create keys, change budgets, change routing rules, or write coding-agent config when the user requested or approved that outward-facing action. Prefer dry runs and interactive previews when available.
+
+## Routing invariants
+
+- Model IDs use the exact `provider/model` strings returned by `/v1/models`.
+- `order` controls provider preference, `only` restricts providers, and `sort` ranks providers by a supported metric.
+- `models` lists fallback models after the primary model.
+- `caching: 'auto'` manages provider prompt-cache markers. It is not an HTTP response cache.
+- `providerTimeouts` applies to BYOK provider attempts and measures time until the provider starts responding.
+- A reasoning entry in `providerOptions` overrides the AI SDK top-level `reasoning` value entirely; the two never merge.
+- `user` and `tags` attach reporting dimensions. They do not create per-user rate limits.
+- Request-scoped provider credentials belong under `providerOptions.gateway.byok` and must remain secret.
+- Virtual Model IDs use `vmc/<slug>`. A Virtual Model can pin routing and provider options server-side; settings it defines generally override the corresponding request settings, while unset settings remain request-configurable.
+
+Read [references/routing.md](references/routing.md) before adding any of these fields.
+
+## Verification checklist
+
+- [ ] A direct model ID exists in the full live model response, or a Virtual Model exists for the authenticated team and resolves as `vmc/<slug>`.
+- [ ] The selected API or SDK supports the requested modality and feature.
+- [ ] Authentication works in the actual runtime, including `.env.local` loading where relevant.
+- [ ] The example prints or returns a result instead of discarding the response.
+- [ ] The project type checker and focused tests pass.
+- [ ] A live request was made only when authorized, and its cost was understood.
+- [ ] Provider routing or fallback behavior is visible in response metadata or Logs.
+- [ ] No secret appears in source, logs, diffs, or the final response.
+- [ ] The final report distinguishes code verification from live and dashboard verification.
+
+## Current documentation
+
+- Getting started: <https://vercel.com/docs/ai-gateway/getting-started>
+- Models and providers: <https://vercel.com/docs/ai-gateway/models-and-providers>
+- Virtual Models: <https://vercel.com/docs/ai-gateway/models-and-providers/virtual-models>
+- SDKs and APIs: <https://vercel.com/docs/ai-gateway/sdks-and-apis>
+- Authentication and BYOK: <https://vercel.com/docs/ai-gateway/authentication-and-byok>
+- Observability and spend: <https://vercel.com/docs/ai-gateway/observability-and-spend>
+- Modalities: <https://vercel.com/docs/ai-gateway/modalities>
+- Evaluation: <https://vercel.com/docs/ai-gateway/modalities/evaluation>
+- TypeSafe API: <https://vercel.com/docs/ai-gateway/sdks-and-apis/typesafe>
+- REST API reference: <https://vercel.com/docs/ai-gateway/sdks-and-apis/rest-api>
+- FAQ: <https://vercel.com/docs/ai-gateway/faq>
+- Coding agents: <https://vercel.com/docs/ai-gateway/coding-agents>
+- AI SDK provider: <https://ai-sdk.dev/providers/ai-sdk-providers/ai-gateway>
Full snapshot data
{
  "description": "Vercel AI Gateway guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and coding-agent configuration. Use when adding AI Gateway to an app, migrating provider calls, choosing models or providers, centralizing model configuration, evaluating application state, debugging gateway requests, or running `vercel ai-gateway` commands.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 97
    },
    {
      "relative_path": "references/coding-agents.md",
      "size_in_bytes": 5817
    },
    {
      "relative_path": "references/evaluation.md",
      "size_in_bytes": 5442
    },
    {
      "relative_path": "references/routing.md",
      "size_in_bytes": 14459
    },
    {
      "relative_path": "references/setup.md",
      "size_in_bytes": 7651
    },
    {
      "relative_path": "references/spend-observability.md",
      "size_in_bytes": 8914
    },
    {
      "relative_path": "references/virtual-models.md",
      "size_in_bytes": 4513
    }
  ],
  "name": "ai-gateway",
  "skill_md_contents": "---\nname: ai-gateway\ndescription: Vercel AI Gateway guidance for setup, model discovery, authentication, routing, fallbacks, virtual models, evaluation models, BYOK, budgets, spend reporting, observability, compatible APIs, and coding-agent configuration. Use when adding AI Gateway to an app, migrating provider calls, choosing models or providers, centralizing model configuration, evaluating application state, debugging gateway requests, or running `vercel ai-gateway` commands.\nsummary: Set up and operate Vercel AI Gateway with current models, virtual models, evaluation, authentication, routing, spend controls, and verification.\nmetadata:\n  priority: 7\n  docs:\n    - \"https://vercel.com/docs/ai-gateway\"\n    - \"https://vercel.com/docs/ai-gateway/getting-started\"\n    - \"https://vercel.com/docs/ai-gateway/models-and-providers/virtual-models\"\n    - \"https://vercel.com/docs/ai-gateway/modalities/evaluation\"\n    - \"https://vercel.com/docs/ai-gateway/sdks-and-apis/typesafe\"\n    - \"https://ai-sdk.dev/providers/ai-sdk-providers/ai-gateway\"\n  sitemap: \"https://vercel.com/docs/sitemap.md\"\n  pathPatterns: []\n  importPatterns:\n    - 'ai'\n    - '@ai-sdk/gateway'\n  bashPatterns:\n    - '\\bvercel\\s+ai-gateway\\b'\n    - '\\bvercel\\s+env\\s+pull\\b'\n    - '\\bnpm\\s+(install|i|add)\\s+[^\\n]*@ai-sdk/gateway\\b'\n    - '\\bpnpm\\s+(install|i|add)\\s+[^\\n]*@ai-sdk/gateway\\b'\n    - '\\bbun\\s+(install|i|add)\\s+[^\\n]*@ai-sdk/gateway\\b'\n    - '\\byarn\\s+add\\s+[^\\n]*@ai-sdk/gateway\\b'\n  promptSignals:\n    phrases:\n      - \"ai gateway\"\n      - \"vercel ai gateway\"\n      - \"ai-gateway\"\n      - \"ai-gateway.vercel.sh\"\n      - \"virtual model\"\n      - \"evaluation model\"\n      - \"vmc/\"\n      - \"typesafe api\"\n      - \"typesafe compat\"\n    allOf:\n      - [model, routing]\n      - [provider, failover]\n      - [gateway, budget]\n      - [gateway, logs]\n      - [gateway, oidc]\n      - [coding, gateway]\n    anyOf:\n      - \"provider ordering\"\n      - \"model fallback\"\n      - \"byok\"\n      - \"spend tracking\"\n      - \"gateway key\"\n      - \"credit balance\"\n      - \"safety identifier\"\n      - \"reasoning effort\"\n      - \"tool calling\"\n      - \"structured outputs\"\n      - \"experimental_evaluate\"\n      - \"central model configuration\"\n      - \"systemone\"\n      - \"v1/evaluate\"\n    noneOf:\n      - \"cloudflare ai gateway\"\n      - \"aws api gateway\"\n    minScore: 6\nvalidate:\n  -\n    pattern: '\\bclaude-(sonnet|opus|haiku)-\\d+-\\d+\\b'\n    message: 'Claude model version uses a hyphen where the AI Gateway slug uses a dot. Fetch /v1/models and use the returned provider/model ID.'\n    severity: error\n  -\n    pattern: gateway\\(['\"][^'\"/]+['\"]\\)\n    message: 'AI Gateway model string is missing its provider prefix. Fetch /v1/models and use a provider/model ID.'\n    severity: error\n  -\n    pattern: (OPENAI_API_KEY|ANTHROPIC_API_KEY|GOOGLE_API_KEY)\n    message: 'Provider key detected. AI Gateway request authentication uses AI_GATEWAY_API_KEY or VERCEL_OIDC_TOKEN; provider keys belong only in an intentional BYOK configuration.'\n    severity: recommended\n    skipIfFileContains: '[Bb][Yy][Oo][Kk]|providerOptions\\s*:\\s*\\{[^}]*gateway'\n  -\n    pattern: gateway\\s*:\\s*\\{[^}]*cacheControl\n    message: \"AI Gateway does not cache whole responses through cacheControl. Use caching: 'auto' for provider prompt caching and verify the current caching docs.\"\n    severity: error\n  -\n    pattern: ANTHROPIC_BASE_URL\\s*=\\s*[\"']?https://ai-gateway\\.vercel\\.sh\n    message: 'Claude Code through AI Gateway needs ANTHROPIC_API_KEY set to an empty value and the gateway key in ANTHROPIC_AUTH_TOKEN. A non-empty ANTHROPIC_API_KEY is used instead of the gateway token.'\n    severity: recommended\n    skipIfFileContains: 'ANTHROPIC_AUTH_TOKEN'\nchainTo:\n  -\n    pattern: 'from\\s+[''\"]ai[''\"]|require\\([''\"]ai[''\"]\\)|\\b(generateText|streamText|ToolLoopAgent)\\b'\n    targetSkill: ai-sdk\n    message: 'AI SDK code detected. Load the AI SDK skill and read the installed package docs before writing or changing SDK code.'\nretrieval:\n  aliases:\n    - model router\n    - ai proxy\n    - provider failover\n    - llm gateway\n    - gateway credits\n    - virtual model\n    - evaluation model\n    - typesafe compatibility\n  intents:\n    - add Vercel AI Gateway to an application\n    - route AI models across providers\n    - configure provider or model fallbacks\n    - authenticate AI Gateway requests\n    - track AI model costs and set budgets\n    - debug AI Gateway requests and routing\n    - connect coding agents to AI Gateway\n    - give coding agents centrally managed model and provider configuration\n    - create or update an AI Gateway virtual model\n    - evaluate application state with typed questions\n    - migrate an existing TypeSafe evaluation client to AI Gateway\n    - find a model by modality, capability, price, or data retention\n    - check AI Gateway credit balance or generation cost\n    - configure reasoning or extended thinking across providers and API formats\n    - add tool calling or function calling across API formats\n    - get structured JSON output matching a schema\n    - send images or PDFs to a model\n  entities:\n    - AI Gateway\n    - AI Gateway Credits\n    - providerOptions.gateway\n    - AI_GATEWAY_API_KEY\n    - VERCEL_OIDC_TOKEN\n    - model routing\n    - provider failover\n    - BYOK\n    - spend reporting\n    - safetyIdentifier\n    - Usage & Billing API\n    - Virtual Models\n    - vmc/<slug>\n    - experimental_evaluate\n    - POST /v1/evaluate\n    - TypeSafe API\n    - /typesafe/v1/systemone\n---\n\n# Vercel AI Gateway\n\nAI Gateway exposes models from multiple providers through shared authentication, model IDs, routing, billing, and observability. Model availability, SDK APIs, CLI commands, prices, and product capabilities change frequently. Verify them from current sources before changing code.\n\n## Start with current sources\n\nBefore implementing:\n\n1. Inspect the project's language, package manager, installed AI SDK version, and existing provider integration.\n2. Read the relevant Vercel page under <https://vercel.com/docs/ai-gateway>. Use the page's `.md` form when a tool needs Markdown.\n3. Fetch the complete live model list. Do not construct model variants by analogy:\n\n   ```bash\n   curl -fsSL https://ai-gateway.vercel.sh/v1/models\n   ```\n\n4. If the code uses the `ai` package, load the `ai-sdk` skill when available. Read version-matched docs under `node_modules/ai/docs/` and source under `node_modules/ai/src/`. If the skill is not installed, use those bundled files directly.\n5. Run `vercel ai-gateway <command> --help` before documenting or scripting CLI flags.\n\nThe live model endpoint and installed package take precedence over model names or SDK syntax remembered from training data.\n\n## Vercel CLI inventory\n\nThe `vercel ai-gateway` command manages gateway resources for the current team. The `setup` subcommand connects local coding agents; the rest of the CLI covers the jobs that previously required dashboard work:\n\n| Command | What it does |\n| --- | --- |\n| `api-keys create/list/inspect/remove` | Create and manage AI Gateway API keys, with budgets, spend alerts, expiry, and restriction exemptions |\n| `budgets set/list/inspect/remove` | Set metered spend limits for the team, a project, a user, or an API key |\n| `budgets defaults set/list/remove` | Set per-scope default limits covering projects, keys, or members without a custom budget |\n| `models list` / `models endpoints <model>` | List the model catalog and one model's provider endpoints from the CLI |\n| `virtual-models create/list/inspect/edit/remove/restore` | Manage reusable, team-scoped model configurations addressed as `vmc/<slug>` |\n| `rules add/list/edit/remove` | Manage routing rules; the CLI marks rules beta, so check `--help` before relying on them. REST CRUD exists under `/v1/ai-gateway/rules` |\n| `setup` | Configure supported coding agents; see [references/coding-agents.md](references/coding-agents.md) |\n| `leaderboard` | Explore public, anonymized usage leaderboards; rarely needed for implementation work |\n\nUse the CLI for credential and spend management when the user is working from a terminal or in CI. Check `vercel ai-gateway <command> --help` for current flags before scripting; do not copy a flag list from this skill into generated code.\n\n## Route the request to the right guide\n\n| User's job | Read |\n| --- | --- |\n| Ask a coding agent to make one Gateway request, or handle first-request credentials, compatible SDKs, or migration | [references/setup.md](references/setup.md) |\n| Provider selection, model fallbacks, caching, BYOK, or timeouts | [references/routing.md](references/routing.md) |\n| Reusable model configuration, a `vmc/<slug>`, or provider options for a client that cannot send them | [references/virtual-models.md](references/virtual-models.md) |\n| Typed evaluation through AI SDK, `POST /v1/evaluate`, or the TypeSafe-compatible API | [references/evaluation.md](references/evaluation.md) |\n| Credits, budgets, reporting, Logs, or request debugging | [references/spend-observability.md](references/spend-observability.md) |\n| Route Claude Code, Codex, OpenCode, Pi, or another coding agent's own model traffic through Gateway | [references/coding-agents.md](references/coding-agents.md) |\n\nRead each relevant reference before editing. A task can require more than one.\n\n## Choose the integration surface\n\n| Existing project | Default path |\n| --- | --- |\n| JavaScript or TypeScript using AI SDK | Use a plain `provider/model` string with `generateText`, `streamText`, `ToolLoopAgent`, or the relevant modality API |\n| Python using AI SDK for Python | Use `ai.get_model('provider/model')` and the current Python SDK docs |\n| Existing OpenAI SDK | Keep the SDK and point `baseURL` or `base_url` to `https://ai-gateway.vercel.sh/v1` |\n| Existing Anthropic SDK | Keep the SDK and point `baseURL` or `base_url` to `https://ai-gateway.vercel.sh` |\n| Evaluation over provider-neutral HTTP | Send Gateway's evaluation request shape to `POST https://ai-gateway.vercel.sh/v1/evaluate` |\n| Existing TypeSafe evaluation client | Keep `@typesafe-ai/sdk` and point `baseURL` to `https://ai-gateway.vercel.sh/typesafe` |\n| Provider-neutral HTTP | Use an AI Gateway compatible endpoint, such as Chat Completions or OpenResponses |\n| Existing direct-provider AI SDK integration | Replace the provider instance with a live AI Gateway `provider/model` string, then remove provider credentials only after verifying the gateway path |\n| Coding agent | Use `vercel ai-gateway setup`; inspect its help before claiming agent support. Use a Virtual Model when the agent needs reusable routing or provider options it cannot send per request |\n\nAI Gateway also supports OpenAI Responses, Anthropic Messages, OpenResponses, Cohere Rerank, embeddings, image and video generation, speech, transcription, realtime sessions, and evaluation. Modality pages under <https://vercel.com/docs/ai-gateway/modalities> cover each request shape, including background jobs for long-running video generation. Evaluation is available through AI SDK 7 or later, `POST /v1/evaluate`, and a TypeSafe-compatible API under `/typesafe`; it is not available through the OpenAI-, Anthropic-, or Cohere-compatible endpoints. Read [references/evaluation.md](references/evaluation.md) before choosing a surface. Read the relevant modality or API page instead of translating one request shape from memory.\n\n## Minimal AI SDK request\n\nThe current AI SDK requires Node.js 22 or later. Confirm the installed package's `engines` field before enforcing a version in an existing project.\n\n```ts\nimport { generateText } from 'ai';\n\nconst model = process.env.AI_GATEWAY_MODEL;\nif (!model) {\n  throw new Error('Set AI_GATEWAY_MODEL to an ID returned by /v1/models');\n}\n\nconst { text } = await generateText({\n  model,\n  prompt: 'Explain the project in one paragraph.',\n});\n\nconsole.log(text);\n```\n\nHonor an exact model the user or task specifies after confirming it exists. Otherwise fetch `/v1/models`, choose a model that fits the requested modality, capabilities, price, context window, data-retention policy, and team access, and set `AI_GATEWAY_MODEL` to that ID. Do not put a time-sensitive model recommendation in reusable examples.\n\nPlain model strings route through AI Gateway. Add `@ai-sdk/gateway` only when the task needs its exported provider, types, model discovery, generation lookup, or spend-report helpers.\n\n## Authentication decision\n\n- Use an **AI Gateway API key** for local scripts, CI, external servers, and non-Vercel deployments. Store it in `AI_GATEWAY_API_KEY` and never print or commit it.\n- Use **Vercel OIDC** for Vercel deployments and linked local projects. Vercel deployments receive `VERCEL_OIDC_TOKEN`; local development uses `vercel link` and `vercel env pull`.\n- **BYOK provider credentials do not replace AI Gateway request authentication.** They decide how AI Gateway authenticates to a model provider.\n- A plain Node.js script does not automatically load `.env.local`. Export variables in the shell or load that file explicitly. Framework behavior may differ.\n\nDo not ask the user to paste a secret into chat, source code, a committed config file, or a command that will enter shell history unless the repository has an established secure mechanism.\n\n## Implementation workflow\n\n1. Establish the user's job, runtime, deployment target, current provider, and required capabilities.\n2. Select authentication from the rules above. Preserve a working existing method unless the user asked to migrate it.\n3. Fetch live model metadata and choose a compatible model. State why it fits.\n4. Read the matching SDK, API, modality, routing, or coding-agent docs.\n5. Make the smallest end-to-end change. Reuse the current project structure and error handling.\n6. Handle only errors the application can act on. Common gateway outcomes include authentication failure, insufficient credits, budget exhaustion, rate limiting, and provider capacity failure.\n7. Run the project's formatter, type checker, and focused tests.\n8. When the task authorizes a live request, run one and inspect the returned model, text or media, usage, and provider metadata.\n9. Verify the request in AI Gateway Logs when dashboard access is available. Logs can take about 90 seconds to ingest.\n\nOnly spend credits, create keys, change budgets, change routing rules, or write coding-agent config when the user requested or approved that outward-facing action. Prefer dry runs and interactive previews when available.\n\n## Routing invariants\n\n- Model IDs use the exact `provider/model` strings returned by `/v1/models`.\n- `order` controls provider preference, `only` restricts providers, and `sort` ranks providers by a supported metric.\n- `models` lists fallback models after the primary model.\n- `caching: 'auto'` manages provider prompt-cache markers. It is not an HTTP response cache.\n- `providerTimeouts` applies to BYOK provider attempts and measures time until the provider starts responding.\n- A reasoning entry in `providerOptions` overrides the AI SDK top-level `reasoning` value entirely; the two never merge.\n- `user` and `tags` attach reporting dimensions. They do not create per-user rate limits.\n- Request-scoped provider credentials belong under `providerOptions.gateway.byok` and must remain secret.\n- Virtual Model IDs use `vmc/<slug>`. A Virtual Model can pin routing and provider options server-side; settings it defines generally override the corresponding request settings, while unset settings remain request-configurable.\n\nRead [references/routing.md](references/routing.md) before adding any of these fields.\n\n## Verification checklist\n\n- [ ] A direct model ID exists in the full live model response, or a Virtual Model exists for the authenticated team and resolves as `vmc/<slug>`.\n- [ ] The selected API or SDK supports the requested modality and feature.\n- [ ] Authentication works in the actual runtime, including `.env.local` loading where relevant.\n- [ ] The example prints or returns a result instead of discarding the response.\n- [ ] The project type checker and focused tests pass.\n- [ ] A live request was made only when authorized, and its cost was understood.\n- [ ] Provider routing or fallback behavior is visible in response metadata or Logs.\n- [ ] No secret appears in source, logs, diffs, or the final response.\n- [ ] The final report distinguishes code verification from live and dashboard verification.\n\n## Current documentation\n\n- Getting started: <https://vercel.com/docs/ai-gateway/getting-started>\n- Models and providers: <https://vercel.com/docs/ai-gateway/models-and-providers>\n- Virtual Models: <https://vercel.com/docs/ai-gateway/models-and-providers/virtual-models>\n- SDKs and APIs: <https://vercel.com/docs/ai-gateway/sdks-and-apis>\n- Authentication and BYOK: <https://vercel.com/docs/ai-gateway/authentication-and-byok>\n- Observability and spend: <https://vercel.com/docs/ai-gateway/observability-and-spend>\n- Modalities: <https://vercel.com/docs/ai-gateway/modalities>\n- Evaluation: <https://vercel.com/docs/ai-gateway/modalities/evaluation>\n- TypeSafe API: <https://vercel.com/docs/ai-gateway/sdks-and-apis/typesafe>\n- REST API reference: <https://vercel.com/docs/ai-gateway/sdks-and-apis/rest-api>\n- FAQ: <https://vercel.com/docs/ai-gateway/faq>\n- Coding agents: <https://vercel.com/docs/ai-gateway/coding-agents>\n- AI SDK provider: <https://ai-sdk.dev/providers/ai-sdk-providers/ai-gateway>\n"
}

SHA-256 of public snapshot: 4008eeb4aa9681f61fad4670f46c8c3a3c5ecce430cea4916be2595dea8ca702