diff --git a/skills/ai-gateway/SKILL.md b/skills/ai-gateway/SKILL.md index c4e208f..2282688 100644 --- a/skills/ai-gateway/SKILL.md +++ b/skills/ai-gateway/SKILL.md @@ -181,7 +181,7 @@ Use the CLI for credential and spend management when the user is working from a | Ask a coding agent to make one Gateway request, or handle first-request credentials, compatible SDKs, or migration | [references/setup.md](references/setup.md) | | Provider selection, model fallbacks, caching, BYOK, or timeouts | [references/routing.md](references/routing.md) | | Reusable model configuration, a `vmc/`, or provider options for a client that cannot send them | [references/virtual-models.md](references/virtual-models.md) | -| Typed evaluation through AI SDK, `POST /v1/evaluate`, or the TypeSafe-compatible API | [references/evaluation.md](references/evaluation.md) | +| Typed evaluation (decisions) through AI SDK, `POST /v1/evaluate`, the OpenAI-compatible `POST /v1/decisions`, or the TypeSafe-compatible API | [references/evaluation.md](references/evaluation.md) | | Credits, budgets, reporting, Logs, or request debugging | [references/spend-observability.md](references/spend-observability.md) | | Route Claude Code, Codex, OpenCode, Pi, or another coding agent's own model traffic through Gateway | [references/coding-agents.md](references/coding-agents.md) | @@ -201,7 +201,7 @@ Read each relevant reference before editing. A task can require more than one. | Existing direct-provider AI SDK integration | Replace the provider instance with a live AI Gateway `provider/model` string, then remove provider credentials only after verifying the gateway path | | Coding agent | Use `vercel ai-gateway setup`; inspect its help before claiming agent support. Use a Virtual Model when the agent needs reusable routing or provider options it cannot send per request | -AI Gateway also supports OpenAI Responses, Anthropic Messages, OpenResponses, Cohere Rerank, embeddings, image and video generation, speech, transcription, realtime sessions, and evaluation. Modality pages under cover each request shape, including background jobs for long-running video generation. Evaluation is available through AI SDK 7 or later, `POST /v1/evaluate`, and a TypeSafe-compatible API under `/typesafe`; it is not available through the OpenAI-, Anthropic-, or Cohere-compatible endpoints. Read [references/evaluation.md](references/evaluation.md) before choosing a surface. Read the relevant modality or API page instead of translating one request shape from memory. +AI Gateway also supports OpenAI Responses, Anthropic Messages, OpenResponses, Cohere Rerank, embeddings, image and video generation, speech, transcription, realtime sessions, and evaluation. Modality pages under cover each request shape, including background jobs for long-running video generation. Evaluation (now called Decision) is available through AI SDK 7 or later, `POST /v1/evaluate`, the OpenAI-compatible `POST /v1/decisions`, and a TypeSafe-compatible API under `/typesafe`; it is not available through Chat Completions, Responses, or the Anthropic- or Cohere-compatible endpoints. Read [references/evaluation.md](references/evaluation.md) before choosing a surface. Read the relevant modality or API page instead of translating one request shape from memory. ## Minimal AI SDK request @@ -285,7 +285,7 @@ Read [references/routing.md](references/routing.md) before adding any of these f - Authentication and BYOK: - Observability and spend: - Modalities: -- Evaluation: +- Decision (formerly Evaluation): - TypeSafe API: - REST API reference: - FAQ: diff --git a/skills/ai-gateway/references/evaluation.md b/skills/ai-gateway/references/evaluation.md index df384a0..f525ec0 100644 --- a/skills/ai-gateway/references/evaluation.md +++ b/skills/ai-gateway/references/evaluation.md @@ -1,27 +1,28 @@ # Evaluation models -Use evaluation models to assess shared application state against typed questions. They return structured boolean probabilities, choices, or scores instead of free-form text. +Use evaluation models to assess shared application state against typed questions. They return structured boolean probabilities, choices, or scores instead of free-form text. AI Gateway docs now call them decision models; `experimental_evaluate` and `gateway.evaluationModel()` remain as deprecated aliases. ## Choose the evaluation surface | Existing project | Use | | --- | --- | -| JavaScript or TypeScript using AI SDK 7 or later | `experimental_evaluate` from `ai` | +| JavaScript or TypeScript using AI SDK 7 or later | `experimental_decide` from `ai` 7.0.128+; earlier AI SDK 7 releases use `experimental_evaluate` | +| Existing OpenAI SDK using the Decisions API | Keep `decisions.create` and point `baseURL` to `https://ai-gateway.vercel.sh/v1` (`POST /v1/decisions`) | | New non-AI-SDK client | Gateway's vendor-neutral `POST /v1/evaluate` | | Existing TypeSafe client | Keep `@typesafe-ai/sdk` and change its API key and `baseURL` | -Evaluation is not supported through the OpenAI-compatible, Anthropic-compatible, or Cohere-compatible endpoints. Use one of the three surfaces above instead. +Evaluation is not supported through Chat Completions, Responses, or the Anthropic-compatible or Cohere-compatible endpoints. Use one of the four surfaces above instead. -Choose a model from the current [Evaluation model list](https://vercel.com/ai-gateway/models?capabilities=evaluation). Do not assume a text-generation model supports evaluation. Use the [Evaluation quickstart](https://vercel.com/docs/ai-gateway/getting-started/evaluation) for the AI SDK first-run workflow, the [Evaluation modality guide](https://vercel.com/docs/ai-gateway/modalities/evaluation) for AI SDK and HTTP request shapes, and the [TypeSafe API guide](https://vercel.com/docs/ai-gateway/sdks-and-apis/typesafe) when migrating a TypeSafe client. +Choose a model from the current [Decision model list](https://vercel.com/ai-gateway/models?capabilities=decision). Do not assume a text-generation model supports evaluation. Use the [Decision quickstart](https://vercel.com/docs/ai-gateway/getting-started/decision) for the AI SDK first-run workflow, the [Decision modality guide](https://vercel.com/docs/ai-gateway/modalities/decision) for AI SDK and HTTP request shapes, and the [TypeSafe API guide](https://vercel.com/docs/ai-gateway/sdks-and-apis/typesafe) when migrating a TypeSafe client. ## AI SDK Load the `ai-sdk` skill and read the installed `ai` package docs before writing code. ```ts -import { experimental_evaluate as evaluate } from 'ai'; +import { experimental_decide as decide } from 'ai'; -const result = await evaluate({ +const result = await decide({ model: 'typesafe-ai/jev', state: 'The support agent issued a full refund to the customer.', questions: { @@ -35,7 +36,7 @@ const result = await evaluate({ console.log(result.answers); ``` -When using an explicit Gateway provider instance, use `gateway.evaluationModel()`. A plain evaluation model string routes through AI Gateway without installing `@ai-sdk/gateway`. +When using an explicit Gateway provider instance, use `gateway.decisionModel()` (`@ai-sdk/gateway` 4.0.104+, which ships with `ai` 7.0.128+; earlier releases use the deprecated `gateway.evaluationModel()`). A plain evaluation model string routes through AI Gateway without installing `@ai-sdk/gateway`. ## HTTP API @@ -77,7 +78,7 @@ The compatibility surface provides: The TypeSafe-compatible request body also accepts AI Gateway controls under `providerOptions.gateway`. Use the generic `/v1/evaluate` endpoint for new HTTP integrations; use `/typesafe` when preserving an existing TypeSafe client is the goal. -On any evaluation surface, use [Evaluation Fallbacks](https://vercel.com/docs/ai-gateway/models-and-providers/evaluation-fallbacks) to rerun a successful but uncertain evaluation with another model: a conditional entry in `providerOptions.gateway.models`, such as a `confidenceBelow` threshold on a Choice or Score question. A triggered fallback bills both stages. +On any evaluation surface, use [Decision Fallbacks](https://vercel.com/docs/ai-gateway/models-and-providers/decision-fallbacks) to rerun a successful but uncertain evaluation with another model: a conditional entry in `providerOptions.gateway.models`, such as a `confidenceBelow` threshold on a Choice or Score question. A triggered fallback bills both stages. ## Question and state rules