Provider adapter selection and configuration: openaiText, anthropicText, geminiText, ollamaText, grokText, groqText, openRouterText, bedrockText, byteplusText, openaiCompatible. Per-model type safety with modelOptions, reasoning/thinking configuration, runtime adapter switching, extendAdapter() for custom models, createModel(). Generic OpenAI-compatible providers (DeepSeek, Together, Fireworks, etc.) via openaiCompatible({ baseURL, apiKey, models }) from @tanstack/ai-openai/compatible. API key env vars: OPENAI_API_KEY, ANTHROPIC_API_KEY, GOOGLE_API_KEY/GEMINI_API_KEY, XAI_API_KEY, GROQ_API_KEY, OPENROUTER_API_KEY, OLLAMA_HOST, BEDROCK_API_KEY (or AWS_BEARER_TOKEN_BEDROCK). BytePlus needs TWO keys: ARK_API_KEY (ModelArk — chat/video/image) and BYTEPLUS_VOICE_API_KEY (Seed Speech — TTS/transcription); neither is a fallback for the other.
Dependency: This skill builds on ai-core. Read it first for critical rules.
Before implementing: Ask the user which provider and model they want. Then fetch the latest available models from the provider's source code (check the adapter's model metadata file, e.g.
packages/ai-openai/src/model-meta.ts) or from the provider's API/docs to recommend the most current model. The model lists in this skill and its reference files may be outdated. Always verify against the source before recommending a specific model.
Create an adapter and use it with chat():
import { chat, toServerSentEventsResponse } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
export async function POST(request: Request) {
const { messages } = await request.json()
const stream = chat({
adapter: openaiText('gpt-5.2'),
messages,
modelOptions: {
temperature: 0.7,
max_output_tokens: 1000,
},
})
return toServerSentEventsResponse(stream)
}The adapter factory function takes the model name as a string literal and an
optional config object (API key, base URL, etc.). The model name is passed
into the factory, not into chat().
Sampling options (temperature, token limits, top_p/topP, etc.) live
inside modelOptions using each provider's native key — they are not
top-level options on chat(). See the per-provider table in
Configuring Sampling below.
Each provider has a dedicated package with tree-shakeable adapter factories. The text adapter is the primary one for chat/completions:
| Provider | Package | Factory | Env Var |
|---|---|---|---|
| OpenAI | @tanstack/ai-openai | openaiText | OPENAI_API_KEY |
| Anthropic | @tanstack/ai-anthropic | anthropicText | ANTHROPIC_API_KEY |
| Gemini | @tanstack/ai-gemini | geminiText | GOOGLE_API_KEY or GEMINI_API_KEY |
| Grok (xAI) | @tanstack/ai-grok | grokText | XAI_API_KEY |
| Groq | @tanstack/ai-groq | groqText | GROQ_API_KEY |
| OpenRouter | @tanstack/ai-openrouter | openRouterText | OPENROUTER_API_KEY |
| Ollama | @tanstack/ai-ollama | ollamaText | OLLAMA_HOST (default: http://localhost:11434) |
| Bedrock | @tanstack/ai-bedrock | bedrockText | BEDROCK_API_KEY or AWS_BEARER_TOKEN_BEDROCK |
| BytePlus | @tanstack/ai-byteplus | byteplusText | ARK_API_KEY (falls back to BYTEPLUS_API_KEY) |
| OpenAI-compatible | @tanstack/ai-openai/compatible | openaiCompatible / openaiCompatibleText | provider-specific (passed via apiKey) |
| Cloudflare | @tanstack/ai-cloudflare | cloudflareText | CLOUDFLARE_ACCOUNT_ID + CLOUDFLARE_API_TOKEN, or { binding: env.AI } in a Worker |
BytePlus uses two keys.
byteplusText/byteplusVideo/byteplusImagereadARK_API_KEY(ModelArk,Authorization: Bearer), butbyteplusSpeech/byteplusTranscriptionare a separate product and readBYTEPLUS_VOICE_API_KEY(Seed Speech,X-Api-Key). Ark keys are also region-isolated — the default base URL is the ap-southeast endpoint.
// Each factory takes model as first arg, optional config as second
import { openaiText, createOpenaiChat } from '@tanstack/ai-openai'
import { anthropicText } from '@tanstack/ai-anthropic'
import { geminiText } from '@tanstack/ai-gemini'
import { grokText } from '@tanstack/ai-grok'
import { groqText } from '@tanstack/ai-groq'
import { openRouterText } from '@tanstack/ai-openrouter'
import { ollamaText } from '@tanstack/ai-ollama'
import { bedrockText } from '@tanstack/ai-bedrock'
import { byteplusText } from '@tanstack/ai-byteplus'
// Model string is passed to the factory, NOT to chat()
const adapter = openaiText('gpt-5.2')
const adapter2 = anthropicText('claude-sonnet-4-6')
const adapter3 = geminiText('gemini-2.5-pro')
const adapter4 = grokText('grok-4.6')
const adapter5 = groqText('llama-3.3-70b-versatile')
const adapter6 = openRouterText('anthropic/claude-sonnet-4')
const adapter7 = ollamaText('llama3.3:latest')
const adapter8 = bedrockText('us.anthropic.claude-3-7-sonnet-20250219-v1:0')
const adapter9 = byteplusText('seed-2-0-lite-260428')
// Optional: pass an explicit API key via the create* sibling
// (the plain factory reads it from the environment)
const adapterWithKey = createOpenaiChat('gpt-5.2', 'sk-...')@tanstack/ai-bedrock (Amazon Bedrock) branches on config.api:
bedrockText(model) or bedrockText(model, { api: 'converse' }) (the default) — Bedrock's native Converse API via @aws-sdk/client-bedrock-runtime (adapter name bedrock-converse). Reaches the broad catalog: Claude, Nova, Llama, Mistral, DeepSeek, and more.bedrockText(model, { api: 'chat' }) — OpenAI-compatible Chat Completions endpoint (adapter name bedrock). Open-weight models only (gpt-oss, DeepSeek V3.x, Gemma, Qwen, etc.). Does NOT reach Claude, Nova, or Llama.bedrockText(model, { api: 'responses' }) — OpenAI-compatible Responses API, mantle-only (adapter name bedrock-responses). Currently gpt-oss and Gemma 4.Use createBedrockText(model, apiKey, config?) to pass the key explicitly. Auth resolves from BEDROCK_API_KEY / AWS_BEARER_TOKEN_BEDROCK, or SigV4 via the standard AWS credential chain (no extra packages needed — handled by @aws-sdk/client-bedrock-runtime).
Use an adapter factory map to switch providers dynamically based on user input or configuration:
import { chat, toServerSentEventsResponse } from '@tanstack/ai'
import type { ModelMessage } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
import { anthropicText } from '@tanstack/ai-anthropic'
import { geminiText } from '@tanstack/ai-gemini'
// Define a map of provider+model to adapter factory calls
const adapters = {
'openai/gpt-5.2': () => openaiText('gpt-5.2'),
'anthropic/claude-sonnet-4-6': () => anthropicText('claude-sonnet-4-6'),
'gemini/gemini-2.5-pro': () => geminiText('gemini-2.5-pro'),
}
function isKnownProviderModel(key: string): key is keyof typeof adapters {
return key in adapters
}
export function handleChat(
providerModel: string,
messages: Array<ModelMessage>,
) {
if (!isKnownProviderModel(providerModel)) {
throw new Error(`Unknown provider/model: ${providerModel}`)
}
const stream = chat({
adapter: adapters[providerModel](),
messages,
})
return toServerSentEventsResponse(stream)
}Different providers expose reasoning/thinking through their modelOptions:
import { chat } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
import { anthropicText } from '@tanstack/ai-anthropic'
import { geminiText } from '@tanstack/ai-gemini'
const messages = [
{ role: 'user' as const, content: 'Plan a database migration.' },
]
// OpenAI: reasoning with effort and summary
const openaiStream = chat({
adapter: openaiText('gpt-5.2'),
messages,
modelOptions: {
reasoning: {
effort: 'high',
summary: 'auto',
},
},
})
// Anthropic: extended thinking with budget_tokens
const anthropicStream = chat({
adapter: anthropicText('claude-sonnet-4-6'),
messages,
modelOptions: {
max_tokens: 16000,
thinking: {
type: 'enabled',
budget_tokens: 8000, // must be >= 1024 and < max_tokens
},
},
})
// Anthropic: adaptive thinking (Sonnet 5, Fable 5, Opus 4.7+) — depth is
// tuned with output_config.effort instead of a token budget
const adaptiveStream = chat({
adapter: anthropicText('claude-sonnet-5'),
messages,
modelOptions: {
max_tokens: 16000,
thinking: {
type: 'adaptive',
display: 'summarized', // stream the reasoning text (default 'omitted')
},
output_config: { effort: 'high' }, // 'low' | 'medium' | 'high' | 'xhigh' | 'max'
},
})
// Gemini: thinking config with budget or level
const geminiStream = chat({
adapter: geminiText('gemini-2.5-pro'),
messages,
modelOptions: {
thinkingConfig: {
includeThoughts: true,
thinkingBudget: 4096,
},
},
})Use extendAdapter() and createModel() to add custom or fine-tuned models
while preserving type safety for the original models:
import { extendAdapter, createModel } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
// Define custom models
const customModels = [
createModel('ft:gpt-5.2:my-org:custom-model:abc123', ['text', 'image']),
createModel('my-local-proxy-model', ['text']),
] as const
// Create extended factory - original models still fully typed
const myOpenai = extendAdapter(openaiText, customModels)
// Use original models - full type inference preserved
const gpt5 = myOpenai('gpt-5.2')
// Use custom models - accepted by the type system
const custom = myOpenai('ft:gpt-5.2:my-org:custom-model:abc123')
// Type error: 'nonexistent-model' is not a valid model
// myOpenai('nonexistent-model')At runtime, extendAdapter simply passes through to the original factory.
The _customModels parameter is only used for type inference.
Sampling controls (temperature, token limits, nucleus sampling) are passed
inside modelOptions using each provider's native key. They are not
top-level fields on chat()/ai()/generate().
import { chat } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
import { anthropicText } from '@tanstack/ai-anthropic'
import { geminiText } from '@tanstack/ai-gemini'
import { ollamaText } from '@tanstack/ai-ollama'
const messages = [{ role: 'user' as const, content: 'Hello' }]
// OpenAI — native keys
chat({
adapter: openaiText('gpt-5.2'),
messages,
modelOptions: { temperature: 0.7, top_p: 0.9, max_output_tokens: 1000 },
})
// Anthropic
chat({
adapter: anthropicText('claude-sonnet-4-6'),
messages,
modelOptions: { temperature: 0.7, top_p: 0.9, max_tokens: 1000 },
})
// Gemini — camelCase
chat({
adapter: geminiText('gemini-2.5-pro'),
messages,
modelOptions: { temperature: 0.7, topP: 0.9, maxOutputTokens: 1000 },
})
// Ollama — NESTED under modelOptions.options
// (use the `family:tag` id — a bare `llama3.3` falls back to untyped options)
chat({
adapter: ollamaText('llama3.3:latest'),
messages,
modelOptions: {
options: { temperature: 0.7, top_p: 0.9, num_predict: 1000 },
},
})Per-provider sampling keys (all live inside modelOptions):
| Provider | Temperature | Nucleus | Max output tokens |
|---|---|---|---|
| OpenAI | temperature | top_p | max_output_tokens |
| Anthropic | temperature | top_p | max_tokens |
| Gemini | temperature | topP | maxOutputTokens |
| Grok (xAI) | temperature | top_p | max_output_tokens |
| Groq | temperature | top_p | max_completion_tokens |
| OpenRouter (chat) | temperature | topP | maxCompletionTokens |
| Ollama | temperature | top_p | num_predict (nested in options) |
| BytePlus | temperature | top_p | max_tokens (or max_completion_tokens) |
temperature is the one key every provider names identically; token limits and
some sampling options use provider-native names. Ollama nests all sampling under
modelOptions.options.
Anthropic
max_tokensdefault: Anthropic's API requiresmax_tokens, so the adapter always sends one. When you omitmodelOptions.max_tokens, it defaults to the selected model's full output ceiling (itsmax_output_tokensfrom model metadata — e.g. 64K for Sonnet, 128K for Opus), not a low constant.max_tokensis a ceiling, not a reservation (billing is per token generated), so leaving it unset is the right default for codegen / agentic / long-form output and avoids silentstop_reason: "max_tokens"truncation. Set it only to cap output below the model ceiling. Other providers treat token limits as optional and don't apply this flooring.
supportsCombinedToolsAndSchemaAdapters can declare an optional capability method:
import { AnthropicTextAdapter } from '@tanstack/ai-anthropic'
// The TextAdapter contract:
// supportsCombinedToolsAndSchema?: (modelOptions?: TProviderOptions) => boolean
// Subclasses override it to narrow the capability:
class LegacyPathAnthropic extends AnthropicTextAdapter<'claude-sonnet-4-6'> {
override supportsCombinedToolsAndSchema(): boolean {
return false
}
}When true, the engine wires outputSchema into the regular
chatStream call alongside tools and harvests the schema-constrained
JSON from the agent loop's final-turn text — skipping the separate
structuredOutput / structuredOutputStream finalization round-trip.
When false (or the method is omitted), the legacy finalization path
runs.
Current per-adapter status (#605):
| Adapter | Returns |
|---|---|
openaiText / openaiChatCompletions | true (all supported models) |
anthropicText | true for Claude 4.5+ (gated by ANTHROPIC_COMBINED_TOOLS_AND_SCHEMA_MODELS), false otherwise |
geminiText | true for Gemini 3.x (gated by GEMINI_COMBINED_TOOLS_AND_SCHEMA_MODELS), false otherwise |
grokText | true (all chat models — inherits the OpenAI Responses base; no per-model gate) |
groqText | false (Groq API rejects schema + tools + stream) |
openRouterText / openRouterResponsesText | Per model — true only when the model and every modelOptions.models fallback are in OPENROUTER_COMBINED_TOOLS_AND_SCHEMA_MODELS |
ollamaText | false (constrained-decoding vs tool-call grammar conflict) |
byteplusText | Per model — true only for the 10 ids in BYTEPLUS_STRUCTURED_OUTPUT_CHAT_MODELS, false otherwise |
Subclasses can override to narrow the capability. When extending an
adapter for a custom model that doesn't support the combination, return
false explicitly.
BytePlus has no JSON-mode fallback. On the 8 chat models outside
BYTEPLUS_STRUCTURED_OUTPUT_CHAT_MODELS, Ark rejectsjson_schemaandjson_object, sofalsehere does not buy a degraded path — it only keepsresponse_formatout of the streaming chat request.structuredOutput()throws andstructuredOutputStream()emitsRUN_ERRORon those models. Noteseed-2-0-lite-260428(the obvious default) is one of them; useseed-2-0-lite-260228ordola-seed-2-1-turbo-260628for typed output.
Any provider that implements the OpenAI Chat Completions API (DeepSeek,
Moonshot/Kimi, Together, Fireworks, Cerebras, Qwen/DashScope, Perplexity,
NVIDIA NIM, LM Studio, etc.) can be used through the generic
openaiCompatible factory from @tanstack/ai-openai/compatible — no
dedicated package required.
import { openaiCompatible } from '@tanstack/ai-openai/compatible'
import { chat, createModel } from '@tanstack/ai'
const messages = [{ role: 'user' as const, content: 'Hello' }]
// Provider-factory: configure baseURL + apiKey + models ONCE,
// then select a model per call (the model arg is a type-safe union).
const deepseek = openaiCompatible({
name: 'deepseek', // optional label for devtools/errors (default 'openai-compatible')
baseURL: 'https://api.deepseek.com/v1',
apiKey: process.env.DEEPSEEK_API_KEY!,
models: [
'deepseek-chat', // bare string → optimistic defaults: text/image in, streaming, tools, structured output
createModel('deepseek-reasoner', {
// rich def → precise per-model capabilities
input: ['text'],
features: ['reasoning', 'structured_outputs'],
}),
],
})
chat({ adapter: deepseek('deepseek-chat'), messages })
chat({ adapter: deepseek('deepseek-reasoner'), messages })config also accepts any OpenAI SDK ClientOptions (notably defaultHeaders
and defaultQuery) for providers that need extra auth headers or query params.
For a single model, use the one-shot helper:
import { openaiCompatibleText } from '@tanstack/ai-openai/compatible'
import { chat } from '@tanstack/ai'
const messages = [{ role: 'user' as const, content: 'Hello' }]
chat({
adapter: openaiCompatibleText('deepseek-chat', {
baseURL: 'https://api.deepseek.com/v1',
apiKey: process.env.DEEPSEEK_API_KEY!,
}),
messages,
})Pass api: 'responses' to target the OpenAI Responses API instead of Chat
Completions (only for the rare compatible provider that implements it, e.g.
Azure OpenAI); the default is 'chat-completions', which is what nearly all
compatible providers speak.
Verify the provider's current
baseURLand model ids against its live docs — they drift. Seedocs/adapters/openai-compatible.mdfor the full provider table.
Four providers expose a native Files/storage API as a tree-shakeable files
adapter: openaiFiles(), anthropicFiles(), geminiFiles() (each reads the
same env var as the provider's text adapter; create*Files(apiKey) variants
take an explicit key), and falFiles(config). Upload media once with
uploadFile(), then reference the returned FileHandle in messages via a
{ type: 'file' } content source instead of re-sending base64 each request:
import { chat, fileSourceFromHandle, uploadFile } from '@tanstack/ai'
import { openaiFiles, openaiText } from '@tanstack/ai-openai'
import { pdfBase64 } from './pdf-data'
const handle = await uploadFile({
adapter: openaiFiles(),
input: { data: pdfBase64, mimeType: 'application/pdf' },
})
chat({
adapter: openaiText('gpt-5.5'),
messages: [
{
role: 'user',
content: [
{ type: 'text', content: 'Summarize this document' },
{ type: 'document', source: fileSourceFromHandle(handle) },
],
},
],
})Rules agents must respect:
fileSourceFromHandle
builds { type: 'file', value: 'file-…', provider: 'openai' }, matching the
AG-UI FileSource arm. A handle only resolves at the provider that issued
it, so an adapter throws when provider names a different adapter.
provider is optional, as on the AG-UI wire; a source without it is taken
as-is. To use the same bytes with two providers,
upload to each and send the matching handle.supportsFileSources. For adapters that don't (Groq, Bedrock, Mistral, OpenRouter, Ollama, BytePlus, Cohere, and anything
written before this feature), chat() / generateImage() /
generateVideo() / embed() reject file sources in preflight, before any
request is built — pass data/url sources there instead.getFile() / deleteFile() work for OpenAI, Anthropic,
Gemini, and Grok, and accept the handle itself (provider-literal typed, so a
foreign handle is a compile error). fal storage is upload-only, so those
calls throw for falFiles(). grokFiles().get() mints the public URL again,
so do not call it after revokePublicUrl().images/edits + Sora input_reference, Gemini Veo, and Chat Completions
image inputs throw endpoint-specific errors for file sources.fileSourceFromHandle(handle) straight into the sendMessage content, and
the server passes the messages to chat() as usual. fileSourceFromHandle
and the FileHandle type are exported from the browser-safe
@tanstack/ai/client entry.See docs/advanced/files-api.md for the full guide.
Every adapter's client config accepts baseURL and defaultHeaders. Use these
two names to route any adapter through Cloudflare AI Gateway, Vercel AI Gateway,
or a corporate proxy. The adapter maps them onto the vendor SDK's own option
names (Gemini httpOptions, Mistral serverURL, Ollama host, Cohere and
ElevenLabs baseUrl/headers). The vendor names still work; when both are
set, baseURL and defaultHeaders win.
import { createGeminiChat } from '@tanstack/ai-gemini'
const gateway = {
baseURL: 'https://gateway.example.com/google-ai-studio',
defaultHeaders: {
'cf-aig-authorization': `Bearer ${process.env.GATEWAY_TOKEN}`,
},
}
createGeminiChat('gemini-3.8-flash', process.env.GOOGLE_API_KEY!, {
...gateway,
})The legacy openai() (and anthropic(), etc.) monolithic adapters are
deprecated. They take the model in chat(), not in the factory.
// WRONG: Legacy monolithic adapter pattern (no longer exported)
import { openai } from '@tanstack/ai-openai'
chat({ adapter: openai(), model: 'gpt-5.2', messages })// CORRECT: Tree-shakeable adapter, model in factory
import { chat } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'
const messages = [{ role: 'user' as const, content: 'Hello' }]
chat({ adapter: openaiText('gpt-5.2'), messages })Source: docs/migration/migration.md
Each provider uses a specific env var name. Using the wrong one causes a runtime error:
| Provider | Correct Env Var | Common Mistake |
|---|---|---|
| OpenAI | OPENAI_API_KEY | |
| Anthropic | ANTHROPIC_API_KEY | |
| Gemini | GOOGLE_API_KEY or GEMINI_API_KEY | GOOGLE_GENAI_API_KEY (does not work) |
| Grok (xAI) | XAI_API_KEY | GROK_API_KEY (does not work) |
| Groq | GROQ_API_KEY | |
| OpenRouter | OPENROUTER_API_KEY | |
| Ollama | OLLAMA_HOST | No API key needed, just the host URL (default: http://localhost:11434) |
| Bedrock | BEDROCK_API_KEY / AWS_BEARER_TOKEN_BEDROCK | Falls back to SigV4 credentials when no API key is set |
Source: adapter source code (utils/client.ts in each adapter package).
Detailed per-adapter reference files:
HIGH Tension: Type safety vs. quick prototyping -- Per-model type safety
requires specific model string literals. Quick prototyping wants dynamic
selection with string variables. Agents optimizing for quick setup silently
lose type safety. If model names come from user input or config files, use
extendAdapter() to add custom names.
ai-core/chat-experience/SKILL.md -- Adapter choice affects chat setupai-core/structured-outputs/SKILL.md -- outputSchema handles provider differences transparently7fb4a5f
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.