Origin Docs

Language Models

Text generation and chat models on ORGN Gateway — TEE and ZDR execution types and the AI SDK chatModel() method.

Language models handle text generation and chat: single-prompt completions, multi-turn conversations, reasoning, and tool use. They are the most common model type on Gateway.

When to Use

Use a language model when you need to generate or transform text:

  • Chat assistants and multi-turn conversations
  • Single-prompt generation, summarization, and rewriting
  • Reasoning tasks (use a reasoning-capable model with reasoningEffort)
  • Code generation and analysis
  • Tool calling and structured output

For image input, see Vision. For turning speech into text, see Audio.

AI SDK Method

Language models are accessed with chatModel() and used with the AI SDK's generateText and streamText:

language-model.ts
import { createOLLM } from '@orgn/gateway';
import { generateText } from 'ai';

const ollm = createOLLM({ apiKey: process.env.OLLM_API_KEY });

const { text } = await generateText({
  model: ollm.chatModel('near_glm_5_1'),
  prompt: 'What is ORGN Gateway?',
});

For streaming, system messages, multi-turn conversations, and reasoning options, see the Vercel AI SDK integration.

The legacy /v1/completions endpoint is not supported. Every completion task can be expressed as a chat call with chatModel().

TEE Catalog

Language models running in Trusted Execution Environments, on NEAR and Phala infrastructure with Intel TDX + NVIDIA H100 confidential compute, or on Tinfoil infrastructure with AMD SEV-SNP + NVIDIA confidential compute. Every request produces a cryptographic attestation receipt.

ModelProviderInfrastructureContext
DeepSeek V3.1DeepSeeknear128K
DeepSeek V3.1DeepSeekphala164K
GLM 4.7ZAInear205K
GLM 4.7ZAIphala203K
GLM 4.7 FlashZAIphala203K
GLM 5ZAInear203K
GLM 5.1ZAInear203K
Kimi K2.5Moonshotphala262K
GPT-OSS 120BOpenAInear131K
GPT-OSS 120BOpenAIphala131K
GPT-OSS 20BOpenAIphala131K
Qwen3 30BAlibabanear262K
Qwen3 30BAlibabaphala262K
Qwen 2.5 7BAlibabaphala32K
Qwen2.5 7B InstructAlibabaphala33K
Qwen3.5 122BAlibabanear131K
Qwen3.5 27BAlibabaphala262K
Venice Uncensored 24BVenicephala33K
Gemma 3 27BGooglephala53K
Llama 3.3 70BMetaphala131K
Gemma4 31BGoogletinfoil256K
GLM 5.2ZAItinfoil384K
GPT-OSS 120BOpenAItinfoil131K
GPT-OSS Safeguard 120BOpenAItinfoil131K
Kimi K2.6Moonshottinfoil256K
Llama 3.3 70BMetatinfoil128K
Voxtral Small 24BMistraltinfoil32K

ZDR Catalog

Language models running on Vercel's AI infrastructure with zero data retention provider agreements. No attestation receipts are generated.

Anthropic

ModelContext
Claude 3 Haiku200K
Claude 3.5 Haiku200K
Claude 3.7 Sonnet200K
Claude Haiku 4.5200K
Claude Sonnet 41M
Claude Sonnet 4.51M
Claude Sonnet 4.61M
Claude Opus 4200K
Claude Opus 4.1200K
Claude Opus 4.5200K
Claude Opus 4.61M
Claude Opus 4.71M
Claude Opus 4.7 (Fast)1M
Claude Opus 4.81M
Claude Opus 4.8 (Fast)1M
Claude Opus 51M
Claude Opus 5 (Fast)1M
Claude Sonnet 51M

OpenAI

ModelContext
GPT-4o8K
GPT-4o mini8K
GPT-4.18K
GPT-4.1 mini8K
GPT-4.1 nano1M
GPT-5400K
GPT-5 mini400K
GPT-5 nano400K
GPT-5 Codex400K
GPT-5.1 Instant128K
GPT-5.1-Codex400K
GPT 5 Chat128K
GPT 5.1 Codex Max400K
GPT 5.1 Codex Mini400K
GPT 5.1 Thinking400K
GPT 5.2400K
GPT 5.2 Chat128K
GPT 5.2 Codex400K
GPT 5.3 Codex400K
GPT 5.41.1M
GPT 5.4 Mini400K
GPT 5.4 Nano400K
GPT 5.4 Pro1.1M
GPT 5.51M
GPT 5.6 Luna1.1M
GPT 5.6 Sol1.1M
GPT 5.6 Terra1.1M
GPT-Realtime-1.5
GPT-Realtime-2
GPT-Realtime mini
GPT-OSS 20B131K
GPT-OSS 120B131K
GPT OSS Safeguard 20B131K
o1200K
o3-mini
o4-mini

Google

ModelContext
Gemini 2.0 Flash1M
Gemini 2.0 Flash-Lite1M
Gemini 2.5 Flash-Lite1M
Gemini 2.5 Flash1M
Gemini 2.5 Pro1M
Gemini 3 Flash1M
Gemini 3 Pro Preview1M
Gemini 3.1 Flash Lite Preview1M
Gemini 3.1 Pro Preview1M
Gemini 3.5 Flash1M
Gemini 3.5 Flash Lite1M
Gemini 3.6 Flash1M
Gemma 4 26B A4B IT262K
Gemma 4 31B IT262K

xAI

ModelContext
Grok 4.1 Fast Reasoning1M
Grok 4.1 Fast Non-Reasoning1M
Grok 4.20 Reasoning2M
Grok 4.20 Non-Reasoning2M
Grok 4.31M
Grok STT
Grok TTS
Grok Voice Think Fast 1.0

Meta

ModelContext
Llama 3.1 8B131K
Llama 3.1 70B131K
Llama 3.2 1B128K
Llama 3.2 3B128K
Llama 3.3 70B128K
Llama 4 Scout131K
Llama 4 Maverick524K

Mistral

ModelContext
Mistral Small32K
Mistral Medium128K
Mistral Large 3256K
Mistral Nemo131K
Mistral Medium Latest256K
Ministral 3B128K
Ministral 8B128K
Ministral 14B256K
Mixtral MoE 8x22B Instruct66K
Magistral Small128K
Magistral Medium128K
Codestral128K
Devstral 2256K
Devstral Small128K
Devstral Small 2256K

Alibaba (Qwen)

ModelContext
Qwen 3 14B41K
Qwen 3 30B41K
Qwen 3 32B131K
Qwen 3 235B131K
Qwen3 235B Thinking262K
Qwen3 Coder262K
Qwen3 Coder 30B262K
Qwen3 Coder Next256K
Qwen3 Next 80B262K
Qwen 3.6 Plus1M

DeepSeek

ModelContext
DeepSeek R1164K
DeepSeek V3164K
DeepSeek V3.1164K
DeepSeek V3.2164K
DeepSeek V4 Flash1M
DeepSeek V4 Flash 07311M
DeepSeek V4 Pro1M

Moonshot

ModelContext
Kimi K2131K
Kimi K2 Turbo256K
Kimi K2 0905256K
Kimi K2 Thinking262K
Kimi K2 Thinking Turbo262K
Kimi K2.5262K
Kimi K2.6262K
Kimi K2.7 Code256K
Kimi K2.7 Code High Speed262K
Kimi K31M
Kimi K3 Fast1M

ZAI

ModelContext
GLM 4.6205K
GLM 4.7205K
GLM 4.7 Flash200K
GLM 5203K
GLM 5.1203K
GLM 5.21M
GLM 5.2 Fast1M

Other Language Models

ModelProviderContext
MiniMax M2.1MiniMax205K
MiniMax M2.5MiniMax205K
Minimax M2.7MiniMax205K
MiniMax M3MiniMax1M
Morph V3 FastMorph82K
Morph V3 LargeMorph82K
INTELLECT 3PrimeIntellect131K
Nemotron 3 Nano 30BNVIDIA262K
Nemotron Nano 9B v2NVIDIA131K
NVIDIA Nemotron 3 Super 120B A12BNVIDIA256K
Nemotron 3 UltraNVIDIA1M
Nova 2 LiteAmazon1M
Nova LiteAmazon300K
Nova MicroAmazon128K
Nova ProAmazon300K
Mercury Coder Small BetaInception32K
StepFun 3.5 FlashStepFun262K
Hy3Tencent256K
InklingThinking Machines256K
Inkling SmallThinking Machines1M
MiMo M2.5Xiaomi1M
MiMo V2.5 ProXiaomi1M

Several models in this catalog also accept image input. Models with vision capability are listed on the Vision page.

Not confidential

A small set of catalog entries run through Vercel's infrastructure but carry neither a TEE attestation nor a ZDR agreement. See Models overview for what this tier does and does not guarantee.

ModelProviderContext
Claude Fable 5Anthropic1M
DeepSeek V4 Pro 0813DeepSeek1M
Fugu UltraSakana1M
Grok 4.5xAI500K
Grok 4.6xAI500K
Kat Coder Pro V2.5KwaiPilot256K
Laguna S 2.1Poolside1M

On this page