Model catalog
One OpenAI-compatible API for open-source and frontier LLMs, plus Google's Gemini family for multimodal chat, reasoning, and image generation. Prices are listed per 1M tokens in USD, as of July 2026. The model ID is what you pass in the request body.
At a glance
| Model ID | Context | Input | Output | Notes |
|---|---|---|---|---|
| deepseek-v4-flash | 1M | $0.112 | $0.224 | Official 0731 agent model. 1M context with Chat Completions and Responses for Codex. |
| deepseek-v4.1-flash | 1M | $0.30 | $1.20 | official DeepSeek V4.1 Flash — new architecture, faster, 1M context, Responses API |
| deepseek-v4-pro | 1M | $0.70 | $2.10 | Premium reasoning with 1M-token context. Maps to o3-mini at ~6× lower output cost. |
| qwen3.8-max | 1M | $2.00 | $6.00 | Alibaba's Qwen 3.8 flagship, released 2026-08-03: a 2.4T-parameter MoE that Alibaba pitches at autonomous coding and long-horizon agents. Thinks by default — the gateway suppresses thinking by default; pass enable_thinking=true to opt into the reasoning trace. |
| qwen3.8-max-prime | 1M | $4.00 | $12.00 | Qwen 3.8 top-tier flagship, deepest reasoning & autonomous coding, 1M context |
| qwen3.8-omni-flash | 1M | $0.15 | $0.47 | Qwen's first agentic omni model: native image, audio and video understanding with tool use and reasoning, 1M context, at Alibaba's first-party price |
| qwen3.7-max | 1M | $1.25 | $3.75 | Qwen 3.7 flagship. Agent-centric, coding & productivity tasks. Thinks by default — the gateway suppresses thinking by default; pass enable_thinking=true to opt into the reasoning trace. |
| qwen3.7-plus | 1M | $0.256 | $1.024 | Alibaba's hosted multimodal agent flagship. Long-running coding/agent loops, 1M context. Thinks by default; pass reasoning.enabled=true to opt into the trace. |
| qwen3.7-flash | 1M | $0.024 | $0.104 | Fast multimodal agent model for vision, visual coding, search, tool use, and high-volume automation. Selectable reasoning. |
| qwen3.6-plus | 1M | $0.26 | $1.56 | 1T-parameter MoE flagship. 1M context. Top Qwen on OpenRouter by token volume. Thinks by default — the gateway suppresses thinking by default; pass reasoning.enabled=true to opt into the reasoning trace. |
| qwen3.6-35b | 262K | $0.112 | $0.80 | 35B/3B-active MoE. Long-context RAG and summarization with strong reasoning. |
| qwen3.8-27b | 1M | $0.34 | $2.04 | Qwen's fast open 27B reasoning model with a 1M-token context window |
| qwen3.8-flash-next | 1M | $0.12 | $0.376 | Qwen's open Qwen4-architecture preview: 6B-active MoE, fast and cheap, with a 1M-token context window |
| kimi-k2.6 | 256K | $0.5472 | $2.728 | Opus-class agentic / planning. Best fit when your eval picks Claude Opus. |
| kimi-k2.7-code | 256K | $0.584 | $2.80 | Always-thinking K2 tuned for long-horizon agentic coding. Fixed sampling parameters; Moonshot reports ~30% less overthinking than K2.6. |
| kimi-k3 | 1M | $2.55 | $12.75 | Moonshot's 2.8T open-weight multimodal reasoning flagship. Complex coding, knowledge work, and long-horizon agentic workflows. |
| muse-spark-1.3 | 1M | $1.00 | $3.40 | Coding & agentic reasoning model, 1M context |
| muse-spark-1.2 | 1M | $1.00 | $3.40 | Coding-focused reasoning model, 1M context |
| muse-glimmer-30b | 131,072 | $0.28 | $1.20 | Compact agentic multimodal model, distilled from Muse Spark |
| glm-5.3 | 1M | $1.12 | $3.52 | Z.ai latest reasoning flagship, 1M context, complex software engineering & long-horizon agents |
| glm-5.3-prime | 1M | $2.24 | $7.04 | Z.ai GLM 5.3 Prime, top reasoning flagship, 1M context, complex software engineering |
| glm-5.3-flash | 1M | $0.06 | $0.20 | Z.ai fast open-weights reasoning model, 1M context at workhorse pricing |
| glm-5.2 | 1M | $1.12 | $3.52 | Z.ai's large-scale reasoning flagship. 1M context, built for long-horizon agent workflows and project-level software engineering. Thinks by default — the gateway suppresses thinking by default; pass reasoning.enabled=true to opt into the reasoning trace. |
| nemotron-3.5-lightning | 262K | $0.08 | $0.20 | 30B-A3B open MoE for fast, tool-heavy agents |
| gpt-oss-120b | 131,072 | $0.12 | $0.48 | OpenAI's open-weight 117B MoE, high-volume agents & reasoning, 131K context |
| gpt-6-astra | 1M | $10.00 ($20.00 >200K) | $50.00 ($75.00 >200K) | OpenAI's GPT-6 flagship. Long-horizon agentic coding, deep research and document work; accepts image input. |
| gpt-6-luna | 1M | $0.10 ($0.20 >200K) | $0.50 ($0.75 >200K) | OpenAI's cost-efficient GPT-6, high-volume chat & lightweight agentic work, 1M context |
| gpt-6.1-sol | 1M | $2.00 ($4.00 >200K) | $10.00 ($15.00 >200K) | OpenAI's GPT-6.1 Sol, near-Astra agentic coding at Sol pricing, 1M context |
| gpt-6-sol | 1M | $2.00 ($4.00 >200K) | $10.00 ($15.00 >200K) | OpenAI's GPT-6 mid-flagship, strong reasoning & agentic coding, 1M context |
| gpt-5.6-luna | 1M | $0.16 | $0.96 | OpenAI's fast, cost-efficient GPT-5.6 tier. High-volume, latency-sensitive chat and classification. Reasoning model; pass reasoning.enabled=false to suppress the trace. |
| gpt-5.6-terra | 1M | $1.60 | $9.60 | OpenAI's balanced GPT-5.6 mid-tier, between Luna and Sol. Everyday coding, reasoning and agentic work. |
| gpt-5.6-sol | 1M | $1.60 | $8.00 | OpenAI's flagship GPT-5.6. Complex reasoning, coding and agentic workflows, especially on hard multi-step tasks. |
| grok-4.5 | 500K | $1.60 | $4.80 | xAI's frontier model for coding, knowledge work and STEM. Always reasons — reasoning.enabled=false is rejected; control depth with reasoning_effort. |
| grok-4.6 | 500K | $2.00 ($4.00 >200K) | $6.00 ($12.00 >200K) | xAI's frontier model for coding, knowledge work, STEM and image analysis; accepts image input. Always reasons — reasoning.enabled=false is rejected; control depth with reasoning_effort. |
| grok-4.7 | 500K | $2.00 ($4.00 >200K) | $6.00 ($12.00 >200K) | xAI's flagship for coding, agentic tasks and knowledge work, succeeding Grok 4.6. Built for long-running software engineering and self-verification; accepts image input. Always reasons — reasoning.enabled=false is rejected; control depth with reasoning_effort. |
| minimax-m3 | 1M | $0.24 | $0.96 | MiniMax's open-weight frontier model. Text-only input, tuned for long-horizon agentic coding. |
| mimo-v2.6-pro | 1M | $0.348 | $0.696 | Xiaomi's flagship open-weight MiMo-V2.6-Pro (1.02T MoE, 42B active) — top open-weights model on the Artificial Analysis Intelligence Index. Coding and agentic workflows; accepts image input. Runs in non-thinking mode (reasoning off). |
| mimo-v2.6-flash | 1M | $0.112 | $0.224 | Xiaomi's fast, low-cost open-weight MiMo-V2.6-Flash (309B MoE, 15B active). Everyday coding and agents; accepts image input. Runs in non-thinking mode (reasoning off). |
| mimo-v2.5 | 1M | $0.112 | $0.224 | Xiaomi's open-weight MiMo-V2.5. Cost-efficient everyday coding and agentic workflows. |
| hy4-preview | 1M | $0.6672 | $2.0008 | Tencent's Hy4 (preview). Successor to Hy3 for coding and agentic workflows with reasoning. |
| hy3 | 262K | $0.1056 | $0.4224 | Tencent's Hy3. General-purpose coding and agentic workflows with selectable reasoning. |
| jev-1.13 | 32K | $0.042 | Free | TypeSafe's System One model on POST /v1/systemone — not a chat model. Answers typed choice / score / noul questions about a JSON state with calibrated probabilities. Bills input only. |
| claude-opus-5-5 | 1M | $3.20 | $16.00 | Anthropic's most capable Claude and the first of the Claude 5.5 family. Leads Opus 5 and Fable 5.1 on agentic coding, knowledge work and computer use, at a lower per-token price than Opus 5. Accepts image input. Chooses its own tools: tool_choice=auto or tool_choice=none only — forcing one specific function is not supported. |
| claude-opus-5 | 1M | $4.00 | $20.00 | Anthropic's Opus 5 for demanding reasoning, coding and long-horizon agentic work. Accepts image input. |
| claude-fable-5-1 | 1M | $8.00 | $40.00 | Anthropic's Mythos-class model, version 5.1, for deep reasoning and long-horizon agentic work. |
| claude-fable-5 | 1M | $8.00 | $40.00 | Anthropic's Mythos-class model for deep reasoning and long-horizon agentic work. |
| claude-opus-4-8 | 1M | $4.00 | $20.00 | Anthropic flagship. Top-tier reasoning, coding and agentic workloads. |
| claude-sonnet-5-5 | 1M | $2.00 | $10.00 | Anthropic's newest Sonnet, faster and fewer tokens per task, 1M context |
| claude-sonnet-5 | 1M | $2.00 | $10.00 | Anthropic's newest Sonnet. Stronger reasoning and coding at a lower price than the previous mid-tier. |
| claude-haiku-4-5 | 200K | $0.80 | $4.00 | Fast, low-cost Anthropic tier for high-volume work. Does not emit a reasoning trace. |
| gemini-3.7-flash | 1M | $0.6375 | $3.1875 | fast multimodal Flash for agents; promotional pricing through 2026-12-31 |
| gemini-3.8-flash | 1M | $0.6375 | $3.1875 | most capable Flash for agents; promotional pricing through 2026-12-31 |
| gemini-3.6-flash | 1M | $1.275 | $6.375 | Current general-purpose Flash GA. Recommended for new Flash integrations. |
| gemini-3.5-flash-lite | 1M | $0.255 | $2.125 | Current low-cost Flash-Lite GA. Recommended for high-volume workloads. |
| gemini-3.5-flash | 1M | $1.275 | $7.65 | Next-gen Flash GA from Google. 1M context, thinks by default. Sits between 3 Flash Preview and 3.1 Pro on capability and price. |
| gemini-3.1-pro-preview | 1M | $1.70 | $10.20 | Google's flagship reasoning model. 1M context, thinks deeply. Preview API; semantics may shift before GA. |
| gemini-3-pro-image | 1M | $1.70 | $10.20 | GA pro-grade image generation. Replaces the retired preview model ID. |
| gemini-3-flash-preview | 1M | $0.425 | $2.55 | Legacy compatibility only. Migrate new and existing workloads to gemini-3.6-flash. |
| gemini-3.1-flash-lite | 1M | $0.2125 | $1.275 | Legacy compatibility only. Migrate new and existing workloads to gemini-3.5-flash-lite. |
| flux.2-pro | - | — | $0.027 per image | Flagship image generation via /v1/images/generations. Billed per generated image. |
| flux.1-schnell | - | — | $0.003 per image | ultra-fast open image generation, billed per image |
| sdxl-turbo | - | — | $0.003 per image | fast open image generation (SDXL Turbo), billed per image |
| flux.2-klein | - | — | $0.02 per image | open FLUX.2 image generation, billed per image |
| qwen-image-max | - | — | $0.1 per image | Alibaba's Qwen-Image flagship: high-fidelity generation with strong typography, billed per image |
| seedream-5.0-pro | - | — | $0.07 per image | ByteDance's Seedream 5.0 Pro: flagship photorealistic image generation, billed per image |
| seedream-4 | - | — | $0.055 per image | ByteDance Seedream 4: fast, high-quality image generation, billed per image |
| bria-fibo-1.5 | - | — | $0.055 per image | Bria FIBO 1.5: image generation trained on fully licensed data (commercially safe), billed per image |
Which model should I use?
- Code-first default -
deepseek-v4-flash. A practical starting point for code, chat, and production agents. 1M context; input $0.112, output $0.224. - Hard reasoning -
deepseek-v4-pro. Use for math, theorem proving, and difficult multi-step problems. 1M context; input $0.70, output $2.10. - Long-document RAG -
qwen3.6-35b. A compact MoE choice for retrieval and long-document summarization. 262K context; input $0.112, output $0.80. - Agentic planning -
kimi-k2.6. A strong option when your evaluations favor Opus-class planning behavior. 256K context; input $0.5472, output $2.728. - Multimodal workloads -
gemini-3.6-flash. Use when prompts include images as well as text. 1M context; input $1.275, output $6.375. - Image generation -
gemini-3-pro-image. Use for pro-grade generated image output. 1M context; input $1.70, output $10.20. - High-volume short turns -
gemini-3.5-flash-lite. Use for cost-sensitive routing, classification, and extraction. 1M context; input $0.255, output $2.125. - Routing & classification decisions -
jev-1.13. A System One model, not chat: typed answers with calibrated probabilities on POST /v1/systemone. 32K context; input $0.042, output free.
Thinking vs. non-thinking
The V4 wave (V4 Flash, V4 Pro, Kimi K2.6) and Qwen 3.6 emit a chain-of-thought trace before the final answer. A one-token "Hi" can return ~175 reasoning tokens. To get non-thinking low-cost chat behavior on these models, pass:
Some always-reasoning models (e.g. Grok 4.5) ignore reasoning.enabled=false— reasoning IS the model. Use V4 Flash if you don't want the trace.
Claude models: structured output and tool choice
Claude models do not enforce response_format: {type: "json_schema"} — the request succeeds and the reply comes back as ordinary prose. When you need schema-shaped output from a Claude model, declare the schema as a tool with strict: true and read the tool-call arguments; that path is enforced.
claude-opus-5-5 also chooses its own tools: tool_choice accepts "auto" or "none", and forcing one specific function is not supported on that model. The other Claude models accept a forced tool choice.
System One models (typed decisions)
System One models — Jev 1.13 today — are not chat models. You send a JSON state plus typed questions (choice, score or noul) to POST /v1/systemone and get calibrated probabilities back from a single forward pass, with no generated text to parse. Use them for routing, classification, triage, scoring and guardrail decisions. They bill input tokens only — output is free — and a Chat Completions call to one is rejected.
Per-model deep dives
Each model has its own page with pricing comparisons, FAQs, and quickstart code. Linked here for convenience:
- DeepSeek V4 Flash
- DeepSeek V4.1 Flash
- DeepSeek V4 Pro
- Qwen3.8 Max
- Qwen3.8 Omni Flash
- Qwen3.7 Max
- Qwen3.7 Plus
- Qwen3.7 Flash
- Qwen3.6 Plus
- Qwen3.6-35B-A3B
- Qwen3.8 27B
- Qwen3.8 Flash Next
- Kimi K2.6
- Kimi K2.7 Code
- Kimi K3
- Muse Spark 1.3
- Muse Spark 1.2
- Muse Glimmer 30B
- GLM 5.3
- GLM 5.3 Flash
- GLM 5.2
- Nemotron 3.5 Lightning
- GPT-OSS 120B
- GPT-6 Astra
- Grok 4.6
- Grok 4.7
- MiniMax M3
- MiMo-V2.6-Pro
- MiMo-V2.6-Flash
- MiMo-V2.5
- Hy4 Preview
- Hy3
- Jev 1.13
- Claude Opus 5.5
- Gemini 3.7 Flash
- Gemini 3.8 Flash
- Gemini 3.6 Flash
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash
- Gemini 3.1 Pro Preview
- Gemini 3 Flash Preview
- Gemini 3.1 Flash Lite
- FLUX.2 Pro
- FLUX.1 Schnell
- SDXL Turbo
- FLUX.2 Klein
- Qwen-Image Max
- Seedream 5.0 Pro
- Seedream 4
- Bria FIBO 1.5
Per-model deep-dive pages for the Gemini family are in progress — the catalog table above carries the canonical pricing and capability summary in the meantime.