Models

Model catalog

One OpenAI-compatible API for open-source and frontier LLMs, plus Google's Gemini family for multimodal chat, reasoning, and image generation. Prices are listed per 1M tokens in USD, as of July 2026. The model ID is what you pass in the request body.

At a glance

Model IDContextInputOutputNotes
deepseek-v4-flash1M$0.112$0.224Official 0731 agent model. 1M context with Chat Completions and Responses for Codex.
deepseek-v4.1-flash1M$0.30$1.20official DeepSeek V4.1 Flash — new architecture, faster, 1M context, Responses API
deepseek-v4-pro1M$0.70$2.10Premium reasoning with 1M-token context. Maps to o3-mini at ~6× lower output cost.
qwen3.8-max1M$2.00$6.00Alibaba's Qwen 3.8 flagship, released 2026-08-03: a 2.4T-parameter MoE that Alibaba pitches at autonomous coding and long-horizon agents. Thinks by default — the gateway suppresses thinking by default; pass enable_thinking=true to opt into the reasoning trace.
qwen3.8-max-prime1M$4.00$12.00Qwen 3.8 top-tier flagship, deepest reasoning & autonomous coding, 1M context
qwen3.8-omni-flash1M$0.15$0.47Qwen's first agentic omni model: native image, audio and video understanding with tool use and reasoning, 1M context, at Alibaba's first-party price
qwen3.7-max1M$1.25$3.75Qwen 3.7 flagship. Agent-centric, coding & productivity tasks. Thinks by default — the gateway suppresses thinking by default; pass enable_thinking=true to opt into the reasoning trace.
qwen3.7-plus1M$0.256$1.024Alibaba's hosted multimodal agent flagship. Long-running coding/agent loops, 1M context. Thinks by default; pass reasoning.enabled=true to opt into the trace.
qwen3.7-flash1M$0.024$0.104Fast multimodal agent model for vision, visual coding, search, tool use, and high-volume automation. Selectable reasoning.
qwen3.6-plus1M$0.26$1.561T-parameter MoE flagship. 1M context. Top Qwen on OpenRouter by token volume. Thinks by default — the gateway suppresses thinking by default; pass reasoning.enabled=true to opt into the reasoning trace.
qwen3.6-35b262K$0.112$0.8035B/3B-active MoE. Long-context RAG and summarization with strong reasoning.
qwen3.8-27b1M$0.34$2.04Qwen's fast open 27B reasoning model with a 1M-token context window
qwen3.8-flash-next1M$0.12$0.376Qwen's open Qwen4-architecture preview: 6B-active MoE, fast and cheap, with a 1M-token context window
kimi-k2.6256K$0.5472$2.728Opus-class agentic / planning. Best fit when your eval picks Claude Opus.
kimi-k2.7-code256K$0.584$2.80Always-thinking K2 tuned for long-horizon agentic coding. Fixed sampling parameters; Moonshot reports ~30% less overthinking than K2.6.
kimi-k31M$2.55$12.75Moonshot's 2.8T open-weight multimodal reasoning flagship. Complex coding, knowledge work, and long-horizon agentic workflows.
muse-spark-1.31M$1.00$3.40Coding & agentic reasoning model, 1M context
muse-spark-1.21M$1.00$3.40Coding-focused reasoning model, 1M context
muse-glimmer-30b131,072$0.28$1.20Compact agentic multimodal model, distilled from Muse Spark
glm-5.31M$1.12$3.52Z.ai latest reasoning flagship, 1M context, complex software engineering & long-horizon agents
glm-5.3-prime1M$2.24$7.04Z.ai GLM 5.3 Prime, top reasoning flagship, 1M context, complex software engineering
glm-5.3-flash1M$0.06$0.20Z.ai fast open-weights reasoning model, 1M context at workhorse pricing
glm-5.21M$1.12$3.52Z.ai's large-scale reasoning flagship. 1M context, built for long-horizon agent workflows and project-level software engineering. Thinks by default — the gateway suppresses thinking by default; pass reasoning.enabled=true to opt into the reasoning trace.
nemotron-3.5-lightning262K$0.08$0.2030B-A3B open MoE for fast, tool-heavy agents
gpt-oss-120b131,072$0.12$0.48OpenAI's open-weight 117B MoE, high-volume agents & reasoning, 131K context
gpt-6-astra1M$10.00 ($20.00 >200K)$50.00 ($75.00 >200K)OpenAI's GPT-6 flagship. Long-horizon agentic coding, deep research and document work; accepts image input.
gpt-6-luna1M$0.10 ($0.20 >200K)$0.50 ($0.75 >200K)OpenAI's cost-efficient GPT-6, high-volume chat & lightweight agentic work, 1M context
gpt-6.1-sol1M$2.00 ($4.00 >200K)$10.00 ($15.00 >200K)OpenAI's GPT-6.1 Sol, near-Astra agentic coding at Sol pricing, 1M context
gpt-6-sol1M$2.00 ($4.00 >200K)$10.00 ($15.00 >200K)OpenAI's GPT-6 mid-flagship, strong reasoning & agentic coding, 1M context
gpt-5.6-luna1M$0.16$0.96OpenAI's fast, cost-efficient GPT-5.6 tier. High-volume, latency-sensitive chat and classification. Reasoning model; pass reasoning.enabled=false to suppress the trace.
gpt-5.6-terra1M$1.60$9.60OpenAI's balanced GPT-5.6 mid-tier, between Luna and Sol. Everyday coding, reasoning and agentic work.
gpt-5.6-sol1M$1.60$8.00OpenAI's flagship GPT-5.6. Complex reasoning, coding and agentic workflows, especially on hard multi-step tasks.
grok-4.5500K$1.60$4.80xAI's frontier model for coding, knowledge work and STEM. Always reasons — reasoning.enabled=false is rejected; control depth with reasoning_effort.
grok-4.6500K$2.00 ($4.00 >200K)$6.00 ($12.00 >200K)xAI's frontier model for coding, knowledge work, STEM and image analysis; accepts image input. Always reasons — reasoning.enabled=false is rejected; control depth with reasoning_effort.
grok-4.7500K$2.00 ($4.00 >200K)$6.00 ($12.00 >200K)xAI's flagship for coding, agentic tasks and knowledge work, succeeding Grok 4.6. Built for long-running software engineering and self-verification; accepts image input. Always reasons — reasoning.enabled=false is rejected; control depth with reasoning_effort.
minimax-m31M$0.24$0.96MiniMax's open-weight frontier model. Text-only input, tuned for long-horizon agentic coding.
mimo-v2.6-pro1M$0.348$0.696Xiaomi's flagship open-weight MiMo-V2.6-Pro (1.02T MoE, 42B active) — top open-weights model on the Artificial Analysis Intelligence Index. Coding and agentic workflows; accepts image input. Runs in non-thinking mode (reasoning off).
mimo-v2.6-flash1M$0.112$0.224Xiaomi's fast, low-cost open-weight MiMo-V2.6-Flash (309B MoE, 15B active). Everyday coding and agents; accepts image input. Runs in non-thinking mode (reasoning off).
mimo-v2.51M$0.112$0.224Xiaomi's open-weight MiMo-V2.5. Cost-efficient everyday coding and agentic workflows.
hy4-preview1M$0.6672$2.0008Tencent's Hy4 (preview). Successor to Hy3 for coding and agentic workflows with reasoning.
hy3262K$0.1056$0.4224Tencent's Hy3. General-purpose coding and agentic workflows with selectable reasoning.
jev-1.1332K$0.042FreeTypeSafe's System One model on POST /v1/systemone — not a chat model. Answers typed choice / score / noul questions about a JSON state with calibrated probabilities. Bills input only.
claude-opus-5-51M$3.20$16.00Anthropic's most capable Claude and the first of the Claude 5.5 family. Leads Opus 5 and Fable 5.1 on agentic coding, knowledge work and computer use, at a lower per-token price than Opus 5. Accepts image input. Chooses its own tools: tool_choice=auto or tool_choice=none only — forcing one specific function is not supported.
claude-opus-51M$4.00$20.00Anthropic's Opus 5 for demanding reasoning, coding and long-horizon agentic work. Accepts image input.
claude-fable-5-11M$8.00$40.00Anthropic's Mythos-class model, version 5.1, for deep reasoning and long-horizon agentic work.
claude-fable-51M$8.00$40.00Anthropic's Mythos-class model for deep reasoning and long-horizon agentic work.
claude-opus-4-81M$4.00$20.00Anthropic flagship. Top-tier reasoning, coding and agentic workloads.
claude-sonnet-5-51M$2.00$10.00Anthropic's newest Sonnet, faster and fewer tokens per task, 1M context
claude-sonnet-51M$2.00$10.00Anthropic's newest Sonnet. Stronger reasoning and coding at a lower price than the previous mid-tier.
claude-haiku-4-5200K$0.80$4.00Fast, low-cost Anthropic tier for high-volume work. Does not emit a reasoning trace.
gemini-3.7-flash1M$0.6375$3.1875fast multimodal Flash for agents; promotional pricing through 2026-12-31
gemini-3.8-flash1M$0.6375$3.1875most capable Flash for agents; promotional pricing through 2026-12-31
gemini-3.6-flash1M$1.275$6.375Current general-purpose Flash GA. Recommended for new Flash integrations.
gemini-3.5-flash-lite1M$0.255$2.125Current low-cost Flash-Lite GA. Recommended for high-volume workloads.
gemini-3.5-flash1M$1.275$7.65Next-gen Flash GA from Google. 1M context, thinks by default. Sits between 3 Flash Preview and 3.1 Pro on capability and price.
gemini-3.1-pro-preview1M$1.70$10.20Google's flagship reasoning model. 1M context, thinks deeply. Preview API; semantics may shift before GA.
gemini-3-pro-image1M$1.70$10.20GA pro-grade image generation. Replaces the retired preview model ID.
gemini-3-flash-preview1M$0.425$2.55Legacy compatibility only. Migrate new and existing workloads to gemini-3.6-flash.
gemini-3.1-flash-lite1M$0.2125$1.275Legacy compatibility only. Migrate new and existing workloads to gemini-3.5-flash-lite.
flux.2-pro-—$0.027 per imageFlagship image generation via /v1/images/generations. Billed per generated image.
flux.1-schnell-—$0.003 per imageultra-fast open image generation, billed per image
sdxl-turbo-—$0.003 per imagefast open image generation (SDXL Turbo), billed per image
flux.2-klein-—$0.02 per imageopen FLUX.2 image generation, billed per image
qwen-image-max-—$0.1 per imageAlibaba's Qwen-Image flagship: high-fidelity generation with strong typography, billed per image
seedream-5.0-pro-—$0.07 per imageByteDance's Seedream 5.0 Pro: flagship photorealistic image generation, billed per image
seedream-4-—$0.055 per imageByteDance Seedream 4: fast, high-quality image generation, billed per image
bria-fibo-1.5-—$0.055 per imageBria FIBO 1.5: image generation trained on fully licensed data (commercially safe), billed per image

Which model should I use?

  • Code-first default - deepseek-v4-flash. A practical starting point for code, chat, and production agents. 1M context; input $0.112, output $0.224.
  • Hard reasoning - deepseek-v4-pro. Use for math, theorem proving, and difficult multi-step problems. 1M context; input $0.70, output $2.10.
  • Long-document RAG - qwen3.6-35b. A compact MoE choice for retrieval and long-document summarization. 262K context; input $0.112, output $0.80.
  • Agentic planning - kimi-k2.6. A strong option when your evaluations favor Opus-class planning behavior. 256K context; input $0.5472, output $2.728.
  • Multimodal workloads - gemini-3.6-flash. Use when prompts include images as well as text. 1M context; input $1.275, output $6.375.
  • Image generation - gemini-3-pro-image. Use for pro-grade generated image output. 1M context; input $1.70, output $10.20.
  • High-volume short turns - gemini-3.5-flash-lite. Use for cost-sensitive routing, classification, and extraction. 1M context; input $0.255, output $2.125.
  • Routing & classification decisions - jev-1.13. A System One model, not chat: typed answers with calibrated probabilities on POST /v1/systemone. 32K context; input $0.042, output free.

Thinking vs. non-thinking

The V4 wave (V4 Flash, V4 Pro, Kimi K2.6) and Qwen 3.6 emit a chain-of-thought trace before the final answer. A one-token "Hi" can return ~175 reasoning tokens. To get non-thinking low-cost chat behavior on these models, pass:

See quickstart for full code →

Some always-reasoning models (e.g. Grok 4.5) ignore reasoning.enabled=false— reasoning IS the model. Use V4 Flash if you don't want the trace.

Claude models: structured output and tool choice

Claude models do not enforce response_format: {type: "json_schema"} — the request succeeds and the reply comes back as ordinary prose. When you need schema-shaped output from a Claude model, declare the schema as a tool with strict: true and read the tool-call arguments; that path is enforced.

claude-opus-5-5 also chooses its own tools: tool_choice accepts "auto" or "none", and forcing one specific function is not supported on that model. The other Claude models accept a forced tool choice.

System One models (typed decisions)

System One models — Jev 1.13 today — are not chat models. You send a JSON state plus typed questions (choice, score or noul) to POST /v1/systemone and get calibrated probabilities back from a single forward pass, with no generated text to parse. Use them for routing, classification, triage, scoring and guardrail decisions. They bill input tokens only — output is free — and a Chat Completions call to one is rejected.

System One docs →

Per-model deep dives