मॉडल

मॉडल catalog

एक OpenAI-compatible API: open-source LLMs (DeepSeek V4, Qwen — 3.7 Max और 3.6 Plus सहित, Kimi, GLM) और Google का Gemini family — multimodal chat, reasoning तथा image generation के लिए। सभी prices per 1M tokens, USD में, July 2026 के अनुसार। Request body में model field में जो pass करते हैं वो है Model ID।

एक नज़र में

Model IDContextInputOutputNotes
deepseek-v4-flash1M$0.086$0.173सस्ता chat और coding workhorse। Default में सोचता है — V3-style replies के लिए reasoning.enabled=false पास करें।
deepseek-v4.1-flash1M$0.125$0.55official DeepSeek V4.1 Flash — new architecture, faster, 1M context, Responses API
deepseek-v4-pro1M$0.70$2.10Premium reasoning + 1M token context। o3-mini के बराबर quality, output cost लगभग 6× कम।
qwen3.8-max1M$2.00$6.00Alibaba का Qwen 3.8 flagship, 2026-08-03 को release हुआ: 2.4T-parameter MoE जिसे Alibaba autonomous coding और long-horizon agents के लिए position करता है। Default में सोचता है — gateway default रूप से thinking दबा देता है; reasoning trace चाहिए तो enable_thinking=true पास करें।
qwen3.8-max-prime1M$4.00$12.00Qwen 3.8 top-tier flagship, deepest reasoning & autonomous coding, 1M context
qwen3.8-omni-flash1M$0.15$0.47Qwen's first agentic omni model: native image, audio and video understanding with tool use and reasoning, 1M context, at Alibaba's first-party price
qwen3.7-max1M$1.25$3.75Qwen 3.7 flagship। Agent-centric workloads, coding और productivity tasks में strong। Default में सोचता है — gateway default रूप से thinking दबा देता है; reasoning trace चाहिए तो enable_thinking=true पास करें।
qwen3.7-plus1M$0.256$1.024Alibaba का multimodal agent flagship। Long-running coding/agent loops और 1M context। Reasoning के लिए reasoning.enabled=true पास करें।
qwen3.7-flash1M$0.024$0.104Visual coding, search, tools और high-volume automation के लिए तेज़ multimodal agent model।
qwen3.6-plus1M$0.26$1.561T-parameter MoE flagship। 1M context। OpenRouter पर token volume में top Qwen। Default में सोचता है — gateway default रूप से thinking दबा देता है; reasoning trace चाहिए तो reasoning.enabled=true पास करें।
qwen3.6-35b262K$0.112$0.8035B/3B-active MoE। Long-context RAG और summarization, stronger reasoning के साथ।
qwen3.8-27b1M$0.34$2.04Qwen's fast open 27B reasoning model with a 1M-token context window
qwen3.8-flash-next1M$0.12$0.376Qwen's open Qwen4-architecture preview: 6B-active MoE, fast and cheap, with a 1M-token context window
kimi-k2.6256K$0.5472$2.728Opus-class agentic / planning। जब आपका eval Claude Opus चुनता है, तब best fit।
kimi-k2.7-code256K$0.584$2.80Moonshot का K2, long-horizon agentic coding के लिए tuned। Moonshot के अनुसार हर task पर K2.6 से ~30% कम reasoning tokens।
kimi-k31M$2.55$12.75Moonshot का 2.8T open-weight multimodal reasoning flagship। Complex coding, knowledge work और long-horizon agentic workflows।
muse-spark-1.31M$1.00$3.40Coding & agentic reasoning model, 1M context
muse-spark-1.21M$1.00$3.40Coding-focused reasoning model, 1M context
muse-glimmer-30b131,072$0.28$1.20Compact agentic multimodal model, distilled from Muse Spark
glm-5.31M$1.12$3.52Z.ai latest reasoning flagship, 1M context, complex software engineering & long-horizon agents
glm-5.3-prime1M$2.24$7.04Z.ai GLM 5.3 Prime, top reasoning flagship, 1M context, complex software engineering
glm-5.3-flash1M$0.06$0.20Z.ai fast open-weights reasoning model, 1M context at workhorse pricing
glm-5.21M$1.12$3.52Z.ai का large-scale reasoning flagship। 1M context, long-horizon agent workflows और project-level software engineering के लिए बना। Default में सोचता है — gateway default रूप से thinking दबा देता है; reasoning trace चाहिए तो reasoning.enabled=true पास करें।
nemotron-3.5-lightning262K$0.066$0.17630B-A3B open MoE for fast, tool-heavy agents
gpt-oss-120b131,072$0.041$0.187OpenAI's open-weight 117B MoE, high-volume agents & reasoning, 131K context
gpt-6-astra1M$10.00 ($20.00 >200K)$50.00 ($75.00 >200K)OpenAI का GPT-6 flagship। Long-horizon agentic coding, deep research और document work; image input स्वीकार करता है।
gpt-6-luna1M$0.10 ($0.20 >200K)$0.50 ($0.75 >200K)OpenAI's cost-efficient GPT-6, high-volume chat & lightweight agentic work, 1M context
gpt-6.1-sol1M$2.00 ($4.00 >200K)$10.00 ($15.00 >200K)OpenAI's GPT-6.1 Sol, near-Astra agentic coding at Sol pricing, 1M context
gpt-6-sol1M$2.00 ($4.00 >200K)$10.00 ($15.00 >200K)OpenAI's GPT-6 mid-flagship, strong reasoning & agentic coding, 1M context
gpt-5.6-luna1M$0.16$0.96OpenAI का तेज़, cost-efficient GPT-5.6 tier। High-volume, latency-sensitive chat और classification के लिए। Reasoning model; trace दबाने के लिए reasoning.enabled=false पास करें।
gpt-5.6-terra1M$1.60$9.60OpenAI का balanced GPT-5.6 mid-tier, Luna और Sol के बीच। रोज़मर्रा की coding, reasoning और agentic work के लिए।
gpt-5.6-sol1M$1.60$8.00OpenAI का flagship GPT-5.6। Complex reasoning, coding और agentic workflows, खासकर कठिन multi-step tasks पर।
grok-4.5500K$1.60$4.80xAI का frontier model — coding, knowledge work और STEM के लिए। हमेशा reason करता है — reasoning.enabled=false reject होता है; depth के लिए reasoning_effort इस्तेमाल करें।
grok-4.6500K$2.00 ($4.00 >200K)$6.00 ($12.00 >200K)xAI का frontier model — coding, knowledge work, STEM और image analysis के लिए; image input स्वीकार करता है। हमेशा reason करता है — reasoning.enabled=false reject होता है; depth के लिए reasoning_effort इस्तेमाल करें।
grok-4.7500K$2.00 ($4.00 >200K)$6.00 ($12.00 >200K)xAI का flagship — coding, agentic tasks और knowledge work के लिए, Grok 4.6 का उत्तराधिकारी। लंबे चलने वाले software engineering और self-verification के लिए बना; image input स्वीकार करता है। हमेशा reason करता है — reasoning.enabled=false reject होता है; depth के लिए reasoning_effort इस्तेमाल करें।
minimax-m31M$0.24$0.96MiniMax का open-weight frontier model। Text-only input, long-horizon agentic coding के लिए tuned।
mimo-v2.6-pro1M$0.348$0.696Xiaomi का flagship open-weight MiMo-V2.6-Pro (1.02T MoE, 42B active) — Artificial Analysis Intelligence Index पर नंबर 1 open-weights model। Coding और agentic workflows; image input स्वीकार करता है। Non-thinking mode में चलता है (reasoning बंद)।
mimo-v2.6-flash1M$0.112$0.224Xiaomi का तेज़, कम लागत वाला open-weight MiMo-V2.6-Flash (309B MoE, 15B active)। रोज़मर्रा coding और agents; image input स्वीकार करता है। Non-thinking mode में चलता है (reasoning बंद)।
mimo-v2.51M$0.112$0.224Xiaomi का open-weight MiMo-V2.5। Cost-efficient रोज़मर्रा coding और agentic workflows।
hy4-preview1M$0.6672$2.0008Tencent's Hy4 (preview). Successor to Hy3 for coding and agentic workflows with reasoning.
hy3262K$0.091$0.363Tencent का Hy3। General-purpose coding और agentic workflows, selectable reasoning के साथ।
jev-1.1332K$0.042मुफ़्तTypeSafe का System One model, POST /v1/systemone पर — chat model नहीं। JSON state के बारे में typed choice / score / noul questions का जवाब calibrated probabilities के साथ देता है। केवल input का billing।
claude-opus-5-51M$1.60$8.00Anthropic का सबसे capable Claude और Claude 5.5 family का पहला model। Agentic coding, knowledge work और computer use में Opus 5 और Fable 5.1 से आगे, और per-token price Opus 5 से कम। Image input स्वीकार करता है। इस model पर Anthropic forced tool_choice स्वीकार नहीं करता; QuickSilver Pro इसे auto के साथ उस tool को call करने के निर्देश के रूप में भेजता है — यह steer करता है, call की guarantee नहीं देता।
claude-opus-51M$2.00$10.00Anthropic का Opus 5 — कठिन reasoning, coding और long-horizon agentic work के लिए। Image input स्वीकार करता है।
claude-fable-5-11M$4.00$20.00Anthropic का Mythos-class model, version 5.1 — deep reasoning और long-horizon agentic work के लिए।
claude-fable-51M$4.00$20.00Anthropic का Mythos-class model — deep reasoning और long-horizon agentic work के लिए।
claude-opus-4-81M$2.00$10.00Anthropic flagship। Top-tier reasoning, coding और agentic workloads।
claude-sonnet-5-51M$1.00$5.00Anthropic's newest Sonnet, faster and fewer tokens per task, 1M context
claude-sonnet-51M$1.00$5.00Anthropic का नया Sonnet। बेहतर reasoning और coding, पिछले mid-tier से कम price पर।
claude-haiku-4-5200K$0.40$2.00Anthropic का तेज़, low-cost tier — high-volume work के लिए। Reasoning trace नहीं देता।
gemini-3.7-flash1M$0.6375$3.1875fast multimodal Flash for agents; promotional pricing through 2026-12-31
gemini-3.8-flash1M$0.6375$3.1875most capable Flash for agents; promotional pricing through 2026-12-31
gemini-3.6-flash1M$0.6375$3.1875Current general-purpose Flash GA। नए Flash integrations के लिए recommended।
gemini-3.5-flash-lite1M$0.255$2.125Current low-cost Flash-Lite GA। High-volume workloads के लिए recommended।
gemini-3.5-flash1M$1.275$7.65Google का next-gen Flash GA। 1M context, default में thinks। Capability और price में 3 Flash Preview और 3.1 Pro के बीच।
gemini-3.1-pro-preview1M$1.70$10.20Google का flagship reasoning model। 1M context, deep thinking। Preview API; GA से पहले semantics बदल सकती हैं।
gemini-3-pro-image1M$1.70$10.20GA pro-grade image generation। Retired preview model ID को replace करता है।
gemini-3-flash-preview1M$0.425$2.55केवल legacy compatibility। नए और मौजूदा workloads को gemini-3.6-flash पर migrate करें।
gemini-3.1-flash-lite1M$0.2125$1.275केवल legacy compatibility। नए और मौजूदा workloads को gemini-3.5-flash-lite पर migrate करें।
flux.2-pro-—$0.027 प्रति image/v1/images/generations के ज़रिए flagship image generation। हर image पर billing।
flux.1-schnell-—$0.003 प्रति imageultra-fast open image generation, billed per image
sdxl-turbo-—$0.003 प्रति imagefast open image generation (SDXL Turbo), billed per image
flux.2-klein-—$0.02 प्रति imageopen FLUX.2 image generation, billed per image
qwen-image-max-—$0.1 प्रति imageAlibaba's Qwen-Image flagship: high-fidelity generation with strong typography, billed per image
seedream-5.0-pro-—$0.07 प्रति imageByteDance's Seedream 5.0 Pro: flagship photorealistic image generation, billed per image
seedream-4-—$0.055 प्रति imageByteDance Seedream 4: fast, high-quality image generation, billed per image
bria-fibo-1.5-—$0.055 प्रति imageBria FIBO 1.5: image generation trained on fully licensed data (commercially safe), billed per image

कौन सा मॉडल इस्तेमाल करें?

  • Code-first default - deepseek-v4-flash. Code, chat और production agents के लिए practical starting point। 1M context; input $0.086, output $0.173.
  • Hard reasoning - deepseek-v4-pro. Math, theorem proving और कठिन multi-step problems के लिए। 1M context; input $0.70, output $2.10.
  • Long-document RAG - qwen3.6-35b. Retrieval और long-document summarization के लिए compact MoE। 262K context; input $0.112, output $0.80.
  • Agentic planning - kimi-k2.6. जब evaluations Opus-class planning behavior को पसंद करें। 256K context; input $0.5472, output $2.728.
  • Multimodal workloads - gemini-3.6-flash. जब prompt में text के साथ images भी हों। 1M context; input $0.6375, output $3.1875.
  • Image generation - gemini-3-pro-image. Pro-grade generated image output के लिए। 1M context; input $1.70, output $10.20.
  • High-volume short turns - gemini-3.5-flash-lite. Cost-sensitive routing, classification और extraction के लिए। 1M context; input $0.255, output $2.125.
  • Routing & classification decisions - jev-1.13. System One model, chat नहीं: POST /v1/systemone पर calibrated probabilities के साथ typed answers। 32K context; input $0.042, output मुफ़्त.

Thinking vs non-thinking

V4 wave (V4 Flash, V4 Pro, Kimi K2.6) और Qwen 3.6 final answer से पहले एक chain-of-thought trace emit करते हैं। एक single-token "Hi" भी लगभग 175 reasoning tokens return कर सकता है। इन models पर non-thinking सस्ता chat behavior पाने के लिए pass करें:

पूरा code देखें Quickstart में →

कुछ always-reasoning models (जैसे Grok 4.5) reasoning.enabled=falseको ignore करते हैं — reasoning ही model है। Trace नहीं चाहिए तो V4 Flash use करें।

Claude models: structured output और tool choice

Claude models response_format: {type: "json_schema"} को enforce नहीं करते — request सफल होती है, लेकिन जवाब schema वाले JSON के बजाय सामान्य prose में आता है। Claude model से schema के अनुसार output चाहिए तो schema को strict: true वाले tool की तरह define करें और tool call के arguments पढ़ें; वह रास्ता enforce होता है।

claude-opus-5-5, claude-sonnet-5-5 और claude-fable-5-1 पर Anthropic forced tool_choice स्वीकार नहीं करता। QuickSilver Pro फिर भी इसे स्वीकार करता है और upstream को "auto" के साथ उस tool को call करने का निर्देश भेजता है, इसलिए model को tool की ओर मज़बूती से steer किया जाता है, पर call की guarantee नहीं होती — बिना tool_calls वाले reply को handle करें। बाकी Claude models forced tool choice natively स्वीकार करते हैं।

System One models (typed decisions)

System One models — अभी Jev 1.13 — chat models नहीं हैं। आप एक JSON state और typed questions (choice, score या noul) भेजते हैं POST /v1/systemone पर, और एक ही forward pass में calibrated probabilities वापस मिलती हैं — parse करने के लिए कोई generated text नहीं। Routing, classification, triage, scoring और guardrail decisions के लिए उपयोग करें। केवल input tokens का billing होता है — output मुफ़्त है — और इन पर Chat Completions call reject हो जाती है।

System One docs (English) →

Per-model deep dives