मॉडल catalog
एक OpenAI-compatible API: open-source LLMs (DeepSeek V4, Qwen — 3.7 Max और 3.6 Plus सहित, Kimi, GLM) और Google का Gemini family — multimodal chat, reasoning तथा image generation के लिए। सभी prices per 1M tokens, USD में, July 2026 के अनुसार। Request body में model field में जो pass करते हैं वो है Model ID।
एक नज़र में
| Model ID | Context | Input | Output | Notes |
|---|---|---|---|---|
| deepseek-v4-flash | 1M | $0.086 | $0.173 | सस्ता chat और coding workhorse। Default में सोचता है — V3-style replies के लिए reasoning.enabled=false पास करें। |
| deepseek-v4.1-flash | 1M | $0.125 | $0.55 | official DeepSeek V4.1 Flash — new architecture, faster, 1M context, Responses API |
| deepseek-v4-pro | 1M | $0.70 | $2.10 | Premium reasoning + 1M token context। o3-mini के बराबर quality, output cost लगभग 6× कम। |
| qwen3.8-max | 1M | $2.00 | $6.00 | Alibaba का Qwen 3.8 flagship, 2026-08-03 को release हुआ: 2.4T-parameter MoE जिसे Alibaba autonomous coding और long-horizon agents के लिए position करता है। Default में सोचता है — gateway default रूप से thinking दबा देता है; reasoning trace चाहिए तो enable_thinking=true पास करें। |
| qwen3.8-max-prime | 1M | $4.00 | $12.00 | Qwen 3.8 top-tier flagship, deepest reasoning & autonomous coding, 1M context |
| qwen3.8-omni-flash | 1M | $0.15 | $0.47 | Qwen's first agentic omni model: native image, audio and video understanding with tool use and reasoning, 1M context, at Alibaba's first-party price |
| qwen3.7-max | 1M | $1.25 | $3.75 | Qwen 3.7 flagship। Agent-centric workloads, coding और productivity tasks में strong। Default में सोचता है — gateway default रूप से thinking दबा देता है; reasoning trace चाहिए तो enable_thinking=true पास करें। |
| qwen3.7-plus | 1M | $0.256 | $1.024 | Alibaba का multimodal agent flagship। Long-running coding/agent loops और 1M context। Reasoning के लिए reasoning.enabled=true पास करें। |
| qwen3.7-flash | 1M | $0.024 | $0.104 | Visual coding, search, tools और high-volume automation के लिए तेज़ multimodal agent model। |
| qwen3.6-plus | 1M | $0.26 | $1.56 | 1T-parameter MoE flagship। 1M context। OpenRouter पर token volume में top Qwen। Default में सोचता है — gateway default रूप से thinking दबा देता है; reasoning trace चाहिए तो reasoning.enabled=true पास करें। |
| qwen3.6-35b | 262K | $0.112 | $0.80 | 35B/3B-active MoE। Long-context RAG और summarization, stronger reasoning के साथ। |
| qwen3.8-27b | 1M | $0.34 | $2.04 | Qwen's fast open 27B reasoning model with a 1M-token context window |
| qwen3.8-flash-next | 1M | $0.12 | $0.376 | Qwen's open Qwen4-architecture preview: 6B-active MoE, fast and cheap, with a 1M-token context window |
| kimi-k2.6 | 256K | $0.5472 | $2.728 | Opus-class agentic / planning। जब आपका eval Claude Opus चुनता है, तब best fit। |
| kimi-k2.7-code | 256K | $0.584 | $2.80 | Moonshot का K2, long-horizon agentic coding के लिए tuned। Moonshot के अनुसार हर task पर K2.6 से ~30% कम reasoning tokens। |
| kimi-k3 | 1M | $2.55 | $12.75 | Moonshot का 2.8T open-weight multimodal reasoning flagship। Complex coding, knowledge work और long-horizon agentic workflows। |
| muse-spark-1.3 | 1M | $1.00 | $3.40 | Coding & agentic reasoning model, 1M context |
| muse-spark-1.2 | 1M | $1.00 | $3.40 | Coding-focused reasoning model, 1M context |
| muse-glimmer-30b | 131,072 | $0.28 | $1.20 | Compact agentic multimodal model, distilled from Muse Spark |
| glm-5.3 | 1M | $1.12 | $3.52 | Z.ai latest reasoning flagship, 1M context, complex software engineering & long-horizon agents |
| glm-5.3-prime | 1M | $2.24 | $7.04 | Z.ai GLM 5.3 Prime, top reasoning flagship, 1M context, complex software engineering |
| glm-5.3-flash | 1M | $0.06 | $0.20 | Z.ai fast open-weights reasoning model, 1M context at workhorse pricing |
| glm-5.2 | 1M | $1.12 | $3.52 | Z.ai का large-scale reasoning flagship। 1M context, long-horizon agent workflows और project-level software engineering के लिए बना। Default में सोचता है — gateway default रूप से thinking दबा देता है; reasoning trace चाहिए तो reasoning.enabled=true पास करें। |
| nemotron-3.5-lightning | 262K | $0.066 | $0.176 | 30B-A3B open MoE for fast, tool-heavy agents |
| gpt-oss-120b | 131,072 | $0.041 | $0.187 | OpenAI's open-weight 117B MoE, high-volume agents & reasoning, 131K context |
| gpt-6-astra | 1M | $10.00 ($20.00 >200K) | $50.00 ($75.00 >200K) | OpenAI का GPT-6 flagship। Long-horizon agentic coding, deep research और document work; image input स्वीकार करता है। |
| gpt-6-luna | 1M | $0.10 ($0.20 >200K) | $0.50 ($0.75 >200K) | OpenAI's cost-efficient GPT-6, high-volume chat & lightweight agentic work, 1M context |
| gpt-6.1-sol | 1M | $2.00 ($4.00 >200K) | $10.00 ($15.00 >200K) | OpenAI's GPT-6.1 Sol, near-Astra agentic coding at Sol pricing, 1M context |
| gpt-6-sol | 1M | $2.00 ($4.00 >200K) | $10.00 ($15.00 >200K) | OpenAI's GPT-6 mid-flagship, strong reasoning & agentic coding, 1M context |
| gpt-5.6-luna | 1M | $0.16 | $0.96 | OpenAI का तेज़, cost-efficient GPT-5.6 tier। High-volume, latency-sensitive chat और classification के लिए। Reasoning model; trace दबाने के लिए reasoning.enabled=false पास करें। |
| gpt-5.6-terra | 1M | $1.60 | $9.60 | OpenAI का balanced GPT-5.6 mid-tier, Luna और Sol के बीच। रोज़मर्रा की coding, reasoning और agentic work के लिए। |
| gpt-5.6-sol | 1M | $1.60 | $8.00 | OpenAI का flagship GPT-5.6। Complex reasoning, coding और agentic workflows, खासकर कठिन multi-step tasks पर। |
| grok-4.5 | 500K | $1.60 | $4.80 | xAI का frontier model — coding, knowledge work और STEM के लिए। हमेशा reason करता है — reasoning.enabled=false reject होता है; depth के लिए reasoning_effort इस्तेमाल करें। |
| grok-4.6 | 500K | $2.00 ($4.00 >200K) | $6.00 ($12.00 >200K) | xAI का frontier model — coding, knowledge work, STEM और image analysis के लिए; image input स्वीकार करता है। हमेशा reason करता है — reasoning.enabled=false reject होता है; depth के लिए reasoning_effort इस्तेमाल करें। |
| grok-4.7 | 500K | $2.00 ($4.00 >200K) | $6.00 ($12.00 >200K) | xAI का flagship — coding, agentic tasks और knowledge work के लिए, Grok 4.6 का उत्तराधिकारी। लंबे चलने वाले software engineering और self-verification के लिए बना; image input स्वीकार करता है। हमेशा reason करता है — reasoning.enabled=false reject होता है; depth के लिए reasoning_effort इस्तेमाल करें। |
| minimax-m3 | 1M | $0.24 | $0.96 | MiniMax का open-weight frontier model। Text-only input, long-horizon agentic coding के लिए tuned। |
| mimo-v2.6-pro | 1M | $0.348 | $0.696 | Xiaomi का flagship open-weight MiMo-V2.6-Pro (1.02T MoE, 42B active) — Artificial Analysis Intelligence Index पर नंबर 1 open-weights model। Coding और agentic workflows; image input स्वीकार करता है। Non-thinking mode में चलता है (reasoning बंद)। |
| mimo-v2.6-flash | 1M | $0.112 | $0.224 | Xiaomi का तेज़, कम लागत वाला open-weight MiMo-V2.6-Flash (309B MoE, 15B active)। रोज़मर्रा coding और agents; image input स्वीकार करता है। Non-thinking mode में चलता है (reasoning बंद)। |
| mimo-v2.5 | 1M | $0.112 | $0.224 | Xiaomi का open-weight MiMo-V2.5। Cost-efficient रोज़मर्रा coding और agentic workflows। |
| hy4-preview | 1M | $0.6672 | $2.0008 | Tencent's Hy4 (preview). Successor to Hy3 for coding and agentic workflows with reasoning. |
| hy3 | 262K | $0.091 | $0.363 | Tencent का Hy3। General-purpose coding और agentic workflows, selectable reasoning के साथ। |
| jev-1.13 | 32K | $0.042 | मुफ़्त | TypeSafe का System One model, POST /v1/systemone पर — chat model नहीं। JSON state के बारे में typed choice / score / noul questions का जवाब calibrated probabilities के साथ देता है। केवल input का billing। |
| claude-opus-5-5 | 1M | $1.60 | $8.00 | Anthropic का सबसे capable Claude और Claude 5.5 family का पहला model। Agentic coding, knowledge work और computer use में Opus 5 और Fable 5.1 से आगे, और per-token price Opus 5 से कम। Image input स्वीकार करता है। इस model पर Anthropic forced tool_choice स्वीकार नहीं करता; QuickSilver Pro इसे auto के साथ उस tool को call करने के निर्देश के रूप में भेजता है — यह steer करता है, call की guarantee नहीं देता। |
| claude-opus-5 | 1M | $2.00 | $10.00 | Anthropic का Opus 5 — कठिन reasoning, coding और long-horizon agentic work के लिए। Image input स्वीकार करता है। |
| claude-fable-5-1 | 1M | $4.00 | $20.00 | Anthropic का Mythos-class model, version 5.1 — deep reasoning और long-horizon agentic work के लिए। |
| claude-fable-5 | 1M | $4.00 | $20.00 | Anthropic का Mythos-class model — deep reasoning और long-horizon agentic work के लिए। |
| claude-opus-4-8 | 1M | $2.00 | $10.00 | Anthropic flagship। Top-tier reasoning, coding और agentic workloads। |
| claude-sonnet-5-5 | 1M | $1.00 | $5.00 | Anthropic's newest Sonnet, faster and fewer tokens per task, 1M context |
| claude-sonnet-5 | 1M | $1.00 | $5.00 | Anthropic का नया Sonnet। बेहतर reasoning और coding, पिछले mid-tier से कम price पर। |
| claude-haiku-4-5 | 200K | $0.40 | $2.00 | Anthropic का तेज़, low-cost tier — high-volume work के लिए। Reasoning trace नहीं देता। |
| gemini-3.7-flash | 1M | $0.6375 | $3.1875 | fast multimodal Flash for agents; promotional pricing through 2026-12-31 |
| gemini-3.8-flash | 1M | $0.6375 | $3.1875 | most capable Flash for agents; promotional pricing through 2026-12-31 |
| gemini-3.6-flash | 1M | $0.6375 | $3.1875 | Current general-purpose Flash GA। नए Flash integrations के लिए recommended। |
| gemini-3.5-flash-lite | 1M | $0.255 | $2.125 | Current low-cost Flash-Lite GA। High-volume workloads के लिए recommended। |
| gemini-3.5-flash | 1M | $1.275 | $7.65 | Google का next-gen Flash GA। 1M context, default में thinks। Capability और price में 3 Flash Preview और 3.1 Pro के बीच। |
| gemini-3.1-pro-preview | 1M | $1.70 | $10.20 | Google का flagship reasoning model। 1M context, deep thinking। Preview API; GA से पहले semantics बदल सकती हैं। |
| gemini-3-pro-image | 1M | $1.70 | $10.20 | GA pro-grade image generation। Retired preview model ID को replace करता है। |
| gemini-3-flash-preview | 1M | $0.425 | $2.55 | केवल legacy compatibility। नए और मौजूदा workloads को gemini-3.6-flash पर migrate करें। |
| gemini-3.1-flash-lite | 1M | $0.2125 | $1.275 | केवल legacy compatibility। नए और मौजूदा workloads को gemini-3.5-flash-lite पर migrate करें। |
| flux.2-pro | - | — | $0.027 प्रति image | /v1/images/generations के ज़रिए flagship image generation। हर image पर billing। |
| flux.1-schnell | - | — | $0.003 प्रति image | ultra-fast open image generation, billed per image |
| sdxl-turbo | - | — | $0.003 प्रति image | fast open image generation (SDXL Turbo), billed per image |
| flux.2-klein | - | — | $0.02 प्रति image | open FLUX.2 image generation, billed per image |
| qwen-image-max | - | — | $0.1 प्रति image | Alibaba's Qwen-Image flagship: high-fidelity generation with strong typography, billed per image |
| seedream-5.0-pro | - | — | $0.07 प्रति image | ByteDance's Seedream 5.0 Pro: flagship photorealistic image generation, billed per image |
| seedream-4 | - | — | $0.055 प्रति image | ByteDance Seedream 4: fast, high-quality image generation, billed per image |
| bria-fibo-1.5 | - | — | $0.055 प्रति image | Bria FIBO 1.5: image generation trained on fully licensed data (commercially safe), billed per image |
कौन सा मॉडल इस्तेमाल करें?
- Code-first default -
deepseek-v4-flash. Code, chat और production agents के लिए practical starting point। 1M context; input $0.086, output $0.173. - Hard reasoning -
deepseek-v4-pro. Math, theorem proving और कठिन multi-step problems के लिए। 1M context; input $0.70, output $2.10. - Long-document RAG -
qwen3.6-35b. Retrieval और long-document summarization के लिए compact MoE। 262K context; input $0.112, output $0.80. - Agentic planning -
kimi-k2.6. जब evaluations Opus-class planning behavior को पसंद करें। 256K context; input $0.5472, output $2.728. - Multimodal workloads -
gemini-3.6-flash. जब prompt में text के साथ images भी हों। 1M context; input $0.6375, output $3.1875. - Image generation -
gemini-3-pro-image. Pro-grade generated image output के लिए। 1M context; input $1.70, output $10.20. - High-volume short turns -
gemini-3.5-flash-lite. Cost-sensitive routing, classification और extraction के लिए। 1M context; input $0.255, output $2.125. - Routing & classification decisions -
jev-1.13. System One model, chat नहीं: POST /v1/systemone पर calibrated probabilities के साथ typed answers। 32K context; input $0.042, output मुफ़्त.
Thinking vs non-thinking
V4 wave (V4 Flash, V4 Pro, Kimi K2.6) और Qwen 3.6 final answer से पहले एक chain-of-thought trace emit करते हैं। एक single-token "Hi" भी लगभग 175 reasoning tokens return कर सकता है। इन models पर non-thinking सस्ता chat behavior पाने के लिए pass करें:
कुछ always-reasoning models (जैसे Grok 4.5) reasoning.enabled=falseको ignore करते हैं — reasoning ही model है। Trace नहीं चाहिए तो V4 Flash use करें।
Claude models: structured output और tool choice
Claude models response_format: {type: "json_schema"} को enforce नहीं करते — request सफल होती है, लेकिन जवाब schema वाले JSON के बजाय सामान्य prose में आता है। Claude model से schema के अनुसार output चाहिए तो schema को strict: true वाले tool की तरह define करें और tool call के arguments पढ़ें; वह रास्ता enforce होता है।
claude-opus-5-5, claude-sonnet-5-5 और claude-fable-5-1 पर Anthropic forced tool_choice स्वीकार नहीं करता। QuickSilver Pro फिर भी इसे स्वीकार करता है और upstream को "auto" के साथ उस tool को call करने का निर्देश भेजता है, इसलिए model को tool की ओर मज़बूती से steer किया जाता है, पर call की guarantee नहीं होती — बिना tool_calls वाले reply को handle करें। बाकी Claude models forced tool choice natively स्वीकार करते हैं।
System One models (typed decisions)
System One models — अभी Jev 1.13 — chat models नहीं हैं। आप एक JSON state और typed questions (choice, score या noul) भेजते हैं POST /v1/systemone पर, और एक ही forward pass में calibrated probabilities वापस मिलती हैं — parse करने के लिए कोई generated text नहीं। Routing, classification, triage, scoring और guardrail decisions के लिए उपयोग करें। केवल input tokens का billing होता है — output मुफ़्त है — और इन पर Chat Completions call reject हो जाती है।
Per-model deep dives
हर model का अपना page है — pricing comparisons, FAQs, और quickstart code के साथ। Links यहाँ:
- DeepSeek V4 Flash
- DeepSeek V4.1 Flash
- DeepSeek V4 Pro
- Qwen3.8 Max
- Qwen3.8 Omni Flash
- Qwen3.7 Max
- Qwen3.7 Plus
- Qwen3.7 Flash
- Qwen3.6 Plus
- Qwen3.6-35B-A3B
- Qwen3.8 27B
- Qwen3.8 Flash Next
- Kimi K2.6
- Kimi K2.7 Code
- Kimi K3
- Muse Spark 1.3
- Muse Spark 1.2
- Muse Glimmer 30B
- GLM 5.3
- GLM 5.3 Flash
- GLM 5.2
- Nemotron 3.5 Lightning
- GPT-OSS 120B
- GPT-6 Astra
- Grok 4.6
- Grok 4.7
- MiniMax M3
- MiMo-V2.6-Pro
- MiMo-V2.6-Flash
- MiMo-V2.5
- Hy4 Preview
- Hy3
- Jev 1.13
- Claude Opus 5.5
- Gemini 3.7 Flash
- Gemini 3.8 Flash
- Gemini 3.6 Flash
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash
- Gemini 3.1 Pro Preview
- Gemini 3 Flash Preview
- Gemini 3.1 Flash Lite
- FLUX.2 Pro
- FLUX.1 Schnell
- SDXL Turbo
- FLUX.2 Klein
- Qwen-Image Max
- Seedream 5.0 Pro
- Seedream 4
- Bria FIBO 1.5
Gemini family के per-model deep-dive pages अभी work in progress हैं — तब तक ऊपर का catalog table canonical pricing और capability summary है।