模型目录
一个 OpenAI 兼容 API:开源 LLM(DeepSeek V4、Qwen 含 3.7 Max 与 3.6 Plus、Kimi、GLM)外加 Google Gemini 全家桶,支持多模态聊天、推理与图像生成。所有价格均为每百万 token 美元报价,截至 2026 年 7 月。请求 body 中的 model 字段就是下表中的 Model ID。
一览
| Model ID | 上下文 | 输入 | 输出 | 说明 |
|---|---|---|---|---|
| deepseek-v4-flash | 1M | $0.086 | $0.173 | 便宜的聊天 & 代码主力。默认思考——若想要 V3 风格的快速回复,传 reasoning.enabled=false。 |
| deepseek-v4.1-flash | 1M | $0.125 | $0.55 | official DeepSeek V4.1 Flash — new architecture, faster, 1M context, Responses API |
| deepseek-v4-pro | 1M | $0.70 | $2.10 | 高端推理 + 1M token 上下文。对标 o3-mini,输出价格约低 6 倍。 |
| qwen3.8-max | 1M | $2.00 | $6.00 | 阿里 Qwen 3.8 旗舰,2026-08-03 发布:2.4T 参数 MoE,阿里官方定位为自主编程与长程 agent。默认思考——网关默认已抑制思考;如需思考链,传 enable_thinking=true 开启。 |
| qwen3.8-max-prime | 1M | $4.00 | $12.00 | Qwen 3.8 top-tier flagship, deepest reasoning & autonomous coding, 1M context |
| qwen3.8-omni-flash | 1M | $0.15 | $0.47 | Qwen's first agentic omni model: native image, audio and video understanding with tool use and reasoning, 1M context, at Alibaba's first-party price |
| qwen3.7-max | 1M | $1.25 | $3.75 | Qwen 3.7 旗舰。面向 agent 工作负载,编程与生产力任务表现突出。默认思考——网关默认已抑制思考;如需思考链,传 enable_thinking=true 开启。 |
| qwen3.7-plus | 1M | $0.256 | $1.024 | 阿里托管的多模态 agent 旗舰。长时运行的编程/agent 循环,1M 上下文。默认已抑制思考;传 reasoning.enabled=true 开启。 |
| qwen3.7-flash | 1M | $0.024 | $0.104 | 快速多模态 agent 模型,适合视觉编程、搜索、工具调用和高吞吐自动化;支持切换思考。 |
| qwen3.6-plus | 1M | $0.26 | $1.56 | 1T 参数 MoE 旗舰。1M 上下文。OpenRouter 上 token 量第一的 Qwen。默认思考——网关默认已抑制思考;如需思考链,传 reasoning.enabled=true 开启。 |
| qwen3.6-35b | 262K | $0.112 | $0.80 | 35B/3B-active MoE。长上下文 RAG 与摘要,推理能力强。 |
| qwen3.8-27b | 1M | $0.34 | $2.04 | Qwen's fast open 27B reasoning model with a 1M-token context window |
| qwen3.8-flash-next | 1M | $0.12 | $0.376 | Qwen's open Qwen4-architecture preview: 6B-active MoE, fast and cheap, with a 1M-token context window |
| kimi-k2.6 | 256K | $0.5472 | $2.728 | Opus 级 agentic / 规划场景。当你的评测选出 Claude Opus 时,这就是最合适的开源对标。 |
| kimi-k2.7-code | 256K | $0.584 | $2.80 | 为长时 agentic 编程调优的始终思考型 K2;采样参数固定。Moonshot 称过度思考比 K2.6 少约 30%。 |
| kimi-k3 | 1M | $2.55 | $12.75 | Moonshot 2.8T 开源权重多模态推理旗舰。复杂编程、知识工作与长时 agentic 工作流。 |
| muse-spark-1.3 | 1M | $1.00 | $3.40 | Coding & agentic reasoning model, 1M context |
| muse-spark-1.2 | 1M | $1.00 | $3.40 | Coding-focused reasoning model, 1M context |
| muse-glimmer-30b | 131,072 | $0.28 | $1.20 | Compact agentic multimodal model, distilled from Muse Spark |
| glm-5.3 | 1M | $1.12 | $3.52 | Z.ai latest reasoning flagship, 1M context, complex software engineering & long-horizon agents |
| glm-5.3-prime | 1M | $2.24 | $7.04 | Z.ai GLM 5.3 Prime, top reasoning flagship, 1M context, complex software engineering |
| glm-5.3-flash | 1M | $0.06 | $0.20 | Z.ai fast open-weights reasoning model, 1M context at workhorse pricing |
| glm-5.2 | 1M | $1.12 | $3.52 | Z.ai 的大规模推理旗舰。1M 上下文,专为长时 agent 工作流和项目级软件工程打造。默认思考——网关默认已抑制思考;如需思考链,传 reasoning.enabled=true 开启。 |
| nemotron-3.5-lightning | 262K | $0.066 | $0.176 | 30B-A3B open MoE for fast, tool-heavy agents |
| gpt-oss-120b | 131,072 | $0.041 | $0.187 | OpenAI's open-weight 117B MoE, high-volume agents & reasoning, 131K context |
| gpt-6-astra | 1M | $10.00 ($20.00 >200K) | $50.00 ($75.00 >200K) | OpenAI GPT-6 旗舰。长程 agentic 编程、深度研究与文档工作,支持图像输入。 |
| gpt-6-luna | 1M | $0.10 ($0.20 >200K) | $0.50 ($0.75 >200K) | OpenAI's cost-efficient GPT-6, high-volume chat & lightweight agentic work, 1M context |
| gpt-6.1-sol | 1M | $2.00 ($4.00 >200K) | $10.00 ($15.00 >200K) | OpenAI's GPT-6.1 Sol, near-Astra agentic coding at Sol pricing, 1M context |
| gpt-6-sol | 1M | $2.00 ($4.00 >200K) | $10.00 ($15.00 >200K) | OpenAI's GPT-6 mid-flagship, strong reasoning & agentic coding, 1M context |
| gpt-5.6-luna | 1M | $0.16 | $0.96 | OpenAI 快速、高性价比的 GPT-5.6 档位。适合高并发、低延迟的对话与分类。推理模型;传 reasoning.enabled=false 可抑制思考链。 |
| gpt-5.6-terra | 1M | $1.60 | $9.60 | OpenAI 均衡的 GPT-5.6 中档,介于 Luna 与 Sol 之间。适合日常编程、推理与 agent 工作。 |
| gpt-5.6-sol | 1M | $1.60 | $8.00 | OpenAI GPT-5.6 旗舰。复杂推理、编程与 agent 工作流,尤其擅长多步骤难题。 |
| grok-4.5 | 500K | $1.60 | $4.80 | xAI 前沿模型,擅长编程、知识工作与 STEM。始终推理——不接受 reasoning.enabled=false;请用 reasoning_effort 控制思考深度。 |
| grok-4.6 | 500K | $2.00 ($4.00 >200K) | $6.00 ($12.00 >200K) | xAI 前沿模型,擅长编程、知识工作、STEM 与图像分析;支持图像输入。始终推理——不接受 reasoning.enabled=false;请用 reasoning_effort 控制思考深度。 |
| grok-4.7 | 500K | $2.00 ($4.00 >200K) | $6.00 ($12.00 >200K) | xAI 旗舰模型,接替 Grok 4.6,擅长编程、agentic 任务与知识工作。专为长时软件工程与自我校验打造,支持图像输入。始终推理——不接受 reasoning.enabled=false;请用 reasoning_effort 控制思考深度。 |
| minimax-m3 | 1M | $0.24 | $0.96 | MiniMax 开源权重前沿模型。仅支持文本输入,面向长时 agentic 编程调优。 |
| mimo-v2.6-pro | 1M | $0.348 | $0.696 | 小米旗舰开源权重 MiMo-V2.6-Pro(1.02T MoE,激活 42B),Artificial Analysis 智能指数排名第一的开源权重模型。擅长编程与 agent 工作流,支持图像输入。以非推理模式提供(推理关闭)。 |
| mimo-v2.6-flash | 1M | $0.112 | $0.224 | 小米开源权重 MiMo-V2.6-Flash(309B MoE,激活 15B),快速低价。适合日常编程与 agent,支持图像输入。以非推理模式提供(推理关闭)。 |
| mimo-v2.5 | 1M | $0.112 | $0.224 | 小米开源权重 MiMo-V2.5。高性价比的日常编程与 agent 工作流。 |
| hy4-preview | 1M | $0.6672 | $2.0008 | Tencent's Hy4 (preview). Successor to Hy3 for coding and agentic workflows with reasoning. |
| hy3 | 262K | $0.091 | $0.363 | 腾讯 Hy3。通用编程与 agent 工作流,支持切换思考。 |
| jev-1.13 | 32K | $0.042 | 免费 | TypeSafe 的 System One 模型,走 POST /v1/systemone —— 不是聊天模型。针对 JSON 状态回答 choice / score / noul 类型化问题,返回校准概率。只按输入计费。 |
| claude-opus-5-5 | 1M | $1.60 | $8.00 | Anthropic 目前能力最强的 Claude,也是 Claude 5.5 系列的首款模型。在 agentic 编程、知识工作与电脑操作上领先 Opus 5 和 Fable 5.1,每 token 价格低于 Opus 5。支持图像输入。Anthropic 官方不支持在该模型上强制指定 tool_choice;QuickSilver Pro 会改成 auto 并附上「调用该工具」的指令,能强烈引导但不保证一定调用。 |
| claude-opus-5 | 1M | $2.00 | $10.00 | Anthropic Opus 5,面向高难推理、编程与长时 agentic 工作。支持图像输入。 |
| claude-fable-5-1 | 1M | $4.00 | $20.00 | Anthropic Mythos 级模型 5.1 版,面向深度推理与长时 agentic 工作。 |
| claude-fable-5 | 1M | $4.00 | $20.00 | Anthropic Mythos 级模型,面向深度推理与长时 agentic 工作。 |
| claude-opus-4-8 | 1M | $2.00 | $10.00 | Anthropic 旗舰。顶级推理、编程与 agent 工作负载。 |
| claude-sonnet-5-5 | 1M | $1.00 | $5.00 | Anthropic's newest Sonnet, faster and fewer tokens per task, 1M context |
| claude-sonnet-5 | 1M | $1.00 | $5.00 | Anthropic 最新 Sonnet。推理与编码更强,价格低于上一代中档。 |
| claude-haiku-4-5 | 200K | $0.40 | $2.00 | Anthropic 快速低成本档位,适合高并发场景。不输出思考链。 |
| gemini-3.7-flash | 1M | $0.6375 | $3.1875 | fast multimodal Flash for agents; promotional pricing through 2026-12-31 |
| gemini-3.8-flash | 1M | $0.6375 | $3.1875 | most capable Flash for agents; promotional pricing through 2026-12-31 |
| gemini-3.6-flash | 1M | $0.6375 | $3.1875 | 当前推荐的通用 Flash 正式版(GA),新 Flash 集成应优先使用。 |
| gemini-3.5-flash-lite | 1M | $0.255 | $2.125 | 当前推荐的低成本 Flash-Lite 正式版(GA),适合高并发任务。 |
| gemini-3.5-flash | 1M | $1.275 | $7.65 | Google 下一代 Flash 正式版(GA)。1M 上下文,默认思考。能力与价格介于 3 Flash Preview 与 3.1 Pro 之间。 |
| gemini-3.1-pro-preview | 1M | $1.70 | $10.20 | Google 旗舰推理模型。1M 上下文,深度思考。Preview API;正式版前语义可能仍有调整。 |
| gemini-3-pro-image | 1M | $1.70 | $10.20 | 正式版专业级图像生成,替代已停用的 preview 模型 ID。 |
| gemini-3-flash-preview | 1M | $0.425 | $2.55 | 仅保留旧集成兼容。新旧工作负载请迁移到 gemini-3.6-flash。 |
| gemini-3.1-flash-lite | 1M | $0.2125 | $1.275 | 仅保留旧集成兼容。新旧工作负载请迁移到 gemini-3.5-flash-lite。 |
| flux.2-pro | - | — | $0.027 / 张 | 旗舰图像生成,走 /v1/images/generations,按张计费。 |
| flux.1-schnell | - | — | $0.003 / 张 | ultra-fast open image generation, billed per image |
| sdxl-turbo | - | — | $0.003 / 张 | fast open image generation (SDXL Turbo), billed per image |
| flux.2-klein | - | — | $0.02 / 张 | open FLUX.2 image generation, billed per image |
| qwen-image-max | - | — | $0.1 / 张 | Alibaba's Qwen-Image flagship: high-fidelity generation with strong typography, billed per image |
| seedream-5.0-pro | - | — | $0.07 / 张 | ByteDance's Seedream 5.0 Pro: flagship photorealistic image generation, billed per image |
| seedream-4 | - | — | $0.055 / 张 | ByteDance Seedream 4: fast, high-quality image generation, billed per image |
| bria-fibo-1.5 | - | — | $0.055 / 张 | Bria FIBO 1.5: image generation trained on fully licensed data (commercially safe), billed per image |
该用哪个模型?
- 代码优先的默认选择 -
deepseek-v4-flash. 适合作为代码、聊天和生产 agent 的起点。 1M 上下文; 输入 $0.086, 输出 $0.173. - 高难推理 -
deepseek-v4-pro. 适合数学、定理证明和复杂多步骤问题。 1M 上下文; 输入 $0.70, 输出 $2.10. - 长文档 RAG -
qwen3.6-35b. 适合检索和长文档摘要的紧凑 MoE 选择。 262K 上下文; 输入 $0.112, 输出 $0.80. - Agentic 规划 -
kimi-k2.6. 当评测偏好 Opus 级规划行为时使用。 256K 上下文; 输入 $0.5472, 输出 $2.728. - 多模态任务 -
gemini-3.6-flash. 当 prompt 同时包含图像和文本时使用。 1M 上下文; 输入 $0.6375, 输出 $3.1875. - 图像生成 -
gemini-3-pro-image. 用于专业级图像输出。 1M 上下文; 输入 $1.70, 输出 $10.20. - 高频短对话 -
gemini-3.5-flash-lite. 适合成本敏感的路由、分类和抽取。 1M 上下文; 输入 $0.255, 输出 $2.125. - 路由与分类决策 -
jev-1.13. System One 模型,不是聊天:在 POST /v1/systemone 上返回带校准概率的类型化答案。 32K 上下文; 输入 $0.042, 输出 免费.
思考模式 vs 非思考模式
V4 系列(V4 Flash、V4 Pro、Kimi K2.6)以及 Qwen 3.6 会在最终答案前输出一段 chain-of-thought 轨迹。一句简单的 "Hi" 也可能产生约 175 个推理 token。如果想在这些模型上拿到 非思考的廉价聊天行为,传:
部分始终推理的模型(如 Grok 4.5)会忽略 reasoning.enabled=false —— 推理本身就是模型。如果不想要 thinking trace,请使用 V4 Flash。
Claude 模型:结构化输出与工具选择
Claude 系列模型不会强制执行 response_format: {type: "json_schema"}——请求会成功,但返回的是普通文本,而不是符合 schema 的 JSON。如果需要结构化输出,请把 schema 定义成带 strict: true 的 tool,然后读取 tool call 的参数;这条路径是受约束的。
claude-opus-5-5、claude-sonnet-5-5 和 claude-fable-5-1 在 Anthropic 官方不支持强制指定 tool_choice。QuickSilver Pro 仍然接受这类请求,会改成 "auto"并附上「调用该工具」的指令发给上游,因此模型会被强烈引导去调用该工具,但不能保证一定调用——请处理回复中没有 tool_calls 的情况。其他 Claude 模型原生支持强制指定。
System One 模型(类型化决策)
System One 模型(目前为 Jev 1.13)不是聊天模型。把一个 JSON 状态和若干类型化问题(choice、score 或 noul)发到 POST /v1/systemone,一次前向计算即返回校准概率,没有需要解析的生成文本。适合路由、分类、分诊、打分和护栏类决策。只按输入 token 计费 —— 输出免费;对它发起 Chat Completions 调用会被拒绝。
单模型详细页
每个模型都有自己的页面,包含定价对比、FAQ 与 quickstart 代码。链接如下:
- DeepSeek V4 Flash
- DeepSeek V4.1 Flash
- DeepSeek V4 Pro
- Qwen3.8 Max
- Qwen3.8 Omni Flash
- Qwen3.7 Max
- Qwen3.7 Plus
- Qwen3.7 Flash
- Qwen3.6 Plus
- Qwen3.6-35B-A3B
- Qwen3.8 27B
- Qwen3.8 Flash Next
- Kimi K2.6
- Kimi K2.7 Code
- Kimi K3
- Muse Spark 1.3
- Muse Spark 1.2
- Muse Glimmer 30B
- GLM 5.3
- GLM 5.3 Flash
- GLM 5.2
- Nemotron 3.5 Lightning
- GPT-OSS 120B
- GPT-6 Astra
- Grok 4.6
- Grok 4.7
- MiniMax M3
- MiMo-V2.6-Pro
- MiMo-V2.6-Flash
- MiMo-V2.5
- Hy4 Preview
- Hy3
- Jev 1.13
- Claude Opus 5.5
- Gemini 3.7 Flash
- Gemini 3.8 Flash
- Gemini 3.6 Flash
- Gemini 3.5 Flash-Lite
- Gemini 3.5 Flash
- Gemini 3.1 Pro Preview
- Gemini 3 Flash Preview
- Gemini 3.1 Flash Lite
- FLUX.2 Pro
- FLUX.1 Schnell
- SDXL Turbo
- FLUX.2 Klein
- Qwen-Image Max
- Seedream 5.0 Pro
- Seedream 4
- Bria FIBO 1.5
Gemini 全家桶的单模型详细页仍在制作中 —— 在此期间,上表是定价与能力的权威摘要。