GPT-OSS 120B on QuickSilver Pro
GPT-OSS 120B is OpenAI's open-weight 117B mixture-of-experts model — the most-used open model in production agent stacks. On QuickSilver Pro it's $0.041 input / $0.187 output per million tokens, roughly 70% below the $0.15 / $0.60 that brand-name hosts list, with a 131K-token context window. Reasoning is always on for this endpoint and cannot be disabled — control depth with `reasoning_effort`.
At a glance
High-volume production agents on OpenAI's open-weight 117B MoE — proven at hundreds of billions of tokens.
Pricing comparison ($/1M tokens)
| Provider | Input | Output | vs QSP |
|---|---|---|---|
| QuickSilver Pro | $0.041 | $0.187 | — |
| OpenRouter low rate (openai/gpt-oss-120b) | $0.037 | $0.17 | 10% more expensive |
| OpenAI (GPT-4o mini) | $0.15 | $0.60 | 69% lower |
When to use
Reach for GPT-OSS 120B when you want OpenAI-grade instruction following at open-model prices: high-volume agent loops, tool-heavy automation, background summarization and classification at scale. Its 117B MoE activates only 5.1B parameters per token, so it's fast, and it is battle-tested — it powers some of the largest public agent deployments running today. Reasoning is always on; use `reasoning_effort` to trade depth against latency and cost.
When to use something else
Its 131K-token window is mid-sized — for whole-codebase context, GLM 5.3 Flash and Qwen3.8 27B carry the same price class with far larger windows. For frontier-difficulty reasoning, step up to GPT-5.6 Sol or Claude Opus 5. Reasoning is mandatory here, so trivial high-frequency calls spend thinking tokens — DeepSeek V4 Flash ($0.086/$0.173) replies directly without the overhead.
Quickstart (curl)
curl https://api.quicksilverpro.io/v1/chat/completions \
-H "Authorization: Bearer $QSP_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-oss-120b",
"messages": [{"role": "user", "content": "Hello!"}]
}'OpenAI-compatible. One-line migration via base_url.
FAQ
Reasoning is always on — it is mandatory on this endpoint and cannot be disabled (a request that tries to turn it off is rejected upstream). Use `reasoning_effort` to trade depth against latency and cost. Reasoning tokens bill as output tokens and draw from your `max_tokens` budget before the visible answer, so set `max_tokens` generously or the reply can come back truncated.
Yes — the open-weight 117B mixture-of-experts model OpenAI published, served as a hosted OpenAI-compatible endpoint. You get the model without running 120B-class infrastructure yourself, with streaming, tool calling, and per-token billing on your QuickSilver Pro balance.
Yes — GPT-OSS 120B is an OpenAI-compatible chat completions endpoint on QuickSilver Pro (the Responses API works too). Set base_url=https://api.quicksilverpro.io/v1, paste your QSP key, and use model="gpt-oss-120b". Streaming, tool calling, and usage.cost accounting all work.
Brand-name hosts list GPT-OSS 120B at $0.15 input / $0.60 output per million tokens; QuickSilver Pro is $0.041 / $0.187 — roughly 70% below that tier, and about 10% above the cheapest OpenRouter listings. Same OpenAI-compatible surface; migration is a base_url + key swap, dropping the `openai/` provider prefix from the model ID.