Home/Models/GPT-OSS 120B
131K contextReasoningResponses API

GPT-OSS 120B on QuickSilver Pro

GPT-OSS 120B is OpenAI's open-weight 117B mixture-of-experts model — the most-used open model in production agent stacks. On QuickSilver Pro it's $0.041 input / $0.187 output per million tokens, roughly 70% below the $0.15 / $0.60 that brand-name hosts list, with a 131K-token context window. Reasoning is always on for this endpoint and cannot be disabled — control depth with `reasoning_effort`.

$0.041 input · $0.187 output per 1M tokens
ByRaullen Chai·Updated

At a glance

Context
131K tokens
Input / 1M
$0.041
Output / 1M
$0.187
Thinks by default
Yes

High-volume production agents on OpenAI's open-weight 117B MoE — proven at hundreds of billions of tokens.

Pricing comparison ($/1M tokens)

ProviderInputOutputvs QSP
QuickSilver Pro$0.041$0.187—
OpenRouter low rate (openai/gpt-oss-120b)$0.037$0.1710% more expensive
OpenAI (GPT-4o mini)$0.15$0.6069% lower

When to use

Reach for GPT-OSS 120B when you want OpenAI-grade instruction following at open-model prices: high-volume agent loops, tool-heavy automation, background summarization and classification at scale. Its 117B MoE activates only 5.1B parameters per token, so it's fast, and it is battle-tested — it powers some of the largest public agent deployments running today. Reasoning is always on; use `reasoning_effort` to trade depth against latency and cost.

When to use something else

Its 131K-token window is mid-sized — for whole-codebase context, GLM 5.3 Flash and Qwen3.8 27B carry the same price class with far larger windows. For frontier-difficulty reasoning, step up to GPT-5.6 Sol or Claude Opus 5. Reasoning is mandatory here, so trivial high-frequency calls spend thinking tokens — DeepSeek V4 Flash ($0.086/$0.173) replies directly without the overhead.

Quickstart (curl)

curl https://api.quicksilverpro.io/v1/chat/completions \
  -H "Authorization: Bearer $QSP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-oss-120b",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

OpenAI-compatible. One-line migration via base_url.

FAQ

Reasoning is always on — it is mandatory on this endpoint and cannot be disabled (a request that tries to turn it off is rejected upstream). Use `reasoning_effort` to trade depth against latency and cost. Reasoning tokens bill as output tokens and draw from your `max_tokens` budget before the visible answer, so set `max_tokens` generously or the reply can come back truncated.

Yes — the open-weight 117B mixture-of-experts model OpenAI published, served as a hosted OpenAI-compatible endpoint. You get the model without running 120B-class infrastructure yourself, with streaming, tool calling, and per-token billing on your QuickSilver Pro balance.

Yes — GPT-OSS 120B is an OpenAI-compatible chat completions endpoint on QuickSilver Pro (the Responses API works too). Set base_url=https://api.quicksilverpro.io/v1, paste your QSP key, and use model="gpt-oss-120b". Streaming, tool calling, and usage.cost accounting all work.

Brand-name hosts list GPT-OSS 120B at $0.15 input / $0.60 output per million tokens; QuickSilver Pro is $0.041 / $0.187 — roughly 70% below that tier, and about 10% above the cheapest OpenRouter listings. Same OpenAI-compatible surface; migration is a base_url + key swap, dropping the `openai/` provider prefix from the model ID.

Try GPT-OSS 120B with double credits — up to $50 in bonus credits

Get API Key