GLM 5.3 Flash on QuickSilver Pro
GLM 5.3 Flash is Z.ai's fast open-weights sibling of GLM 5.3 — the same 2026 reasoning family distilled to workhorse pricing, with the full 1M-token context window and 131K max output. On QuickSilver Pro it's $0.06 input / $0.20 output per million tokens, ~20% below Z.ai's own list price of $0.075 / $0.25. Like GLM 5.3, reasoning is always on for this endpoint and cannot be disabled — control depth with `reasoning_effort` instead.
At a glance
High-volume agents and coding at workhorse pricing — the fast open-weights member of Z.ai's newest reasoning family, with a 1M-token context window.
Pricing comparison ($/1M tokens)
| Provider | Input | Output | vs QSP |
|---|---|---|---|
| QuickSilver Pro | $0.06 | $0.20 | lowest-cost |
| Z.ai (z-ai/glm-5.3-flash) | $0.075 | $0.25 | 20% lower |
| OpenAI (GPT-4o mini) | $0.15 | $0.60 | 67% lower |
When to use
Reach for GLM 5.3 Flash when you want the GLM 5.3 family's reasoning-and-tools behavior at a fraction of the price: high-volume agent loops, background codegen, summarization over huge working sets in its 1M-token context, and batch pipelines where per-call cost dominates. At $0.06 / $0.20 per million tokens it competes with the cheapest tier of the catalog while keeping a 1M-token window — a combination the budget models don't offer. Reasoning is always on; use `reasoning_effort` to trade depth against latency and cost.
When to use something else
For the hardest repo-scale engineering and long-horizon plans, GLM 5.3 is the stronger family pick — same window, deeper model. Reasoning is mandatory here too, so it spends thinking tokens even on trivial turns; if you want a cheap model that replies directly, DeepSeek V4 Flash ($0.086/$0.173) answers without the overhead. If you have prompts tuned on GLM 5.2 or GLM 5.3, A/B before switching down — traces and tool-call cadence differ across the family.
Quickstart (curl)
curl https://api.quicksilverpro.io/v1/chat/completions \
-H "Authorization: Bearer $QSP_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3-flash",
"messages": [{"role": "user", "content": "Hello!"}]
}'OpenAI-compatible. One-line migration via base_url.
FAQ
Reasoning is always on — it is mandatory on this endpoint and cannot be disabled (a request that tries to turn it off is rejected upstream). Use `reasoning_effort` to trade depth against latency and cost. Reasoning tokens bill as output tokens and are drawn from your `max_tokens` budget before the visible answer, so set `max_tokens` generously or the reply can come back truncated. If you want a GLM without always-on thinking, use GLM 5.2.
Same family, different weight class: GLM 5.3 Flash is the fast, open-weights variant released 2026-08-26, priced at $0.06 input / $0.20 output per million tokens versus $1.12 / $3.52 for GLM 5.3 — roughly 18x cheaper on input. Both share the 1M-token context window, the 131K max-output limit, and the always-on reasoning contract. Use Flash for volume; step up to GLM 5.3 when task difficulty, not throughput, is the bottleneck.
Yes — GLM 5.3 Flash is an OpenAI-compatible chat completions endpoint on QuickSilver Pro (the Responses API works too). Set base_url=https://api.quicksilverpro.io/v1, paste your QSP key, and use model="glm-5.3-flash". Streaming, tool calling (`tool_choice: "auto"` — this endpoint rejects forcing a specific function), and usage.cost accounting all work. `response_format: json_schema` is not enforced on this endpoint yet; if you need strict structured output, use GLM 5.2.
Z.ai lists GLM 5.3 Flash at $0.075 input / $0.25 output per million tokens; QuickSilver Pro is $0.06 / $0.20, ~20% below on both legs. Same OpenAI-compatible surface; migration is a base_url + key swap, dropping the `z-ai/` provider prefix from the model ID.