Home/Models/GLM 5.3 Flash
1M contextReasoningResponses API

GLM 5.3 Flash on QuickSilver Pro

GLM 5.3 Flash is Z.ai's fast open-weights sibling of GLM 5.3 — the same 2026 reasoning family distilled to workhorse pricing, with the full 1M-token context window and 131K max output. On QuickSilver Pro it's $0.06 input / $0.20 output per million tokens, ~20% below Z.ai's own list price of $0.075 / $0.25. Like GLM 5.3, reasoning is always on for this endpoint and cannot be disabled — control depth with `reasoning_effort` instead.

$0.06 input · $0.20 output per 1M tokens
ByRaullen Chai·Updated

At a glance

Context
1M tokens
Input / 1M
$0.06
Output / 1M
$0.20
Thinks by default
Yes

High-volume agents and coding at workhorse pricing — the fast open-weights member of Z.ai's newest reasoning family, with a 1M-token context window.

Pricing comparison ($/1M tokens)

ProviderInputOutputvs QSP
QuickSilver Pro$0.06$0.20lowest-cost
Z.ai (z-ai/glm-5.3-flash)$0.075$0.2520% lower
OpenAI (GPT-4o mini)$0.15$0.6067% lower

When to use

Reach for GLM 5.3 Flash when you want the GLM 5.3 family's reasoning-and-tools behavior at a fraction of the price: high-volume agent loops, background codegen, summarization over huge working sets in its 1M-token context, and batch pipelines where per-call cost dominates. At $0.06 / $0.20 per million tokens it competes with the cheapest tier of the catalog while keeping a 1M-token window — a combination the budget models don't offer. Reasoning is always on; use `reasoning_effort` to trade depth against latency and cost.

When to use something else

For the hardest repo-scale engineering and long-horizon plans, GLM 5.3 is the stronger family pick — same window, deeper model. Reasoning is mandatory here too, so it spends thinking tokens even on trivial turns; if you want a cheap model that replies directly, DeepSeek V4 Flash ($0.086/$0.173) answers without the overhead. If you have prompts tuned on GLM 5.2 or GLM 5.3, A/B before switching down — traces and tool-call cadence differ across the family.

Quickstart (curl)

curl https://api.quicksilverpro.io/v1/chat/completions \
  -H "Authorization: Bearer $QSP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

OpenAI-compatible. One-line migration via base_url.

FAQ

Reasoning is always on — it is mandatory on this endpoint and cannot be disabled (a request that tries to turn it off is rejected upstream). Use `reasoning_effort` to trade depth against latency and cost. Reasoning tokens bill as output tokens and are drawn from your `max_tokens` budget before the visible answer, so set `max_tokens` generously or the reply can come back truncated. If you want a GLM without always-on thinking, use GLM 5.2.

Same family, different weight class: GLM 5.3 Flash is the fast, open-weights variant released 2026-08-26, priced at $0.06 input / $0.20 output per million tokens versus $1.12 / $3.52 for GLM 5.3 — roughly 18x cheaper on input. Both share the 1M-token context window, the 131K max-output limit, and the always-on reasoning contract. Use Flash for volume; step up to GLM 5.3 when task difficulty, not throughput, is the bottleneck.

Yes — GLM 5.3 Flash is an OpenAI-compatible chat completions endpoint on QuickSilver Pro (the Responses API works too). Set base_url=https://api.quicksilverpro.io/v1, paste your QSP key, and use model="glm-5.3-flash". Streaming, tool calling (`tool_choice: "auto"` — this endpoint rejects forcing a specific function), and usage.cost accounting all work. `response_format: json_schema` is not enforced on this endpoint yet; if you need strict structured output, use GLM 5.2.

Z.ai lists GLM 5.3 Flash at $0.075 input / $0.25 output per million tokens; QuickSilver Pro is $0.06 / $0.20, ~20% below on both legs. Same OpenAI-compatible surface; migration is a base_url + key swap, dropping the `z-ai/` provider prefix from the model ID.

Try GLM 5.3 Flash with double credits — up to $50 in bonus credits

Get API Key