Nemotron 3 Ultra on QuickSilver Pro
Nemotron 3 Ultra is NVIDIA's open frontier-reasoning model: a 550B-parameter Mixture-of-Experts with 55B active parameters per token, built on a hybrid Transformer-Mamba architecture for deep reasoning and agent orchestration. QuickSilver Pro provides 262K-token context at $0.50 input / $2.20 output per million tokens, with zero data retention — the same as the lowest stable paid market rate.
At a glance
NVIDIA's largest open model for hard reasoning and multi-step agents — 550B total parameters, 55B active per token, and 262K context.
Pricing comparison ($/1M tokens)
| Provider | Input | Output | vs QSP |
|---|---|---|---|
| QuickSilver Pro | $0.50 | $2.20 | — |
| Market price (public paid rate) | $0.50 | $2.20 | same |
When to use
Use Nemotron 3 Ultra when an open-weight model has to carry difficult work: multi-step planning, orchestrating other agents and tools, long analytical answers, and coding tasks where a small efficiency model runs out of depth. It supports function calling and JSON Schema structured output, and cached input costs $0.10 per million tokens.
When to use something else
It is a large model and generates more slowly than efficiency models, so it is a poor fit for latency-sensitive chat or high-volume simple extraction — use Nemotron 3.5 Lightning for those. A single response is capped at 16,384 output tokens. The QSP endpoint is text-only and currently guarantees 262K-token context with zero data retention.
Quickstart (curl)
curl https://api.quicksilverpro.io/v1/chat/completions \
-H "Authorization: Bearer $QSP_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nemotron-3-ultra",
"messages": [{"role": "user", "content": "Hello!"}]
}'OpenAI-compatible. One-line migration via base_url.
FAQ
$0.50 per million input tokens, $2.20 per million output tokens, and $0.10 per million cached-input tokens — the same as the lowest stable paid market rate, with no markup.
They sit at opposite ends of the same family. Lightning is a 30B model with 3B active parameters, tuned for speed and cost. Ultra is a 550B model with 55B active parameters, tuned for reasoning depth and orchestration. Start with Lightning and escalate to Ultra when your evaluations show the smaller model failing.
Yes. It supports streaming, function calling (including a forced tool choice), and JSON Schema structured output. QSP serves it with reasoning off, so responses are direct and carry no separate reasoning trace or hidden reasoning tokens.
Set base_url=https://api.quicksilverpro.io/v1, use your QSP key, and set model="nemotron-3-ultra". Both Chat Completions and Responses API clients use the same public model ID.