Home/Migrate/From Together AI
Migration guide · 5 minutes

Together AI → QuickSilver Pro

Both APIs speak the OpenAI chat-completions shape, so the move is a base-URL swap. QuickSilver Pro serves the latest DeepSeek V4 wave — V4 Flash for low-cost chat and V4 Pro for premium reasoning — well below Together's own-GPU rates. For the full side-by-side analysis, see /vs/together-ai.

The steps

  1. 1

    Get a QuickSilver Pro API key

    Sign up at quicksilverpro.io/dashboard. First top-up bonus: we match your first top-up 100%, up to $50 — added to your balance when you top up again (any amount from $5).

  2. 2

    Change the base URL

    In your OpenAI SDK init, swap the base_url.

    - base_url="https://api.together.xyz/v1"
    + base_url="https://api.quicksilverpro.io/v1"
  3. 3

    Swap the API key

    Replace your Together token with a QuickSilver Pro key.

    - api_key=os.environ["TOGETHER_API_KEY"],
    + api_key=os.environ["QSP_KEY"],
  4. 4

    Rename model IDs

    Together prefixes model IDs with the originating org. Drop the prefix and use the QuickSilver Pro short name.

    Together AIQuickSilver Pro
    deepseek-ai/DeepSeek-V4-Prodeepseek-v4-pro
  5. 5

    Test your core flows end-to-end

    Run one representative request for each feature you use — chat, streaming, tool / function calling, and json_schema strict mode. Apart from the known difference below, any behavioral difference is a bug — report it.

    One known difference: json_schema strict mode is model-dependent. The Claude models do not support it, so a schema there is advisory and the enforced path is a tool with strict: true — structured output.

Full before/after

Before · Together AI
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.together.xyz/v1",
    api_key=os.environ["TOGETHER_API_KEY"],
)

r = client.chat.completions.create(
    model="deepseek-ai/DeepSeek-V4-Pro",
    messages=[{"role": "user", "content": "Hi"}],
)
After · QuickSilver Pro
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.quicksilverpro.io/v1",
    api_key=os.environ["QSP_KEY"],
)

r = client.chat.completions.create(
    model="deepseek-v4-pro",
    messages=[{"role": "user", "content": "Hi"}],
)

What you'll pay after switching

Per 1M tokens, input / output. QuickSilver Pro rates vs Together AI's published per-token pricing.

ModelQuickSilver ProTogether AISavings
DeepSeek V4 Pro$0.70 / $2.10$1.74 / $3.48~40% output

Common migration pitfalls

⚠
Fine-tunes and dedicated endpoints don't move
QuickSilver Pro serves shared open-source models only. If you call Together fine-tuned models or reserved dedicated GPU endpoints, those have no QSP equivalent — keep them on Together and run both SDKs side-by-side.
⚠
Minimum top-up is lower, not a blocker
Together's minimum top-up is $25; QuickSilver Pro's is $5. Nothing breaks — your first top-up is also matched 100% (up to $50), credited on your next top-up.
⚠
Model generation may differ
Together serves V4 Pro with a 512K context window. QuickSilver Pro serves V4 Pro with its 1M context and also offers the official V4 Flash 0731 agent build. Re-run evals after switching if your prompts are sensitive to provider or quantization differences.
⚠
Rate limits work differently
QuickSilver Pro applies per-key throughput caps (default 600 req/min, 1M tok/min, 8 parallel). For bursty traffic, enable retry-on-429 in your client and request a higher limit if needed.

Migrating from Together AI — FAQ

How much lower is QuickSilver Pro on DeepSeek?
On the same V4 Pro model, QuickSilver Pro is $0.70 / $2.10 per 1M tokens versus Together's current $1.74/$3.48 rate — about 60% lower on input and 40% lower on output. The official V4 Flash 0731 build is $0.086 / $0.173 versus Together's $0.14/$0.28, about 38% lower on both legs.
How do I migrate from Together AI?
Change base_url from api.together.xyz/v1 to api.quicksilverpro.io/v1 and swap the API key. Model ID mappings: deepseek-ai/DeepSeek-V4-Pro -> deepseek-v4-pro.
When should I stay on Together AI?
If you fine-tune custom models, reserve dedicated GPU endpoints, use Llama or Mistral, or need embeddings — we do not serve those. Image generation we do: Gemini 3 Pro Image, FLUX.2 Pro, FLUX.1 Schnell, SDXL Turbo, FLUX.2 Klein, Qwen-Image Max, Seedream 5.0 Pro, Seedream 4 and Bria FIBO 1.5.
Same OpenAI features?
Yes for chat: streaming, tools, and usage.cost all work through the official OpenAI SDK. json_schema strict mode is model-dependent — the Claude models do not support it, so use a tool with `strict: true` there.

Other migration guides

Need help?

Email [email protected] — a human replies usually within 4 hours. For the broader analysis, see QuickSilver Pro vs Together AI.

Start saving in 5 minutes

First top-up matched 100%, up to $50 in bonus credits, added on your next top-up. Keep your code on the OpenAI SDK — only the base URL and key change.

Get API Key