Claude Haiku 5.5 on QuickSilver Pro
Claude Haiku 5.5: $0.05 input / $0.25 output per 1M tokens, 50% below Anthropic list ($0.10 / $0.50). Context: 1M tokens. Maximum output: 128K tokens. Available through Chat Completions and Responses. Capabilities: text input, text output, image input, streaming and tool calling.
Try it free: included in the $0.05 trial credit, no card
At a glance
For developers building subagents and high-volume chat or extraction workflows.
Pricing comparison ($/1M tokens)
| Provider | Input | Output | vs QSP |
|---|---|---|---|
| QuickSilver Pro | $0.05 | $0.25 | lowest-cost |
| Anthropic list price (Anthropic API) | $0.10 | $0.50 | 50% lower |
When to use
Use Claude Haiku 5.5 for subagents and high-volume workflows. Context: 1M tokens; maximum output: 128K tokens. The gateway drops `temperature` and `top_p` and rewrites assistant prefill as a user continuation.
When to use something else
For reasoning support, use Claude Sonnet 5.5: 1M context / 128K maximum output tokens at 20× this model’s base output price. Use Claude Haiku 4.5 when you need `temperature` or assistant prefill without a rewrite; its context is 200K tokens. Above 200K prompt tokens, this model bills the whole request at $0.25 input / $1.25 output per million tokens.
Quickstart (curl)
curl https://api.quicksilverpro.io/v1/chat/completions \
-H "Authorization: Bearer $QSP_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-haiku-5-5",
"messages": [{"role": "user", "content": "Hello!"}]
}'import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.quicksilverpro.io/v1",
api_key=os.environ["QSP_KEY"],
)
response = client.chat.completions.create(
model="claude-haiku-5-5",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)OpenAI-compatible. One-line migration via base_url.
FAQ
Claude Haiku 5.5 costs $0.05 input / $0.25 output per 1M tokens. Anthropic list is $0.10 / $0.50; the discount is 50% on both base rates.
For Claude Haiku 5.5, prompts above 200K tokens use $0.25 input / $1.25 output per 1M tokens for the request. Below or at that threshold, base rates are $0.05 / $0.25. The 1M-token context limit is separate from the billing threshold.
The gateway drops `temperature` and `top_p`. This model rejects a trailing assistant prefill, so the gateway rewrites it as a user continuation turn carrying the prefix.
The gateway retains a forced `tool_choice` on this model. Claude `response_format` with `json_schema` is advisory. Use a tool with `strict: true` for enforced argument shape; forced tool choice can require that tool call.
Catalog capabilities: text input, text output, image input, streaming and tool calling. For Claude image requests, the gateway fetches public remote image URLs and inlines the image data. An inaccessible or rejected URL fails the request; inline base64 is also supported.
Set `base_url="https://api.quicksilverpro.io/v1"`, use your QSP API key, and set `model="claude-haiku-5-5"`. Supported APIs: Chat Completions and Responses. See the runnable quickstart on this page and the API docs.