Home/Models/Claude Haiku 4.5
200K contextMultimodalResponses API

Claude Haiku 4.5 on QuickSilver Pro

Claude Haiku 4.5: $0.40 input / $2.00 output per 1M tokens, 60% below Anthropic list ($1.00 / $5.00). Context: 200K tokens.

$0.40 input · $2.00 output per 1M tokens
Get API Key

Unlocks with your first top-up from $5

ByRaullen Chai·Updated

At a glance

Context
200K tokens
Input / 1M
$0.40
Output / 1M
$2.00
Thinks by default
No

For developers building subagents and high-volume chat or extraction workflows.

Pricing comparison ($/1M tokens)

ProviderInputOutputvs QSP
QuickSilver Pro$0.40$2.00lowest-cost
Anthropic list price (Anthropic API)$1.00$5.0060% lower

When to use

Use Claude Haiku 4.5 for high-volume workflows. Context: 200K tokens; the catalog does not publish a maximum-output limit. The gateway retains `temperature` and assistant prefill, but drops `top_p`.

When to use something else

Switch to the successor Claude Haiku 5.5: it has 1M context / 128K maximum output tokens, with base output priced at 0.125× this model. Above 200K prompt tokens, Claude Haiku 5.5 bills the whole request at $0.25 input / $1.25 output per million tokens. The gateway drops `temperature` and `top_p` and rewrites assistant prefill as a user continuation.

Quickstart (curl)

curl https://api.quicksilverpro.io/v1/chat/completions \
  -H "Authorization: Bearer $QSP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-haiku-4-5",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
python (OpenAI SDK)
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.quicksilverpro.io/v1",
    api_key=os.environ["QSP_KEY"],
)
response = client.chat.completions.create(
    model="claude-haiku-4-5",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

OpenAI-compatible. One-line migration via base_url.

FAQ

Claude Haiku 4.5 costs $0.40 input / $2.00 output per 1M tokens. Anthropic list is $1.00 / $5.00; the discount is 60% on both base rates.

The catalog has no separate long-context tier for Claude Haiku 4.5. Base rates apply within its 200K-token context.

The gateway drops `top_p`, but retains `temperature`. This model accepts assistant prefill without the continuation rewrite.

The gateway retains a forced `tool_choice` on this model. Claude `response_format` with `json_schema` is advisory. Use a tool with `strict: true` for enforced argument shape; forced tool choice can require that tool call.

Catalog capabilities: text input, text output, image input, streaming and tool calling. For Claude image requests, the gateway fetches public remote image URLs and inlines the image data. An inaccessible or rejected URL fails the request; inline base64 is also supported.

Claude · Models

Claude Haiku 4.5

Unlocks with your first top-up from $5

Get API Key