Home/Models/Mistral Large 4
524,288 contextMultimodalReasoningResponses API

Mistral Large 4 on QuickSilver Pro

Mistral Large 4 is Mistral AI's frontier multimodal model, released in preview on October 6, 2026: text and image input, a 524,288-token context, and a build aimed at coding, reasoning and agentic work. QuickSilver Pro serves Mistral's own endpoint at $0.68 input / $2.09 output per million tokens — the same as the model's list price on OpenRouter, with no markup.

$0.68 input · $2.09 output per 1M tokens
ByRaullen Chai·Updated

At a glance

Context
524,288 tokens
Input / 1M
$0.68
Output / 1M
$2.09
Thinks by default
No

Mistral's flagship for coding and agents — text and image input, 524,288 context, and reasoning you switch on per request.

Pricing comparison ($/1M tokens)

ProviderInputOutputvs QSP
QuickSilver Pro$0.68$2.09—
OpenRouter list price (mistralai/mistral-large-4-0)$0.68$2.09same

When to use

Use Mistral Large 4 for agent loops, multi-file coding and long-document work where you want a current European flagship: the 524,288-token context holds a large repository or a stack of reports, it reads screenshots and diagrams alongside text, and it supports function calling (including a forced tool choice) and JSON Schema structured output. Answers are direct by default; for harder problems pass `reasoning: {"enabled": true}` and the model thinks before it answers. Cached input costs $0.07 per million tokens.

When to use something else

Reasoning mode is thorough and can spend thousands of output tokens before the answer starts, so leave it off for routine chat and set a generous `max_tokens` when you turn it on. For high-volume simple extraction a flash-tier model such as DeepSeek V4 Flash ($0.086/$0.173) costs far less. A single response is capped at 131,072 output tokens. This is a preview release served from Mistral's own endpoint, which does not offer zero data retention, and behavior may change before general availability.

Quickstart (curl)

curl https://api.quicksilverpro.io/v1/chat/completions \
  -H "Authorization: Bearer $QSP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistral-large-4",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

OpenAI-compatible. One-line migration via base_url.

FAQ

$0.68 per million input tokens, $2.09 per million output tokens, and $0.07 per million cached-input tokens — the same as the model's list price on OpenRouter, with no markup. Reasoning tokens are billed as output tokens.

No. By default it answers directly and bills no reasoning tokens. Pass `reasoning: {"enabled": true}` in the request to turn thinking on for that call; the reasoning trace comes back in a separate field and its tokens count toward `max_tokens` and are billed at the output rate.

Yes. It accepts text and images as input and returns text. It supports streaming, function calling (including a forced tool choice), and JSON Schema structured output. Send images as URLs or inline as base64 data URLs.

Set base_url=https://api.quicksilverpro.io/v1, use your QSP key, and set model="mistral-large-4". Both Chat Completions and Responses API clients use the same public model ID.

Try Mistral Large 4 with double credits — up to $50 in bonus credits

Get API Key