Gemini 3.1 Pro Preview QuickSilver Pro पर
Gemini 3.1 Pro Preview Google का flagship reasoning मॉडल है — 1M-token context window, डिफ़ॉल्ट रूप से गहरी thinking, multimodal foundation। QuickSilver Pro पर इसकी कीमत प्रति 1M tokens $1.70 input / $10.20 output है, यानी Google के Vertex retail और OpenRouter की कीमत से ~15% कम; उनकी कीमत है $2/$12। यह एक preview API है; GA से पहले semantics बदल सकती है।
एक नज़र में
Long-context reasoning, multi-step analysis, जटिल code review — o1-class quality पर और ~6x कम output लागत में।
Pricing तुलना ($/1M tokens)
| Provider | Input | Output | QSP की तुलना में |
|---|---|---|---|
| QuickSilver Pro | $1.70 | $10.20 | सबसे कम कीमत |
| OpenRouter (google/gemini-3.1-pro-preview) | $2.00 | $12.00 | 15% कम |
| OpenAI (o1) | $15.00 | $60.00 | 83% कम |
कब इस्तेमाल करें
Gemini 3.1 Pro Preview तब चुनें जब task को लंबे context पर chain-of-thought से सचमुच फ़ायदा हो: बड़े codebase का review, कई documents का legal/research synthesis, rich memory के साथ जटिल agentic planning, लंबी समस्या पर mathematical reasoning। 1M context में बिना chunking के अच्छा-खासा corpus आ जाता है, और thinking trace कठिन समस्याओं पर जवाब की quality बढ़ाता है। OpenAI o1 के बराबर quality tier, लागत के एक छोटे हिस्से में।
कब कोई और model चुनें
Routine chat या short-prompt classification के लिए 3.1 Pro Preview को default न बनाएँ — यह डिफ़ॉल्ट रूप से सोचता है, इसलिए एक लाइन का सवाल भी दिखने वाले जवाब से पहले 100-200 reasoning tokens खर्च कर सकता है। High-volume low-cost chat के लिए Gemini 3.1 Flash Lite ($0.2125/$1.275) या DeepSeek V4 Flash ($0.086/$0.173) cost-per-task में बेहतर हैं। Gemini से अलग शैली की reasoning के लिए DeepSeek V4 Pro ($0.70/$2.10) input और output दोनों पर सस्ता है।
Quickstart (curl)
curl https://api.quicksilverpro.io/v1/chat/completions \
-H "Authorization: Bearer $QSP_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.1-pro-preview",
"messages": [{"role": "user", "content": "Hello!"}]
}'OpenAI-संगत। base_url बदलकर एक लाइन में migration।
FAQ
यह एक preview API है — Google इसे यही label देता है, और GA से पहले semantics बदल सकती है। Output formats, thinkingConfig handling और rate limits बिना सूचना के बदल सकते हैं। Prototyping और evals के लिए आज ही ship करें। Revenue-critical paths के लिए किसी Gemini GA मॉडल (जैसे Gemini 3.5 Flash) पर pin करें और जब तक Google इसे promote नहीं करता, 3.1 Pro को flag के पीछे का experiment मानें।
Gemini 3.1 Pro Preview डिफ़ॉल्ट रूप से सोचता है — मॉडल अपने अंतिम जवाब से पहले एक लंबा reasoning trace निकालता है, और usage.completion_tokens में reasoning_tokens और text_tokens दोनों शामिल होते हैं। हमारे smoke test में एक साधारण "what is 7+8?" पर 160 reasoning + 3 text tokens आए। max_tokens का budget उसी हिसाब से रखें: ~200 tokens से कम पर यह जोखिम रहता है कि मॉडल का budget thinking phase में ही खत्म हो जाए और वह खाली जवाब लौटाए।
QuickSilver Pro पर अभी नहीं — Gemini 3.1 Pro पर `reasoning: { enabled: false }` चुपचाप drop कर दिया जाता है। बिना thinking वाले Gemini के लिए gemini-3.1-flash-lite ($0.2125/$1.275) इस्तेमाल करें, जो सचमुच non-thinking है। ध्यान दें: Flash tiers डिफ़ॉल्ट रूप से सोचते हैं — केवल Flash Lite बिना reasoning के आता है।
QuickSilver Pro पर Gemini 3.1 Pro Preview की कीमत प्रति 1M tokens $1.70 input / $10.20 output है — Google के Vertex retail और OpenRouter की कीमत से लगभग 15% कम; उनकी कीमत है $2/$12। आप एक ही OpenAI-compatible key पर 14 models के लिए एक bill चुकाते हैं, हर response पर `usage.cost` accounting के साथ। किसी दूसरे provider से switch करने के लिए OpenAI SDK में base_url + key बदलें और `model="gemini-3.1-pro-preview"` रखें।