Gemini 3.1 Flash Lite QuickSilver Pro पर
Gemini 3.1 Flash Lite केवल मौजूदा integrations के साथ compatibility के लिए उपलब्ध है। Production और नए workloads को gemini-3.5-flash-lite पर migrate करें, जो Google का मौजूदा low-cost Flash-Lite GA मॉडल है।
एक नज़र में
Gemini 3.5 Flash-Lite पर migrate हो रहे integrations के लिए अस्थायी compatibility।
Pricing तुलना ($/1M tokens)
| Provider | Input | Output | QSP की तुलना में |
|---|---|---|---|
| QuickSilver Pro | $0.2125 | $1.275 | सबसे कम कीमत |
| OpenRouter (google/gemini-3.1-flash-lite) | $0.25 | $1.50 | 15% कम |
| OpenAI (GPT-4o mini) | $0.15 | $0.60 | 112% महँगा |
कब इस्तेमाल करें
3.1 Flash Lite का उपयोग high-volume, cost-sensitive काम के लिए करें जहाँ आपको reasoning trace नहीं चाहिए: routing और classification, extraction, summarization, साधारण chat, और agent sub-tasks जहाँ latency और कीमत raw reasoning depth से ज़्यादा मायने रखते हैं। डिफ़ॉल्ट रूप से non-thinking होने का मतलब है कि output tokens predictable हैं — बड़े scale पर budget बनाना आसान।
कब कोई और model चुनें
Multi-step reasoning, कठिन coding या analysis के लिए 3.5 Flash ($1.275/$7.65) या किसी Pro tier पर जाएँ — Flash Lite लागत के बदले depth छोड़ता है। अगर आपको खास तौर पर कम लागत में thinking वाला Gemini चाहिए, तो 3 Flash Preview ($0.425/$2.55) डिफ़ॉल्ट रूप से reasoning करता है। Image generation के लिए Gemini image models या FLUX इस्तेमाल करें।
Quickstart (curl)
curl https://api.quicksilverpro.io/v1/chat/completions \
-H "Authorization: Bearer $QSP_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.1-flash-lite",
"messages": [{"role": "user", "content": "Hello!"}]
}'OpenAI-संगत। base_url बदलकर एक लाइन में migration।
FAQ
नहीं — Flash Lite non-thinking tier है, इसलिए यह reasoning trace के बिना सीधे जवाब देता है। इससे output token counts (और लागत) predictable रहते हैं, और high-volume workloads को ठीक यही चाहिए। अगर आपको reasoning चाहिए, तो 3 Flash Preview या 3.5 Flash डिफ़ॉल्ट रूप से सोचते हैं।
हाँ — प्रति 1M tokens $0.2125 input / $1.275 output पर यह catalog का सबसे कम लागत वाला Gemini है और high-volume routing, classification और extraction के लिए स्वाभाविक पसंद है। Gemini के अलावा low-cost chat के लिए DeepSeek V4 Flash ($0.086/$0.173) input और output दोनों पर इससे भी सस्ता है।
QuickSilver Pro पर 3.1 Flash Lite की कीमत प्रति 1M tokens $0.2125 input / $1.275 output है — Vertex retail की कीमत से ~15% कम; उसकी कीमत है $0.25/$1.50। 18 models के लिए एक OpenAI-compatible key, एक bill, और हर response पर `usage.cost` field ताकि आप हर request का खर्च reconcile कर सकें।