होम/Models/Gemini 3.1 Flash Lite
Legacy1M contextMultimodalResponses API

Gemini 3.1 Flash Lite QuickSilver Pro पर

Gemini 3.1 Flash Lite केवल मौजूदा integrations के साथ compatibility के लिए उपलब्ध है। Production और नए workloads को gemini-3.5-flash-lite पर migrate करें, जो Google का मौजूदा low-cost Flash-Lite GA मॉडल है।

प्रति 1M tokens: input $0.2125 · output $1.275
लेखक:Raullen Chai·अपडेट:

एक नज़र में

Context
1M tokens
Input / 1M
$0.2125
Output / 1M
$1.275
Default में सोचता है
नहीं

Gemini 3.5 Flash-Lite पर migrate हो रहे integrations के लिए अस्थायी compatibility।

Pricing तुलना ($/1M tokens)

ProviderInputOutputQSP की तुलना में
QuickSilver Pro$0.2125$1.275सबसे कम कीमत
OpenRouter (google/gemini-3.1-flash-lite)$0.25$1.5015% कम
OpenAI (GPT-4o mini)$0.15$0.60112% महँगा

कब इस्तेमाल करें

3.1 Flash Lite का उपयोग high-volume, cost-sensitive काम के लिए करें जहाँ आपको reasoning trace नहीं चाहिए: routing और classification, extraction, summarization, साधारण chat, और agent sub-tasks जहाँ latency और कीमत raw reasoning depth से ज़्यादा मायने रखते हैं। डिफ़ॉल्ट रूप से non-thinking होने का मतलब है कि output tokens predictable हैं — बड़े scale पर budget बनाना आसान।

कब कोई और model चुनें

Multi-step reasoning, कठिन coding या analysis के लिए 3.5 Flash ($1.275/$7.65) या किसी Pro tier पर जाएँ — Flash Lite लागत के बदले depth छोड़ता है। अगर आपको खास तौर पर कम लागत में thinking वाला Gemini चाहिए, तो 3 Flash Preview ($0.425/$2.55) डिफ़ॉल्ट रूप से reasoning करता है। Image generation के लिए Gemini image models या FLUX इस्तेमाल करें।

Quickstart (curl)

curl https://api.quicksilverpro.io/v1/chat/completions \
  -H "Authorization: Bearer $QSP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.1-flash-lite",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

OpenAI-संगत। base_url बदलकर एक लाइन में migration।

FAQ

नहीं — Flash Lite non-thinking tier है, इसलिए यह reasoning trace के बिना सीधे जवाब देता है। इससे output token counts (और लागत) predictable रहते हैं, और high-volume workloads को ठीक यही चाहिए। अगर आपको reasoning चाहिए, तो 3 Flash Preview या 3.5 Flash डिफ़ॉल्ट रूप से सोचते हैं।

हाँ — प्रति 1M tokens $0.2125 input / $1.275 output पर यह catalog का सबसे कम लागत वाला Gemini है और high-volume routing, classification और extraction के लिए स्वाभाविक पसंद है। Gemini के अलावा low-cost chat के लिए DeepSeek V4 Flash ($0.086/$0.173) input और output दोनों पर इससे भी सस्ता है।

QuickSilver Pro पर 3.1 Flash Lite की कीमत प्रति 1M tokens $0.2125 input / $1.275 output है — Vertex retail की कीमत से ~15% कम; उसकी कीमत है $0.25/$1.50। 18 models के लिए एक OpenAI-compatible key, एक bill, और हर response पर `usage.cost` field ताकि आप हर request का खर्च reconcile कर सकें।

Gemini 3.1 Flash Lite को double credits के साथ आज़माएँ — $50 तक bonus credits

API Key पाएँ