होम/Models/Nemotron 3.5 Lightning
262K contextResponses API

Nemotron 3.5 Lightning QuickSilver Pro पर

Nemotron 3.5 Lightning NVIDIA का open 30B-parameter Mixture-of-Experts मॉडल है, जिसमें प्रति token केवल 3B active parameters हैं। यह तेज़, tool-heavy agents, high-throughput automation, और domain customization के लिए बना है। QuickSilver Pro 262K-token context देता है, $0.066 input / $0.176 output प्रति million tokens पर, zero data retention के साथ — सबसे कम stable paid market rate से लगभग 10% ऊपर।

प्रति 1M tokens: input $0.066 · output $0.176
लेखक:Raullen Chai·अपडेट:

एक नज़र में

Context
262K tokens
Input / 1M
$0.066
Output / 1M
$0.176
Default में सोचता है
नहीं

तेज़, सस्ते tool-using agents — कुल 30B parameters, प्रति token 3B active, और 262K context।

Pricing तुलना ($/1M tokens)

ProviderInputOutputQSP की तुलना में
QuickSilver Pro$0.066$0.176—
बाज़ार मूल्य (public paid rate)$0.06$0.1610% महँगा
OpenAI (GPT-4o-mini)$0.15$0.6071% कम

कब इस्तेमाल करें

Nemotron 3.5 Lightning को high-volume coding assistants, retrieval agents, लंबे समय तक चलने वाले tool loops, structured extraction, और ऐसे workflows के लिए इस्तेमाल करें जो बड़े system prompt या repository context को बार-बार reuse करते हैं। इसका sparse 30B-A3B design generation को सस्ता रखता है, और cached input की कीमत $0.033 प्रति million tokens है।

कब कोई और model चुनें

यह तेज़ efficiency मॉडल है, frontier flagship नहीं। कठिन research, गणितीय reasoning, या complex autonomous coding को DeepSeek V4 Pro, Qwen3.8 Max, या Grok 4.6 पर तब ले जाएँ जब आपकी evaluations ऊँची कीमत को सही ठहराएँ। QSP endpoint केवल text लेता है और फ़िलहाल zero data retention के साथ 262K-token context की गारंटी देता है।

Quickstart (curl)

curl https://api.quicksilverpro.io/v1/chat/completions \
  -H "Authorization: Bearer $QSP_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nemotron-3.5-lightning",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

OpenAI-संगत। base_url बदलकर एक लाइन में migration।

FAQ

$0.066 प्रति million input tokens, $0.176 प्रति million output tokens, और $0.033 प्रति million cached-input tokens — सबसे कम stable paid market rate से लगभग 10% ऊपर।

मॉडल का architecture 1M tokens तक support कर सकता है, जबकि QSP फ़िलहाल production में 262K tokens की गारंटी देता है। हम सैद्धांतिक अधिकतम का प्रचार करने के बजाय वही सीमा प्रकाशित करते हैं जिसे हम test और enforce करते हैं।

हाँ। यह streaming, function calling, और JSON Schema structured output को support करता है। QSP इसे reasoning बंद रखकर serve करता है, इसलिए responses सीधे आते हैं और उनमें कोई अलग reasoning trace या छिपे हुए reasoning tokens नहीं होते।

base_url=https://api.quicksilverpro.io/v1 सेट करें, अपनी QSP key इस्तेमाल करें, और model="nemotron-3.5-lightning" सेट करें। Chat Completions और Responses API दोनों के clients एक ही public model ID इस्तेमाल करते हैं।

Nemotron 3.5 Lightning को double credits के साथ आज़माएँ — $50 तक bonus credits

API Key पाएँ