Llama 3.1 Nemotron Instruct 70B API Pricing
NVIDIA · released 2024-10-15
Pricing
0.120¢ / 1K
0.120¢ / 1K
0.120¢ / 1K
What it actually costs
| Workload | Tokens | Cost |
|---|---|---|
| Short chat reply | 500 in / 300 out | 0.096¢ |
| RAG query | 4,000 in / 500 out | 0.54¢ |
| Long document summary | 50,000 in / 1,000 out | $0.0612 |
| Code review of a 2k-line file | 30,000 in / 2,000 out | $0.0384 |
| 1,000 chat replies/day for 30 days | 30,000 × (500 in / 300 out) | $28.80 |
Cost = (input tokens ÷ 1,000,000 × $1.20) + (output tokens ÷ 1,000,000 × $1.20).
Benchmarks
#388 of 533
#210 of 249
Similar intelligence, lower price
| Model | Intelligence | Blended $/1M |
|---|---|---|
| Llama 3.1 Instruct 8B | 6.9 | $0.028 |
| Gemma 4 E4B (Reasoning) | 8.9 | $0.04 |
| Sarvam 30B (high) | 6.6 | $0.047 |
| Granite 4.2 3B | 9.1 | $0.052 |
| Nova Micro | 5.9 | $0.061 |
Leaderboard neighbours
| Model | Intelligence | Blended $/1M |
|---|---|---|
| Solar Pro 2 (Non-reasoning) | 7.0 | — |
| Grok Beta | 6.9 | — |
| Llama 3.1 Instruct 8B | 6.9 | $0.028 |
| Llama 3.1 Nemotron Instruct 70B | 6.9 | $1.20 |
| Qwen2.5 Instruct 32B | 6.9 | — |
| Qwen3.5 2B (Reasoning) | 6.9 | — |
| Mistral Large 2 (Jul '24) | 6.8 | $3.00 |
More from NVIDIA
| Model | Intelligence | Blended $/1M |
|---|---|---|
| Nemotron 3 Ultra 550B A55B (Reasoning) | 22.9 | $1.05 |
| Nemotron 3.5 Lightning | 12.9 | $0.095 |
| Nemotron 3 Super 120B A12B (Reasoning) | 12.8 | $0.45 |
| Nemotron Cascade 2 30B A3B | 11.7 | — |
| Nemotron 3 Nano Omni 30B A3B Reasoning | 10.3 | $0.42 |
| Llama Nemotron Super 49B v1.5 (Reasoning) | 9.0 | $0.40 |
| Llama 3.3 Nemotron Super 49B v1 (Reasoning) | 8.9 | — |
| NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) | 8.9 | $0.088 |
Count tokens & estimate cost for Llama 3.1 Nemotron Instruct 70B See Llama 3.1 Nemotron Instruct 70B on the leaderboard
Source: Artificial Analysis · last synced 2026-09-22
FAQ
Who develops Llama 3.1 Nemotron Instruct 70B?
Llama 3.1 Nemotron Instruct 70B is developed by NVIDIA, released 2024-10-15.
How much does Llama 3.1 Nemotron Instruct 70B cost per million tokens?
Llama 3.1 Nemotron Instruct 70B costs $1.20 per 1M input tokens and $1.20 per 1M output tokens ($1.20 blended), per Artificial Analysis's official pricing data.
How does Llama 3.1 Nemotron Instruct 70B rank on benchmarks?
Llama 3.1 Nemotron Instruct 70B scores 6.9 on the Artificial Analysis Intelligence Index, ranking #388 of 533 models we track.
Is there a cheaper model with similar intelligence to Llama 3.1 Nemotron Instruct 70B?
Yes — Llama 3.1 Instruct 8B scores similar intelligence (6.9) at $0.028/1M blended, versus $1.20/1M for Llama 3.1 Nemotron Instruct 70B.
Where does this data come from?
Pricing and benchmark data for Llama 3.1 Nemotron Instruct 70B comes from Artificial Analysis (artificialanalysis.ai), synced daily. Last synced 2026-09-22.