GPU Cost Calculator: What AI Inference and Training Really Costs on Every Major Platform
- GPU compute costs vary 10–30× depending on provider, GPU model, and reservation type — an A100 on AWS on-demand is $3.00–$3.50/hour; the same GPU reserved for 1 year drops to $1.60–$1.90/hour; spot/preemptible instances go as low as $0.80–$1.20/hour.
- For AI inference (running a model to generate outputs), GPU cost is rarely the right metric — token cost from managed APIs (OpenAI, Anthropic, Google) is almost always cheaper than self-hosting unless you're at very high volumes.
- The break-even volume where self-hosted GPU beats managed API typically requires 50–200 million tokens per day — equivalent to a large-scale production application, not a startup or side project.
- Indian cloud providers (Jio Cloud, BSNL, Sify) offer GPU instances at ₹400–₹1,200/hour — but GPU availability, reliability, and support lag significantly behind AWS, Azure, and GCP.
Use our free GPU Cost Calculator to compute hourly, daily, and monthly GPU rental cost across providers — and find whether self-hosting or managed API is cheaper for your volume.
GPU Pricing by Provider and Model (2025)
NVIDIA H100 (Flagship — Best for LLM Training and Inference)
Provider On-Demand $/hr 1-Year Reserved Spot/Preemptible
AWS (p5.48xlarge, 8× H100) $98.32/hr (÷8 = $12.29/GPU) ~$7.80/GPU/hr ~$4.50/GPU/hr
Google Cloud (a3-highgpu) $12.29/GPU/hr ~$7.50/GPU/hr ~$3.80/GPU/hr
Azure (NDv5) $11.50/GPU/hr ~$7.20/GPU/hr ~$4.00/GPU/hr
Lambda Labs $3.29/GPU/hr N/A N/A
CoreWeave $4.76/GPU/hr ~$2.80/GPU/hr N/A
Vast.ai (marketplace) $2.50–$4.00/GPU/hr N/A $1.50–$2.50
NVIDIA A100 80GB (Most Common for LLM Work)
Provider On-Demand $/hr 1-Year Reserved Spot/Preemptible
AWS (p4de.24xlarge, per GPU) $3.47/GPU/hr $1.89/GPU/hr $1.04/GPU/hr
Google Cloud (a2-ultragpu) $3.14/GPU/hr $1.73/GPU/hr $0.94/GPU/hr
Azure (NDasrv4) $3.40/GPU/hr $1.85/GPU/hr $1.02/GPU/hr
Lambda Labs $2.49/GPU/hr $1.90/GPU/hr N/A
CoreWeave $2.21/GPU/hr $1.65/GPU/hr N/A
RunPod $1.64–$2.49/GPU/hr N/A $0.79/GPU/hr
Vast.ai $1.20–$2.00/GPU/hr N/A $0.70–$1.30
NVIDIA A10G / A10 (Mid-Range — Good for Inference)
Provider On-Demand $/hr
AWS (g5.xlarge, 1× A10G) $1.006/hr
Google Cloud (g2-standard, 1× L4) $0.70/hr
Azure (Standard_NV6ads, 1/6 A10) $0.90/hr
Lambda Labs $0.75/hr
RunPod $0.44–$0.74/hr
NVIDIA T4 (Entry-Level — Small Models, Dev Work)
Provider On-Demand $/hr
AWS (g4dn.xlarge) $0.526/hr
Google Cloud (n1 + T4) $0.35/hr
Azure (Standard_NC4as) $0.526/hr
RunPod $0.14–$0.24/hr
GPU Cost in INR (For Indian Developers)
| Provider | On-Demand $/hr | 1-Year Reserved | Spot/Preemptible |
|---|---|---|---|
| AWS (p5.48xlarge, 8× H100) | $98.32/hr (÷8 = $12.29/GPU) | ~$7.80/GPU/hr | ~$4.50/GPU/hr |
| Google Cloud (a3-highgpu) | $12.29/GPU/hr | ~$7.50/GPU/hr | ~$3.80/GPU/hr |
| Azure (NDv5) | $11.50/GPU/hr | ~$7.20/GPU/hr | ~$4.00/GPU/hr |
| Lambda Labs | $3.29/GPU/hr | N/A | N/A |
| CoreWeave | $4.76/GPU/hr | ~$2.80/GPU/hr | N/A |
| Vast.ai (marketplace) | $2.50–$4.00/GPU/hr | N/A | $1.50–$2.50 |
| Provider | On-Demand $/hr | 1-Year Reserved | Spot/Preemptible |
|---|---|---|---|
| AWS (p4de.24xlarge, per GPU) | $3.47/GPU/hr | $1.89/GPU/hr | $1.04/GPU/hr |
| Google Cloud (a2-ultragpu) | $3.14/GPU/hr | $1.73/GPU/hr | $0.94/GPU/hr |
| Azure (NDasrv4) | $3.40/GPU/hr | $1.85/GPU/hr | $1.02/GPU/hr |
| Lambda Labs | $2.49/GPU/hr | $1.90/GPU/hr | N/A |
| CoreWeave | $2.21/GPU/hr | $1.65/GPU/hr | N/A |
| RunPod | $1.64–$2.49/GPU/hr | N/A | $0.79/GPU/hr |
| Vast.ai | $1.20–$2.00/GPU/hr | N/A | $0.70–$1.30 |
| Provider | On-Demand $/hr |
|---|---|
| AWS (g5.xlarge, 1× A10G) | $1.006/hr |
| Google Cloud (g2-standard, 1× L4) | $0.70/hr |
| Azure (Standard_NV6ads, 1/6 A10) | $0.90/hr |
| Lambda Labs | $0.75/hr |
| RunPod | $0.44–$0.74/hr |
| Provider | On-Demand $/hr |
|---|---|
| AWS (g4dn.xlarge) | $0.526/hr |
| Google Cloud (n1 + T4) | $0.35/hr |
| Azure (Standard_NC4as) | $0.526/hr |
| RunPod | $0.14–$0.24/hr |
At USD/INR of ₹83.50 (verify current rate):
| GPU | Provider | $/hr | ₹/hr | ₹/day | ₹/month |
|---|---|---|---|---|---|
| H100 | Lambda Labs | $3.29 | ₹274 | ₹6,580 | ₹1,97,400 |
| A100 80GB | RunPod | $1.64 | ₹137 | ₹3,282 | ₹98,460 |
| A100 80GB | Vast.ai | $1.20 | ₹100 | ₹2,400 | ₹72,000 |
| A10G | AWS | $1.006 | ₹84 | ₹2,009 | ₹60,270 |
| T4 | RunPod | $0.20 | ₹17 | ₹400 | ₹12,000 |
| T4 | Google Cloud | $0.35 | ₹29 | ₹700 | ₹21,000 |
Indian cloud alternatives:
- Jio Cloud GPU: ₹300–₹800/hr (limited availability, V100 class)
- E2E Networks GPU: ₹180–₹600/hr (A100 available, good uptime)
- Yotta Infrastructure: ₹400–₹1,000/hr (enterprise-focused)
GPU Cost Calculator: What Are You Computing?
For AI Training
Training Cost = GPU Hours Required × Cost per GPU Hour
Estimating GPU hours for training:
| Model Size | Dataset | GPUs Needed | Training Time | Cost (A100 @ $2/hr) |
|---|---|---|---|---|
| 7B params (LLaMA-style) | 1T tokens | 64 GPUs | ~21 days | $64,512 |
| 13B params | 1T tokens | 128 GPUs | ~21 days | $1,29,024 |
| 70B params | 2T tokens | 512 GPUs | ~90 days | $22,14,720 |
| Small fine-tune (7B) | 10K examples | 1–2 GPUs | 2–8 hours | $4–$32 |
| LoRA fine-tune (7B) | 50K examples | 1 GPU | 4–12 hours | $8–$24 |
Key insight: Fine-tuning (adapting an existing model) costs 1,000–10,000× less than training from scratch. For 99% of use cases, fine-tuning an open-source model is the right approach, not training from scratch.
For AI Inference (Serving)
Monthly Inference Cost = (Requests/day × Tokens/request × Cost per token) OR (GPU hours × GPU cost)
Self-hosted inference — what you can serve per GPU:
| GPU | Model | Throughput (tokens/sec) | Cost per 1M tokens |
|---|---|---|---|
| T4 (16GB) | Llama-3 8B (Q4) | ~200 t/s | ~$0.05 |
| A10G (24GB) | Llama-3 8B | ~450 t/s | ~$0.03 |
| A100 80GB | Llama-3 70B | ~80 t/s | ~$0.35 |
| H100 80GB | Llama-3 70B | ~180 t/s | ~$0.25 |
| 2× H100 | Llama-3 405B | ~30 t/s | ~$3.00 |
Self-Hosted GPU vs. Managed API: Break-Even Analysis
Scenario: Running Llama-3 70B equivalent quality
Managed API option:
- GPT-4o: $5/1M input tokens + $15/1M output tokens → avg ~$10/1M tokens
- Claude Sonnet: $3/1M input + $15/1M output → avg ~$9/1M tokens
Self-hosted option (A100, RunPod $1.64/hr):
- Throughput: ~80 tokens/sec
- 1M tokens takes: 1,000,000 ÷ 80 = 12,500 seconds = 3.47 hours
- Cost: 3.47 × $1.64 = $5.69 per 1M tokens
Break-even: At 1M+ tokens/day, self-hosting A100 is 40–75% cheaper than GPT-4o.
But add:
- Engineering time to set up and maintain inference server: $2,000–$10,000/month
- Downtime and reliability costs: significant for production apps
- Model updates and fine-tuning: ongoing cost
Realistic break-even for self-hosting: 30–100M tokens/day in production volume, with a dedicated ML engineer. Below this — use managed APIs.
GPU Selection Guide by Use Case
Use Case Recommended GPU Why
Development / prototyping T4 or A10G Cheap, sufficient for testing
Fine-tuning 7B model A10G or A100 40GB Fits full model, LoRA efficient
Fine-tuning 70B model 2–4× A100 80GB Needs multi-GPU for VRAM
Production inference (7–13B) A10G or L4 Best price/performance for inference
Production inference (70B) A100 80GB or H100 VRAM requirement
LLM training from scratch H100 cluster Maximum throughput
Image generation (SD/FLUX) A10G or 3090 Good VRAM/price balance
Video generation H100 or multi-A100 Extremely compute-intensive
Spot/Preemptible Instances: 60–70% Discount With Trade-offs
| Use Case | Recommended GPU | Why |
|---|---|---|
| Development / prototyping | T4 or A10G | Cheap, sufficient for testing |
| Fine-tuning 7B model | A10G or A100 40GB | Fits full model, LoRA efficient |
| Fine-tuning 70B model | 2–4× A100 80GB | Needs multi-GPU for VRAM |
| Production inference (7–13B) | A10G or L4 | Best price/performance for inference |
| Production inference (70B) | A100 80GB or H100 | VRAM requirement |
| LLM training from scratch | H100 cluster | Maximum throughput |
| Image generation (SD/FLUX) | A10G or 3090 | Good VRAM/price balance |
| Video generation | H100 or multi-A100 | Extremely compute-intensive |
Spot instances can be interrupted with 2-minute warning (AWS) or 30-second warning (GCP). This makes them:
Suitable for:
- Batch training jobs (with checkpointing every 10–15 minutes)
- Offline inference batch processing
- Data preprocessing and experimentation
Not suitable for:
- Real-time inference APIs
- Interactive applications
- Long training runs without robust checkpointing
Best practice: Use spot for training (checkpoint frequently), use on-demand or reserved for inference serving.
FAQ
Try the Free GPU Cost Calculator
Use ToolMira's calculator — no signup, no ads, works on mobile.
Open AI Cost Calculator →Disclaimer: This article is for educational purposes only and does not constitute financial, investment, or professional advice. Please consult a qualified professional before making any decisions based on this content.