Runpod
LLM APIs Verified May 2026
Runpod deal: Sign up free and pay only for what you use — no commitments
GPU cloud for AI builders — H100s from $2.89/hr, per-second billing, serverless and persistent Pods across 30+ regions.
- H100s for under $3.30/hr
- Per-second billing on Serverless
- Genuine on-demand availability
- Two tiers of trust
How Runpod scored 80/100
6 weighted criteria, each scored out of 10 and published with its reasoning. Featured placements never move a score.
Deal Strength
8.0 /10There is no exclusive coupon here; the real value is the pricing model itself. Sign-up is free with zero subscription fee, and you pay only for compute you actually use — per-second on Serverless, per-minute on Pods — so you can test a GPU workload for the cost of a few minutes.
Value for Money
9.0 /10Runpod is priced well below the hyperscalers, often up to 80% cheaper for equivalent GPUs. H100 PCIe starts around $2.89/hr and A100 80GB near $1.39/hr, versus far pricier on-demand instances on AWS or GCP. For on-demand AI training and inference, the value is excellent.
Capability
8.0 /10The platform is broad and genuinely capable: persistent Pods with SSH and Jupyter, auto-scaling Serverless endpoints with sub-200ms cold starts, and multi-GPU Clusters scaling past 200 GPUs over InfiniBand. With 30+ GPU SKUs spanning B200, H200, H100 and A100, it covers most LLM, diffusion and fine-tuning workloads with few gaps.
Time to Value
7.0 /10Provisioning is fast — popular GPUs are usually bookable in minutes and preconfigured templates get common frameworks running quickly. Standing up a custom container or a bespoke Serverless handler still takes some setup work, so most teams are productive within hours rather than instantly.
Trust & Reliability
6.0 /10Runpod is backed by real credentials: SOC 2 Type II, ISO 27001 and HIPAA compliance, a 99.99% uptime SLA on Secure Cloud, and adoption by over a million developers plus names like Hugging Face, Perplexity and Replit. The caveat is that lower-cost Community Cloud capacity runs on distributed third-party hosts, so individual-host reliability is less consistent than dedicated Secure Cloud.
Flexibility & Exit
10.0 /10This is as lock-in-free as GPU cloud gets. Billing is pure usage-based with no contracts or minimums, Serverless scales to zero when idle, and you can start, stop or tear down compute whenever you like — giving you full control over spend with nothing to cancel.
Sign up free and pay only for what you use — no commitments
Runpod has no subscription fee — you pay only for compute time used, with per-second billing on Serverless and per-minute billing on Pods. Storage is $0.05–$0.14/GB/mo. Promotional credits for new accounts are awarded at Runpod’s discretion. Verify current pricing at signup.
Affiliate link — same price for you, and it never moves the score.
- H100s for under $3.30/hr
- Per-second billing on Serverless
- Genuine on-demand availability
- Two tiers of trust
About Runpod
Quick answer
Runpod is a usage-based GPU cloud for AI builders. There is no subscription — you rent GPUs by the minute as persistent Pods (containers with SSH/Jupyter) or run auto-scaling Serverless endpoints billed by the second. Headline pricing in 2026: H100 SXM at $3.29/hr, A100 80GB at $1.49/hr, L40 at $0.99/hr, and budget GPUs (RTX A5000, L4) from $0.27–$0.39/hr; storage is $0.05–$0.14/GB/mo. That undercuts an AWS p5 on-demand instance (~$98/hr for 8×H100) by more than half while keeping H100/H200/B200 capacity bookable in minutes. Best for teams who want flagship compute without a one-year reservation. Sign up free through the partner link and pay only for what you use.
The real question: what does a GPU-hour actually cost?
Almost every GPU-cloud comparison gets derailed by branding. The number that matters to an AI team is brutally simple: how many dollars does one hour of an H100 cost, and can I actually get one today? On the hyperscalers the honest answer in 2026 is "expensive and usually reserved." An AWS p5 instance — eight H100s — lists near $98/hr on-demand, which works out to roughly $12.25 per H100-hour before you factor in the capacity reservation you almost certainly need to get one at all. Runpod's pitch is that it collapses that to one rentable H100 SXM at $3.29/hr, on demand, with per-minute billing and no commitment. That is the whole story, and it is why Runpod earns a place on most AI teams' shortlist.
The second-order point is billing granularity. A reserved hyperscaler instance bills whether you use it or not. Runpod Pods bill per minute and Serverless bills per second of actual request processing — which is the only honest way to price bursty inference. If your traffic is spiky, scale-to-zero Serverless means you stop paying the moment the queue empties.
Put the two together and the economics get interesting. A team fine-tuning a 7B model overnight on a single A100 80GB pays roughly $1.49/hr — call it about $12 for an eight-hour run — and then shuts the Pod down. The same eight hours on a reserved hyperscaler instance is billed against a commitment you signed weeks earlier, whether the GPU was busy or idle. For research and experimentation, where you spin compute up and down dozens of times a week, that difference compounds into the single largest line item you control. The discipline Runpod rewards is simple: provision when you need it, stop when you don't, and let per-minute and per-second billing do the rest.
There is also a supply story behind the price. H100s have been scarce on the hyperscalers for two years, which is exactly why getting one on AWS often means a capacity reservation and a wait. Runpod's distributed model — a mix of Secure Cloud datacenters and a Community Cloud of peer providers — means flagship GPUs including B200, H200, and H100 are generally bookable in minutes. For a team that needs to start training today, availability is as much a feature as price.
Runpod pricing in 2026 — the full GPU-hour table
| Tier | GPUs | Price (on-demand) | Billing |
|---|---|---|---|
| Budget GPUs (Pods) | L4, RTX A5000, A40, L40 | $0.27–$0.99/hr | Per-minute |
| Pro GPUs (Pods) | A100 80GB, RTX 6000 Ada, RTX Pro 6000 | $1.39–$2.09/hr | Per-minute |
| Flagship GPUs (Pods) | H100, H200, B200 | $2.89–$5.89/hr | Per-minute |
| Serverless | scales to zero, per-request | $0.69–$8.64/hr | Per-second |
| Storage | network volumes / S3-compatible | $0.05–$0.14/GB/mo | Monthly |
There is no platform fee layered on top — the GPU-hour and storage rate is the bill. Promotional credits for new accounts are awarded at Runpod's discretion; verify the current rate at signup, since flagship GPU pricing moves as supply changes.
Runpod vs Lambda Labs, Vast.ai, and AWS
The GPU-cloud market splits into three archetypes, and Runpod deliberately sits between them. The comparison that matters is the effective per-H100-hour cost paired with whether you can actually get the hardware.
| Platform | H100 on-demand | Real serverless? | Availability | Best for |
|---|---|---|---|---|
| Runpod | ~$3.29/hr | Yes (per-second) | Bookable in minutes | Flexible flagship compute, bursty inference |
| Lambda Labs | Comparable on-demand | No | Limited regions, frequent waitlists | Sustained training in a single region |
| Vast.ai | Cheapest (marketplace) | No | Highly variable, peer-sourced | Cost-first dev / non-critical batch |
| AWS p5 | ~$12.25/H100-hr | No (SageMaker only) | Capacity reservation usually required | Teams already locked into AWS |
The takeaway: Vast.ai will sometimes beat Runpod on raw price, but reliability is a coin-flip; Lambda matches Runpod on-demand but has no serverless tier and tighter capacity; AWS is the most expensive and the hardest to provision. Runpod's edge is the combination — marketplace-adjacent pricing, hyperscaler-grade availability, and a genuine serverless option none of the others ship.
Runpod at a glance — the spec sheet
| Billing model | Usage-based — per-minute Pods, per-second Serverless, no subscription |
|---|---|
| Flagship GPUs | H100, H200, B200 (up to 180GB VRAM on B200) |
| Trust tiers | Secure Cloud (tier-3+ DCs, SLA) and Community Cloud (peer-sourced, cheaper) |
| Regions | 30+ worldwide |
| Multi-GPU | Clusters up to 64 GPUs with InfiniBand |
| Storage | Persistent network volumes ($0.07/GB/mo) + S3-compatible |
| Deploy options | 50+ templates (PyTorch, vLLM, Ollama, ComfyUI, A1111), BYOC Docker, Flash (Python-only) |
| Automation | CLI + REST API for CI/CD; public model endpoints |
What you actually get
Pods (persistent containers)
GPU containers with SSH and Jupyter, billed per minute. The right tool for fine-tuning, notebooks, and batch training where you want a stable environment that keeps its state.
Serverless endpoints
Auto-scaling worker pools billed per second of request processing. Scales from zero to N workers, so production inference only costs money while it is doing work.
Runpod Flash
Ship a Python file and get a serverless GPU endpoint — no Dockerfile, no image build. It removes the single biggest friction point of every other serverless-GPU platform.
Two trust tiers
Secure Cloud runs in tier-3+ datacenters with SLA-backed uptime for production; Community Cloud is peer-sourced and cheaper for dev and experiments. You choose the risk/price trade-off per workload.
Real clusters
Multi-GPU clusters up to 64 GPUs over InfiniBand, bookable without an enterprise sales call — rare at this price point.
Templates + BYOC
50+ one-click templates (vLLM, Ollama, ComfyUI, A1111) or bring your own container. CLI and REST API wire it all into your CI/CD.
How to get an H100 running on Runpod in five steps
Who should use Runpod — and who shouldn't
✓ Use Runpod if you
- Need H100/H200/B200 capacity without a one-year reservation.
- Run bursty inference and want to stop paying when idle.
- Want to fine-tune on an A100 for $1.49/hr and shut it down.
- Are comfortable bringing your own MLOps stack (W&B, MLflow).
- Want to ship a serverless GPU endpoint without writing a Dockerfile.
✗ Skip it if you
- Need a turnkey, fully managed MLOps platform with built-in tracking.
- Require five-nines guaranteed uptime on the cheapest Community tier.
- Run latency-critical chat UIs and can't tolerate any cold-start lag.
- Are already deeply committed to a hyperscaler's reserved capacity.
Where Runpod earns its keep
Four workloads cover the bulk of what teams actually run on it. LLM fine-tuning on a budget is the obvious one — rent an A100 for $1.49/hr instead of $3+/hr elsewhere, fine-tune Llama or Mistral in a few hours, and shut it down before the next billing minute ticks over. Production inference at scale is the Serverless story: deploy a vLLM or TGI endpoint that scales from zero to N workers and only bills while it's serving requests, which is the right shape for any product with uneven traffic. AI agents that need persistent GPU state live on Pods, where a persistent network volume keeps context and checkpoints across restarts so a long-running, multi-step pipeline doesn't lose its place. And image and video generation services spin up A1111, ComfyUI, or a video-model template and serve generations to users at marketplace prices — a category where GPU cost directly sets your margin.
The common thread is that none of these workloads wants a yearly reservation. They want flagship hardware on tap, billed by the minute or second, with the freedom to switch GPU class as the model or the traffic changes. That is precisely the gap Runpod fills between the cheap-but-flaky marketplaces and the reliable-but-expensive hyperscalers.
What's included
- Pods: persistent GPU containers with SSH/Jupyter (per-minute billing)
- Serverless: auto-scaling endpoints with per-second billing
- Runpod Flash: serverless GPU with just Python — no Docker required
- Thousands of GPUs across 30+ regions worldwide
- Community Cloud (peer-to-peer) and Secure Cloud (tier-3+ DCs)
- Multi-GPU clusters up to 64 GPUs with InfiniBand
- Public API endpoints for pre-deployed models (LLMs, image, video)
- Persistent network volumes ($0.07/GB/mo) and S3-compatible storage
- One-click templates for PyTorch, vLLM, Ollama, ComfyUI, A1111
- CLI + REST API for full automation and CI/CD
- BYOC (bring-your-own-container) Docker support
- SLA-backed uptime on reserved clusters
Runpod pricing
Verified May 2026. Vendor's published rates at the time we checked — always confirm at checkout.
| Plan | Price | Term | What you get |
|---|---|---|---|
| Budget GPUs (Pods) | $0.27–$0.99/hr | L4, A5000, A40, L40 — per-minute | Great for inference and fine-tuning · 24GB–48GB VRAM tier · Spin up in seconds · Per-minute billing |
| Pro GPUs (Pods) | $1.39–$2.09/hr | A100, RTX 6000 Ada, RTX Pro 6000 | Training-grade memory (48–96GB) · Multi-GPU clusters available · Persistent network volumes · Per-minute billing |
| Flagship GPUs (Pods) | $2.89–$5.89/hr | H100, H200, B200 | Highest training throughput · Up to 180GB VRAM (B200) · Reserved clusters up to 64 GPUs · SLA-backed uptime on reservations |
| Serverless | $0.69–$8.64/hr | per-second billing, scales to zero | Cold starts in seconds (Flash) · Auto-scaling worker pool · Pay only when handling requests · Ideal for production inference |
How to claim it
4 steps. The last one is the part most people skip.
- 1
Open Runpod through the link on this page
It carries our referral tag. The price you pay is identical either way, and it never changes the score on this page.
- 2
Pick the plan that matches your usage
This offer applies automatically through the link — there is no code to enter.
- 3
Confirm the discount before you pay
The order summary should show the reduced amount. If it does not, stop and tell us — we re-test listings that stop working.
- 4
Check what happens at renewal
Runpod has no subscription fee — you pay only for compute time used, with per-second billing on Serverless and per-minute billing on Pods. Storage is $0.05–$0.14/GB/mo. Promotional credits for new accounts are awarded at Runpod’s discretion. Verify current pricing at signup.
Where Runpod wins and loses
What works
- H100s for under $3.30/hr H100 SXM at $3.29/hr undercuts AWS p5 and most enterprise clouds by 50%+ while offering true on-demand access — no reservation required.
- Per-second billing on Serverless You only pay while a request is being processed, which is the only honest way to bill production inference workloads with bursty traffic.
- Genuine on-demand availability B200, H200, and H100 capacity is generally bookable in minutes — a real differentiator while H100s remain scarce on hyperscalers.
- Two tiers of trust Secure Cloud for production (tier-3+ datacenters, SLA) and Community Cloud for dev/experiments (cheaper, peer-sourced) — you pick the risk/price trade-off.
- No Docker required (Flash) Runpod Flash lets you ship a Python file and get a serverless GPU endpoint — removes the biggest friction point of every other serverless GPU platform.
- Real multi-GPU and clusters Up to 64-GPU clusters with InfiniBand for training, available without an enterprise sales call.
What doesn't
- Community Cloud has variable reliability Cheaper GPUs in Community Cloud are sourced from independent providers — uptime is generally good but not guaranteed for production workloads. Use Secure Cloud when it matters.
- Cold-start latency on Serverless Even with Flash, the first request to a cold worker can take several seconds — fine for batch, painful for latency-critical chat UIs without min-worker provisioning.
- No built-in MLOps stack Runpod is raw compute. If you want experiment tracking, model registry, or pipelines, you BYO (Weights & Biases, MLflow, etc.).
The bottom line
Runpod is one of the best-value ways to rent serious AI compute: H100s start around $2.89/hr, billing is pure pay-as-you-go with no monthly fee, and you can spin up Pods or scale-to-zero Serverless endpoints in minutes. Reliability on lower-cost community hosts can vary, but for training, fine-tuning and GPU inference it is a clear Strong Buy.
Runpod is a GPU cloud built for AI workloads. You can rent GPUs by the minute as persistent Pods (containers with SSH/Jupyter) or run auto-scaling Serverless endpoints billed per second. It targets developers, researchers, and AI companies who need H100/H200/B200-class compute without committing to a hyperscaler.
It is fully usage-based with no monthly fee. As of 2026: H100 SXM is $3.29/hr, A100 80GB is $1.49/hr, L40 is $0.99/hr, and budget GPUs (RTX A5000, L4) start at $0.27–$0.39/hr. Serverless adds a per-second tier from $0.69/hr to $8.64/hr depending on GPU. Storage is $0.05–$0.14/GB/mo. Verify current pricing at signup.
Lambda Labs has similar on-demand H100s but limited regions and no real serverless. Vast.ai is the cheapest peer-to-peer marketplace but reliability is highly variable. AWS p5 instances are $98/hr on-demand and require capacity reservations. Runpod sits in the sweet spot: hyperscaler-grade availability with marketplace-grade pricing, plus a genuine serverless tier.
Use Pods for interactive work (fine-tuning, notebooks, batch training) where you want a persistent environment. Use Serverless for production inference with variable traffic — it scales to zero when idle and you only pay per request.
Secure Cloud Pods run in tier-3+ datacenters with SLA-backed uptime — appropriate for production. Community Cloud uses peer-sourced infrastructure and is recommended for development, experimentation, and non-critical batch jobs.
Yes — Runpod supports bring-your-own-container (BYOC) Docker images on both Pods and Serverless. You can also start from 50+ pre-built templates (PyTorch, vLLM, Ollama, ComfyUI, A1111, etc.) to skip the image-build step.