Search for "rent GPU for AI" and you get a pile of pages that quote a different price for the same card. The reason is simple: an hour of GPU time is priced by VRAM, by cloud tier and by how long you hold the machine, and a good share of the bill hides in disks and idle time. This guide puts numbers from the providers' own price pages side by side and explains how to read them. The full list of GPU platforms is in our GPU cloud ranking.
A note on method. On 10 October 2026 we opened the price pages of RunPod, DigitalOcean and Cherry Servers and copied the rates as shown. Vultr's price page returned an access error when we checked, so there are no Vultr numbers here. We did not benchmark the cards: our /tests page measures the TTFB of hosters' public websites, which says nothing about GPU speed.
GPU rates change several times a year. Every figure below is a snapshot from the provider's site, October 2026. Open the live price page before you top up.
Hourly cloud or a dedicated GPU server
With hourly cloud you start a machine, pay per second or per hour, and delete it when the job ends. A dedicated GPU server is a whole physical machine rented for a month or longer, and at some providers by the hour too. For most jobs the first is cheaper, and the arithmetic fits on one line: divide the monthly price by the hourly price and you get the number of hours after which the monthly plan starts to win.
Cherry Servers lists an A100 80 GB from $2.362/h and from $1,657.68/month. Divide: about 702 hours, or 96% of a 730-hour month. The monthly discount is nearly invisible unless the card runs without a break. RunPod asks $1.79/h for an A100, which is about $1,307 for 730 hours. On these listed rates the cloud is cheaper even if you never switch it off.
- Hourly cloud fits when you train or generate in bursts, test a model, or your load differs from week to week.
- A dedicated server makes sense when you need a stable machine with a large included traffic allowance (Cherry Servers lists 30–100 TB of free egress), a fixed IP and a disk nobody wipes. It also fits an inference service that runs around the clock for months.
- If the model fits into the graphics card of your own computer, renting only buys speed. Check your own hardware first.
VRAM decides almost everything
A model either fits into video memory or it does not. There is no "slower but works" on a single card. RunPod's documentation gives a working rule of about 2 GB of VRAM per billion parameters at 16-bit precision. That is roughly 14 GB for a 7B model, 26 GB for 13B and 140 GB for 70B, which means several GPUs. Quantisation lowers the requirement: per the same docs, a 4-bit 70B model needs about 35 GB.
Images have a lower bar. The same documentation puts SDXL at about 8 GB minimum and recommends 16–24 GB for SDXL and Flux, which leaves room for bigger batches and LoRA training. For fine-tuning LLMs it suggests 40–80 GB (A100 or H100 class), and memory bandwidth matters most there.
Leave a margin. Context length and batch size also use memory, so a model that "just fits" will crash on the first long prompt.
Which card for which job
| GPU (VRAM) | RunPod, $/h | Fits |
|---|---|---|
| RTX 3090 (24 GB) | 0.50 | SDXL and Flux, 7B LLMs (≈14 GB), quantised models up to 13B |
| RTX 4090 (24 GB) | 0.89 | the same jobs, faster |
| L40S (48 GB) | 1.09 | larger quantised LLMs, big image batches, LoRA training |
| A100 (80 GB) | 1.79 | fine-tuning 7B–13B, 70B in 4-bit (≈35 GB) |
| H100 SXM (80 GB) | 3.99 | training, high-throughput inference |
| H200 (141 GB) | 5.29 | 70B at 16-bit is a squeeze: weights alone are ≈140 GB, little room for long context |
A practical rule: take the cheapest card whose VRAM clears your model with room to spare, and move up only when you hit out-of-memory errors or a job that takes too long. A 24 GB card at $0.50–0.89 an hour answers most questions from hobbyists and small teams.
What an hour costs: providers compared
| Provider | Cards and prices | Billing | Traffic | Best for |
|---|---|---|---|---|
| RunPod | RTX 4090 24 GB — $0.89/h; A100 80 GB — $1.79/h; H100 SXM — $3.99/h | per second; about an hour of credit needed to deploy | no ingress or egress fees | experiments, fine-tuning, SD and Flux, Serverless inference |
| DigitalOcean | RTX 4000 Ada 20 GB — $0.76/h; L40S 48 GB — $1.57/h; H100 80 GB — $4.41/h; MI300X 192 GB — $2.59/h | per second, 5-minute minimum; a stopped Droplet is still billed | 10,000–15,000 GiB included, overage price not stated | teams whose app already runs on DigitalOcean |
| Cherry Servers | A2 16 GB — from $0.239/h; A40 48 GB — from $0.801/h; A100 80 GB — from $2.362/h | hourly, monthly, yearly | 30–100 TB of free egress | a dedicated GPU server; the page lists Lithuania, some models are pre-order or waiting list |
| Vultr | a Cloud GPU line exists, the price page did not open, so no figures | — | — | compare on the live page; $300 credit for new users for 30 days |
Two observations. RunPod is cheaper on cards of the same class: the L40S is $1.09/h there against $1.57 at DigitalOcean, and the H100 is $3.99 against $4.41. But it is a GPU-only platform, so your database and web app have to live somewhere else. DigitalOcean costs more per hour, yet a GPU Droplet sits in the same account, network and API as the rest of your infrastructure. Cherry Servers is the only one of the three that hands you a whole physical machine with a big traffic allowance. Its A40 and A100 are marked as pre-order, so check stock before you pay.
Community, Secure and spot: where the discount comes from
RunPod sells two clouds. Secure Cloud runs in tier 3/4 data centres with redundancy and suits production and sensitive data. Community Cloud is machines owned by third parties: cheaper, but reliability varies. RunPod says it no longer accepts new Community hosts, though existing capacity stays. The price page has a Community/Secure toggle and the list we copied does not say which one it shows. So pick the tier explicitly at deploy time and read the rate on the deploy screen.
The other source of discounts is spot, meaning machines that can be taken back. DigitalOcean lists spot GPU Droplets only for the newest data-centre cards (for example the B300 at $6.17/h). It aims to give at least two hours' notice but may reclaim a machine with less or none, and it locks the price at creation. Spot suits jobs that save checkpoints and resume. It is a poor fit for a demo in front of a customer.
- Community and spot: experiments, batch generation, jobs with checkpoints.
- Secure and on-demand: client work, private data, anything with a deadline.
Hidden costs: disks, traffic and idle time
- Disk. At RunPod a container disk costs $0.10/GB a month and is erased when the pod stops. A volume disk costs $0.10/GB a month while running and $0.20/GB while stopped. A network volume is $0.07/GB a month under 1 TB and does not depend on a pod. For 200 GB of models that is $20 running, $40 stopped, or $14 on a network volume.
- Traffic. RunPod's docs state no fees for data ingress or egress. DigitalOcean includes 10,000–15,000 GiB in a GPU Droplet plan but gives no overage price. Cherry Servers lists 30–100 TB of free egress.
- Idle time. DigitalOcean keeps billing a powered-off Droplet because the disk, CPU, RAM and IP stay reserved; only destroying it stops the meter. An RTX 4090 at $0.89/h forgotten over a weekend (48 hours) costs about $43.
- Balance. At RunPod you need roughly an hour of credit to deploy. When the balance reaches $0, pods stop, and those without a network volume are terminated with their data, which cannot be recovered.
Top up with a margin and keep weights and checkpoints on a network volume or in external storage. A container disk is temporary: stopping the pod takes its contents with it.
Checklist before you pay
- Write down the model size and precision, work out the VRAM (2 GB per billion parameters at 16-bit) and add a 20–30% margin.
- Take the smallest card that fits and give it a test hour.
- Decide the tier: Secure for client work, Community for experiments.
- Put weights and checkpoints on a network volume or in object storage, not on the container disk.
- Check the disk price while stopped and the egress terms.
- Set a reminder or an auto-shutdown, and a balance alert.
- Top up the minimum first. Through our link RunPod gives new users a $5–500 credit bonus.
- For a service that must stay up, test a cold start before launch: a new pod, loading the weights, serving.
Which provider for whom
- RunPod: the cheapest hour of an RTX 4090, A100 and H100 among the three compared, ready templates (PyTorch, ComfyUI, Ollama) and Serverless for inference. The downside is GPU only, and you must tell Community from Secure yourself.
- DigitalOcean: if your app already lives there and you want a GPU Droplet in the same account. It costs more than RunPod for the L40S and H100.
- Cherry Servers: a dedicated GPU server with a big traffic allowance and a monthly contract. Check which models are actually in stock.
- Vultr: has a Cloud GPU line and a $300 credit for new users, but we could not read its price page, so compare on the live page.
Count the cost of the finished result, not of the hour. A card that is $0.40 cheaper but crashes your run at 90% costs more than the pricier one that finishes.
Igor Jazov, Tophosting editor
Bottom line
For experiments, image generation and fine-tuning, start with hourly cloud on a 24–48 GB card and count hours, not months. Take a dedicated server when the GPU runs almost without pauses or you need the traffic and a stable address. If your project does not need a GPU, an ordinary VPS is enough. The full list of GPU clouds with new-user bonuses is in our ranking.
FAQ
How much does it cost to rent a GPU per hour?
Per the providers' sites, October 2026: RunPod lists the RTX 3090 at $0.50/h, RTX 4090 at $0.89/h, A100 80 GB at $1.79/h and H100 SXM at $3.99/h; DigitalOcean has the L40S at $1.57/h and H100 at $4.41/h; Cherry Servers lists the A100 from $2.362/h. Rates change, so open the live page before paying.
What is the cheapest GPU cloud for AI?
Among the three price pages we compared, RunPod shows the lowest rates on cards of the same class: the L40S is $1.09/h there against $1.57 at DigitalOcean. The page does not say whether the list is Community or Secure, so confirm the rate on the deploy screen.
How much VRAM do I need for Stable Diffusion or Flux?
RunPod's documentation puts SDXL at about 8 GB minimum and recommends 16–24 GB for SDXL and Flux. A 24 GB card (RTX 3090, RTX 4090) is the safe choice, with room for batches and LoRA training.
Hourly GPU cloud or a dedicated GPU server?
Divide the monthly price by the hourly one. If you will use fewer hours than that, hourly rental is cheaper. For an A100 at Cherry Servers it is about 702 of 730 hours, so the monthly plan wins only with near round-the-clock load.
Can I rent a GPU with per-second billing?
Yes. RunPod bills Pods per second. DigitalOcean also bills per second, with a 5-minute minimum, and a powered-off Droplet keeps being billed until you destroy it.
Is Community Cloud safe for private data?
RunPod's documentation recommends Secure Cloud for production and sensitive data, since it runs in data centres with redundancy. Community Cloud is third-party machines with variable reliability, so keep it for experiments.
