Starting from
€0.57/hour
Experience exceptional performance across a range of demanding applications with our GPU Servers, from generative AI, LLM inference, and AI model training to 3D rendering and video processing.
On-demand NVIDIA L4, L40S, H100 and B200 GPUs available now!
Mian, Product Manager
GPU resources with the flexibility and easy management of our Cloud Servers.
No shared hardware – each GPU is always dedicated to a single server.
Get started quickly with drivers & tooling pre-installed.
No signup required
Instant access to our live demo. No signup. No commitment.
With UpCloud’s usage-based GPU billing, you only pay for active compute time. This allows for precise scaling of resources, eliminating costs for idle infrastructure. It’s ideal for teams seeking high-end GPUs without long-term commitments.
Experience the power of NVIDIA GPUs!
Cost-effective AI Inference, High-throughput Video Processing, and Edge AI deployments. Optimized for low-latency, energy-efficient operations.
The NVIDIA L4 Tensor Core GPU powered by the NVIDIA Ada Lovelace architecture delivers universal, energy-efficient acceleration for video, AI, visual computing, graphics, virtualization, and more.
Generative AI, Mid-to-Large Scale AI Model Training and Inference, Real-time 3D Rendering, Virtual Production, and High-Performance Computing (HPC) simulations.
The NVIDIA L40S GPU is the most powerful universal GPU for the cloud, delivering end-to-end acceleration for the next generation of AI-enabled applications.
The NVIDIA H100 GPU delivers exceptional performance, scalability, and security for every workload. H100 uses breakthrough innovations based on the NVIDIA Hopper™ architecture to deliver industry-leading conversational AI, speeding up large language models (LLMs) by 30X.
H100 also includes a dedicated Transformer Engine to solve trillion-parameter language models.
Generative AI, Large-Scale AI Model Training and Inference (LLMs), High-Performance Computing (HPC) simulations, Scientific Computing, and Data Analytics.
The NVIDIA B200 GPU, powered by the Blackwell architecture, is the world’s most powerful AI chip, designed to power a new era of computing with up to 4x faster training and 30x faster inference than previous generations.
Introducing GPU Spot Instances
Spot instances run on our spare capacity. If that capacity is needed elsewhere in our network, the instance will be terminated. But if your workloads are stateless, fault-tolerant, or can easily pause and resume, you’re getting top-tier GPU performance at a massive discount.
1 – 3 per server
8 – 32
64 – 384 GB
Starting from
€0.57/hour
| GPU | CPU cores | RAM | Spot Price | Price |
|---|---|---|---|---|
| 1 x NVIDIA L4 | 8 cores | 64 GB | €0.57/h €410/mo |
€0.58/h €418/mo |
| 1 x NVIDIA L4 | 12 cores | 128 GB | €0.69/h €497/mo |
€0.70/h €504/mo |
| 1 x NVIDIA L4 | 16 cores | 192 GB | €0.81/h €583/mo |
€0.82/h €590/mo |
| 1 x NVIDIA L4 | 20 cores | 256 GB | €0.93/h €670/mo |
€0.94/h €677/mo |
| 2 x NVIDIA L4 | 12 cores | 128 GB | €1.17/h €842/mo |
€1.18/h €850/mo |
| 2 x NVIDIA L4 | 16 cores | 192 GB | €1.29/h €929/mo |
€1.30/h €936/mo |
| 2 x NVIDIA L4 | 20 cores | 256 GB | €1.41/h €1015/mo |
€1.42/h €1022/mo |
| 2 x NVIDIA L4 | 32 cores | 384 GB | €1.63/h €1174/mo |
€1.64/h €1181/mo |
| 3 x NVIDIA L4 | 16 cores | 192 GB | €1.79/h €1289/mo |
€1.80/h €1296/mo |
| 3 x NVIDIA L4 | 20 cores | 256 GB | €1.91/h €1375/mo |
€1.92/h €1382/mo |
| 3 x NVIDIA L4 | 32 cores | 384 GB | €2.13/h €1534/mo |
€2.14/h €1541/mo |
1 – 3 per server
8 – 32
64 – 384 GB
Starting from
€0.83/hour
| GPU | CPU cores | RAM | Spot Price | Price |
|---|---|---|---|---|
| 1 x NVIDIA L40S | 8 cores | 64 GB | €0.83/h €600/mo |
€1.11/h €800/mo |
| 1 x NVIDIA L40S | 12 cores | 128 GB | €0.94/h €675/mo |
€1.25/h €900/mo |
| 1 x NVIDIA L40S | 16 cores | 192 GB | €1.15/h €825/mo |
€1.53/h €1100/mo |
| 1 x NVIDIA L40S | 20 cores | 256 GB | €1.35/h €975/mo |
€1.81/h €1300/mo |
| 2 x NVIDIA L40S | 12 cores | 128 GB | €1.56/h €1125/mo |
€2.08/h €1500/mo |
| 2 x NVIDIA L40S | 16 cores | 192 GB | €1.98/h €1425/mo |
€2.64/h €1900/mo |
| 2 x NVIDIA L40S | 20 cores | 256 GB | €2.40/h €1725/mo |
€3.19/h €2300/mo |
| 2 x NVIDIA L40S | 32 cores | 384 GB | €2.81/h €2025/mo |
€3.75/h €2700/mo |
| 3 x NVIDIA L40S | 16 cores | 192 GB | €2.81/h €2025/mo |
€3.75/h €2700/mo |
| 3 x NVIDIA L40S | 20 cores | 256 GB | €3.23/h €2325/mo |
€4.31/h €3100/mo |
| 3 x NVIDIA L40S | 32 cores | 384 GB | €3.65/h €2625/mo |
€4.86/h €3500/mo |
1 – 8 per server
12 – 96
240 – 1920 GB
Starting from
€1.78/hour
| GPU | CPU cores | RAM | Spot Price | Price |
|---|---|---|---|---|
| 1 x NVIDIA H100 | 12 cores | 240 GB | €1.78/h €1282/mo |
€1.79/h €1289/mo |
| 2x NVIDIA H100 | 24 cores | 480 GB | €3.57/h €2570/mo |
€3.58/h €2578/mo |
| 4x NVIDIA H100 | 48 cores | 960 GB | €7.15/h €5148/mo |
€7.16/h €5155/mo |
| 8 x NVIDIA H100 | 96 cores | 1920 GB | €14.31/h €10303/mo |
€14.32/h €10310/mo |
1 – 8 per server
12 – 96
240 – 1920 GB
Starting from
€3.38/hour
| GPU | CPU cores | RAM | Spot Price | Price |
|---|---|---|---|---|
| 1 x NVIDIA B200 | 12 cores | 240 GB | €3.38/h €2430/mo |
€4.50/h €3240/mo |
| 2x NVIDIA B200 | 24 cores | 480 GB | €6.75/h €4860/mo |
€9.00/h €6480/mo |
| 4x NVIDIA B200 | 48 cores | 960 GB | €13.50/h €9720/mo |
€18.00/h €12960/mo |
| 8 x NVIDIA B200 | 96 cores | 1920 GB | €27.00/h €19440/mo |
€36.00/h €25920/mo |
Skip the API fees and data privacy concerns. Our new tutorial shows you how to spin up an UpCloud GPU and run powerful open-weight models like Mistral-7B with Ollama. Go from deployment to inference on your own private, high-performance server.
With our GPU servers, you keep full control over your data, your models, and your infrastructure.
Privacy-first by design: Hosted in Finland, your AI workloads stay in GDPR-compliant data centers with strong jurisdictional protections.
Open-source aligned: Whether you’re fine-tuning LLaMA, deploying open-source LLMs, or building on PyTorch etc, our infrastructure doesn’t impose restrictions or hidden service layers.
No platform dependencies: Unlike hyperscalers, we don’t force you into proprietary ML platforms or opaque orchestration layers.
UpCloud GPU Servers offer a range of GPU models with the unique ability to scale the server and GPUs as needed.
Start with a smaller card for development purposes, then migrate to use more powerful cards for production use, or vice-versa, all without needing to reinstalling the server or losing any data.
It’s as easy as shutting down the server, changing the plan and powering the server back up again!
| NVIDIA L4 | NVIDIA L40S | NVIDIA H100 | NVIDIA B200 | |
|---|---|---|---|---|
| Primary use case | Energy-efficient accelerator for AI inference, video transcoding, graphics/VDI and edge deployments |
Multi-workload “universal” GPU – GenAI, LLM training & inference, 3D graphics, rendering, video |
High-traffic inference, massive batch processing and large-model training |
Trillion-parameter model inference, model training, running complex models in real-time |
| Architecture | Ada Lovelace | Ada Lovelace | Hopper | Blackwell |
| GPU Memory | 24GB GDDR6 | 48GB GDDR6 | 80GB HBM3e | 192GB HBM3e |
| Memory Bandwidth | 300GB/s | 864GB/s | 3.35TB/s | 8.0TB/s |
| Peak compute | FP32 30.3 TFLOPS FP8 Tensor 0.48 PFLOPS |
FP32 91.6 TFLOPS FP8 Tensor 1.46 PFLOPS |
FP32 67 TFLOPS FP8 Tensor 3.90 PFLOPS |
FP32 74.45 TFLOPS FP8 Tensor 9.00 PFLOPS |
99.999% uptime SLA across Public Cloud, Private Cloud, Managed Kubernetes, Managed Databases, and more.
Choose from a range of Linux and Windows distributions for full control and flexibility. Launch in just 45 seconds through our simple control panel.
Global reach spanning four continents and 15 data centers.
European cloud for a global audience. Help satisfy your compliance requirements with a European, ISO 27001 certified cloud provider.
We pride ourselves on providing outstanding customer service – 24 hours a day, 365 days a year.
In 2024, we were proud to continue our outstanding streak, responding to initial queries in just 46 seconds!
Simple, transparent pricing with no hidden fees. Focus on growth, knowing exactly what you’ll pay every month. Discover our plans, starting at just €3 per month.
Experiment with new technologies, build your servers with full root access, and scale resources on demand—all without worrying about unpredictable costs.
99.999% uptime SLA across Public Cloud, Private Cloud, Managed Kubernetes, Managed Databases, and more.
Choose from a range of Linux and Windows distributions for full control and flexibility. Launch in just 45 seconds through our simple control panel.
Global reach spanning four continents and 15 data centers.
European cloud for a global audience. Help satisfy your compliance requirements with a European, ISO 27001 certified cloud provider.
We pride ourselves on providing outstanding customer service – 24 hours a day, 365 days a year.
In 2024, we were proud to continue our outstanding streak, responding to initial queries in just 46 seconds!
Simple, transparent pricing with no hidden fees. Focus on growth, knowing exactly what you’ll pay every month. Discover our plans, starting at just €3 per month.
Experiment with new technologies, build your servers with full root access, and scale resources on demand—all without worrying about unpredictable costs.
Innovate responsibly. Our Helsinki data center routes excess heat generated from our GPUs directly into the city’s district heating network. This makes our GPU offering in Helsinki one of the most environmentally friendly options on the market, contributing to a greener future while powering your creative endeavors.
Our Helsinki data center is powered by 100% renewables and routes excess heat generated from our GPUs into the city’s district heating network. This makes our GPU offering in Helsinki one of the most environmentally friendly options on the market, contributing to a greener future while powering your AI.
High performance and reliability with no upfront costs or commitments.
GPU FAQ
A GPU server is a cloud server equipped with a dedicated graphics processing unit alongside its CPU, memory and storage. While originally designed for rendering graphics, GPUs excel at parallel processing, performing thousands of calculations simultaneously. This makes them well suited to workloads like AI model training, machine learning inference, video transcoding and scientific computing, where the same operation runs across huge amounts of data at once.
A regular cloud server relies on CPUs, which handle a wide range of tasks sequentially and are ideal for general workloads like web hosting, databases and applications. A GPU server adds a graphics processor with thousands of smaller cores built for parallel work. For tasks like training neural networks or processing video, a GPU can be orders of magnitude faster than a CPU alone. Many teams run both, using Cloud Servers for their applications and GPU Servers for compute-heavy jobs, connected over the same private network.
Common uses include training and fine-tuning machine learning models, running inference for AI applications like chatbots and recommendation engines, video transcoding and streaming, 3D rendering, and scientific simulations. Generative AI has driven much of the recent demand, from serving large language models to image and video generation. See our AI/ML solutions and Audio/Video solutions pages for detail on specific workloads.A
For small models and light experimentation, a CPU can be enough. But for training models of any real size, or serving inference at production speed, a GPU is effectively essential. Even running open-weight models like Llama or Mistral locally becomes impractical on CPU beyond the smallest sizes. A single L4 or L40S server is a sensible starting point for most inference workloads, and you can scale up from there. Our Ollama tutorial shows how to get a model running in under an hour.
Yes. GPU Servers are billed hourly with no long-term commitment, so you can spin one up for a training run and delete it when finished. For maximum savings on interruptible workloads, spot pricing offers the same hardware at a significant discount. Monthly costs are capped at 28 days of usage, so a server running all month never costs more than the listed monthly price.
It depends on the workload. The NVIDIA L4 is the most cost-effective option for AI inference, video transcoding and edge deployments. The L40S is a versatile all-rounder suited to generative AI, mid-scale model training and 3D rendering. The H100 handles high-traffic inference and large model training, while the B200 is built for the largest jobs, including trillion-parameter models. You can also switch between GPU models later by shutting down the server and changing the plan, so you are not locked into your first choice.
For more guidance, see our AI/ML GPU solutions and Audio/Video GPU solutions pages.
All GPU Server plans run in our data center in Helsinki, Finland. The facility is powered entirely by renewable energy and recovers up to 90% of waste heat, which is fed into Helsinki’s district heating network. Read more about our approach on the EU Sovereign Cloud page or explore our data centers.
GPU Servers deploy in just a few minutes through the control panel or API. Our AI/ML-ready Ubuntu template comes with NVIDIA drivers and tooling pre-installed, so you can go from deployment to running workloads without manual setup. If you want a practical walkthrough, our tutorial on running LLMs with Ollama takes you from a fresh server to inference on an open-weight model.
GPU Servers can be attached to your infrastructure alongside our other products and connected over SDN private networks. If you are orchestrating containerised workloads, see Managed Kubernetes for running clusters alongside your GPU capacity.
By using our AI/ML-ready GPU Ubuntu template, the necessary NVIDIA drivers are pre-installed. If you choose another operating system, you will need to manually install the appropriate NVIDIA drivers and CUDA toolkit to enable GPU functionality.Q
No. All GPU Server plans include zero-cost egress, so there are no charges for outbound data transfer. This matters for AI workloads where moving datasets, model weights and inference results in and out of the cloud can otherwise become a significant cost.
Our Fair Transfer Policy applies.
Multi-GPU servers are equipped with more than one dedicated GPU. All GPUs are exposed to your server via PCIe passthrough, allowing you to utilize them for parallel processing, distributed training, or other multi-GPU workloads. You can verify the available GPUs using the nvidia-smi tool.
Yes, each GPU is dedicated to a single server and is not shared with other customers. This ensures consistent performance and security for your workloads.
Our advice is to match the GPU to the job. You can find a great comparison between our GPUs >here<
Memory is usually the deciding factor: your model weights, plus overhead, need to fit in GPU memory for inference, and training needs considerably more. If in doubt, start smaller. You can shut down and switch to a bigger GPU without reinstalling.
Beyond the GPU itself, consider how you deploy it. GPU Servers on hourly billing suit most teams, Managed Kubernetes works for containerised workloads across multiple nodes, and Private Cloud GPUs provide dedicated hardware for organisations with strict compliance or isolation requirements.
If you are unsure, start with a single GPU Server and contact our sales team.
Spot pricing is good for time-insensitive workloads, sold at a discount. A spot instance may be terminated at any given time. Read more at GPU Server Spot pricing documentation.
GPU virtualization is how a physical GPU is made available to virtual servers. There are two main approaches. Shared virtualization splits one GPU between multiple customers, which lowers cost but means unpredictable performance. Passthrough gives a virtual server direct, exclusive access to the physical GPU. UpCloud GPU Servers use passthrough, so the full performance of the GPU is yours alone, with no sharing and no contention from other tenants. You get bare-metal GPU performance combined with the flexibility of a cloud server.
Yes. If you need isolated infrastructure for compliance or security reasons, Private Cloud GPUs provide dedicated GPU hardware within your own private cloud environment. This suits organisations with strict data handling requirements, including public sector and regulated industries.
Start your free trial today and discover why thousands of businesses rely on UpCloud