What a dedicated CPU server is good for in an AI pipeline

We do not sell GPU nodes today. If a training run or a high-volume inference workload needs GPU throughput, sales quotes that hardware directly rather than listing it as a catalogue SKU. What ships from the catalogue is single-tenant CPU bare metal, which covers more of a typical AI pipeline than the model call itself: retrieval, embeddings, feature pipelines and the CPU-bound stages around a model. None of it is shared with another tenant, so a neighbor’s batch job cannot show up as latency on your inference endpoint. See current configurations on our dedicated servers page.

CPU inference and embeddings

Quantized inference through llama.cpp, ONNX Runtime or OpenVINO, and embedding generation, run on cores no other tenant shares. Nothing above your own OS is deciding when your process runs.

Vector search and retrieval

A RAG pipeline's retrieval step, a vector index such as FAISS, Qdrant or Milvus, or a feature store, depends on consistent local disk throughput more than on a GPU. Bare metal gives it dedicated NVMe, not a shared volume.

Data preparation and orchestration

Tokenization, cleaning, chunking and the orchestration that feeds a training run are CPU-bound stages. They do not need a GPU, and on a shared host they compete with one for the scheduler's attention.

  • Inference
  • RAG
  • Vector search
  • Data pipelines

Where CPU stops and a GPU quote starts

The catalogue on this page is CPU bare metal: single-tenant, no virtualization, hourly while you test a pipeline against it. Training a model from scratch, or running inference at a scale that needs GPU throughput, is different hardware that we do not list as a SKU. Sales quotes GPU capacity directly; the CPU-bound stages of the same pipeline can go straight to checkout.

What the platform runs on

  • Windows Server and Linux one-click images
  • Full root access, provisioned from the cloud console
  • NVMe local storage for vector indexes and datasets
  • Hourly billing while you test a pipeline, monthly once it is in production

Network and reliability

  • AS55285 is our own network, across seven live metros: Amsterdam, Frankfurt, London, Chicago, Dallas, Los Angeles and New Jersey
  • 100% network uptime SLA: 5% of the monthly fee credited per hour of downtime we caused, up to the full month
  • DDoS mitigation inline on the port, active by default
  • Public looking glass to test the path from five of our seven metros before you order

Private network between your servers

  • Datasets and training traffic move over a layer 2 network between your own servers, off their public interfaces
  • The same private network can extend to your servers in our other datacenters
  • Available at no cost, managed from the cloud console or by API

Frequently asked questions

What ships from the catalogue today, what needs a quote, and how the network and SLA apply to an AI workload.

Not as a catalogue SKU today. What ships from our dedicated servers page is single-tenant CPU bare metal. If your workload needs GPU throughput, for training or high-volume inference, contact sales with the model and scale you are targeting and we will quote hardware directly.

Ask sales about GPU capacity

CPU bare metal in the catalogue, GPU nodes by quote
Order a CPU configuration directly from the catalogue, hourly or monthly. For GPU capacity, tell sales what you are training or serving and the scale you need; that hardware is quoted, not listed.