GPU Platform as a Service

Sovereign GPaaS

Turn the accelerators you already own into a metered service.

Sovereign GPaaS is a control plane that pools the accelerators across your estate into a single schedulable resource. Teams request a fraction of a GPU, a whole card or a multi-node group through one API, and the platform places the work, enforces the quota and records the consumption. It runs on the hardware you already own, inside your own data centre, on upstream Kubernetes with no proprietary fork.

An accelerator that sits idle still takes rack space, power and cooling, and the capital behind it earns nothing while it waits. Pooling, fractional allocation and queueing put more of that capital in front of real work, and per-tenant metering lets you charge the consumption back to a business unit or bill an external tenant for it. Utilisation is largely a platform decision rather than a hardware one: it depends on how work is placed, partitioned and queued above the cards.

Architecture

How it is put together

06

Consumption and chargeback

A self-service portal and API where teams request capacity. Every allocation is metered by tenant, project and duration, and those records feed billing, showback or internal chargeback.

05

Multi-tenant isolation

Namespace, network and device boundaries keep tenants apart, with audit-ready records of who consumed what and when.

04

Workload scheduling

Queues, priorities, fair-share policy and gang scheduling place training, fine-tuning and inference jobs onto the right accelerator without the requesting team needing to know which card it landed on.

03

GPU pooling and partitioning

Software slicing and hardware-level partitioning turn discrete cards into one pool of fractions, whole devices and multi-node groups.

02

Cloud-native substrate

Upstream Kubernetes with accelerator-aware networking, including RDMA, plus observability and lifecycle management across x86 and ARM hosts.

01

Heterogeneous hardware

Accelerators from several vendors and several hardware generations, presented to workloads through one consistent interface.

Capabilities

Fractional GPU allocation

Split a single accelerator across several workloads so small jobs stop holding whole cards they cannot fill.

Unified accelerator pool

One pool spanning several vendors and hardware generations, so teams draw capacity from the estate rather than from a named machine.

Quota and queueing

Per-tenant quotas, priorities, fair-share queues and gang scheduling keep the pool busy without letting one team starve the rest.

Per-tenant metering

Every allocation is recorded by tenant, project and duration, ready for chargeback, showback or an external invoice.

Hard multi-tenancy

Isolation at namespace, network and device level, with memory and compute limits enforced per allocation and audit-ready access records.

Inference-ready serving

Serving runtimes with prefill and decode separation and attention-cache reuse present the pool as a metered inference endpoint rather than raw capacity.

Self-service provisioning

Teams draw capacity from one API and portal instead of raising a ticket for a physical machine.

Open and portable

Standard Kubernetes interfaces, full source access and knowledge transfer, so your own engineers can operate the platform and move it if they choose.

Where it fits

  • A research team needs an accelerator for a few hours a week, but holds a whole card permanently because there is no way to give them less.
  • Risk, fraud and customer-service teams each bought their own GPUs, and none of them can lend spare capacity to the others.
  • Finance cannot say which business unit is responsible for the accelerator line in the infrastructure budget.
  • An operator wants to sell GPU capacity to external tenants but has no way to isolate, meter and invoice them.
  • An auditor asks which tenant had access to which accelerator and which data, and the answer has to come from records rather than recollection.

What you end up with

The customer ends up owning a metered, multi-tenant GPU service on their own accelerators, in their own data centre, with the scheduling policy, isolation boundaries and billing records under their control and their engineers trained to run it.

The rest of the stack

Let's meet each other online!

Easily schedule your desired time to get a FREE 30-minute consultation with our expert team.

Ali Salmaji

Ali Salmaji

DevOps Solution Architect

Do you need more help?

Use the calendar below and choose a free time to arrange a meeting instantly.

Book a meeting