The control plane
for inference fleets
A GPU operating system built for real-time inference — models are the unit, not containers
Enroll NVIDIA and AMD nodes, deploy compiled plowrt plans with replicas, and serve them through one OpenAI-compatible endpoint — with SLO-aware routing, failover, and per-key usage, from one console, on-prem or in your cloud.
General-purpose schedulers weren't built for token deadlines
Kubernetes was built to schedule containers; Slurm to schedule batch jobs. To both, a GPU is a countable device and a model is opaque cargo. A production inference fleet serves many models on many GPUs under tight latency budgets — and real-time workloads such as voice make the budgets explicit: a reply that streams late is a reply that failed. Running that through a general-purpose layer that assigns work as it arrives recovers utilization, but it reintroduces at fleet scale exactly what host orchestration causes inside a single GPU: decisions on the critical path, queueing variance, and tail latency that inherits every scheduling accident.
Static assignment avoids the jitter but wastes capacity; dynamic scheduling recovers capacity but sells the tail. Real-time serving needs a third option: an operating system for the fleet, built purely around inference.
A GPU OS: models as processes, budgets as the contract
The Orchestrator treats the fleet the way an operating system treats a machine — except its processes are models. Plow compiles a model into a fixed execution plan a GPU runs without consulting the host; the Orchestrator deploys those plans as replicas across GPUs, nodes, and vendors deliberately, ahead of time, routes each request against its latency budget, and supervises the fleet with measured telemetry — so that at serve time the fleet, like the GPU, has nothing left to decide.
A GPU OS where the model is the first-class unit
declare a model and its latency budget; the Orchestrator places plowrt plans onto GPU and CPU nodes ahead of time, routes traffic to them, and supervises them — where Kubernetes schedules containers and Slurm schedules batch jobs, this schedules models
It is built around four commitments:
- SLO-aware serving. Replicas are chosen against declared latency budgets, not best-effort queues — predictability is the contract, and capacity-based admission control rejects overload before it queues.
- Mixed fleets, one plane. NVIDIA and AMD nodes managed from a single control plane, because plow's plans already run natively on both.
- On-prem first. Bare metal, private cloud, sovereign and air-gapped deployments — the same Rust-first, no-Python-serving-stack posture as the engine.
- Measured feedback. Fleet telemetry flows through Plow SMI; placement decisions are calibrated by measurement, in the same spirit as the engine's auto-tuner.
What's in the platform
Clusters and nodes
Enroll NVIDIA and AMD nodes with a node agent, group them into clusters, and watch per-GPU utilization, memory, and health live through Plow SMI.
Model deployments
Versioned deployments with replicas, serving configs, and a full rollout history. Scale a model by changing its replica count, not by editing manifests.
One OpenAI-compatible endpoint
Weighted and SLO-aware routing across replicas and providers, capacity limits, circuit breakers, and automatic failover behind a single API.
Keys, usage, and access
Scoped API keys with per-key policies, usage metered on every request, role-based access, and sign-in with Google or passkeys.
Jobs and reservations
Reserve GPUs, launch jobs from templates, follow job health, and open SSH sessions to the nodes you hold.
Evals on the fleet
Run evaluation suites against deployed models on the same hardware that serves them.