The inference platform for NVIDIA and AMD
infervisor — runtime plowrt · OpenAI-compatible
infervisor:~$ plowrt serve
One inference runtime
for every modality.
LLMs, diffusion, and voice on a single engine, rebuilt from first principles in Rust. Models compile ahead of time into a fixed on-device execution plan — every compiler step gated by machine-checked Lean proofs — and run natively on NVIDIA and AMD, on-prem or in your cloud. Measured head-to-head against vLLM.
LLMs, diffusion, and voice flow into the Infervisor engine, which runs them on NVIDIA and AMD accelerators. INFERVISOR LLMs diffusion voice NVIDIA AMD
modalities LLM · diffusion · voice targets NVIDIA · AMD deploy on-prem│cloud serving
WHAT IT IS

Optimized at the system level. Not just the kernel.

The inference layer the category is missing — a single engine, rethought from first principles, that replaces a stack of vendor-locked runtimes.

01

System-level, from first principles

Not a bag of hand-tuned kernels. We’re rebuilding the inference engine end to end — the whole model compiled ahead of time into one fixed execution plan, so the accelerators run it without checking back with the host CPU.

02

One runtime, every modality

LLMs, diffusion, and voice served from a single dataflow engine — one stack to deploy, monitor, and scale instead of three.

03

Any accelerator

One model graph, one compiler, native plans for NVIDIA and AMD — every kernel our own, no vendor libraries underneath. On-prem or in your cloud, with no lock-in.

PRODUCT · PLOW

Plow — a Packet Language
for On-device Workers.

The engine inside Infervisor — where a worker is a warp on NVIDIA and a wave on AMD. A model becomes a validated, target-specific execution artifact before it reaches production, so the serving path can stay small and predictable.

Plow compiles a transformer model into one prepared serving artifact A transformer block — norm, attention with q k v, residual adds, and mlp, repeated across N layers — is funneled through the plow compiler: analyze, optimize, validate. The result is a single prepared artifact that crosses the deploy boundary onto the small plowrt runtime running on NVIDIA or AMD, which emits a steady stream of tokens. COMPILE — ONCE SERVE — EVERY SESSION MODEL CHECKPOINT norm attention q · k · v + mlp + × N layers plow compiler analyze · optimize · validate deploy one artifact plowrt runs the plan NVIDIA · AMD same compiled-plan contract steady tokens · no scheduler jitter
01

Plan before serving

The expensive work of understanding and optimizing a model happens before deployment, leaving fewer decisions on the request path.

02

Keep the runtime small

A lightweight Rust runtime loads a prepared artifact and executes it with the host CPU outside the steady-state model loop.

03

Measure every claim

Correctness gates come first. Latency, serving overhead, startup, and multi-model behavior are reported as separate results.

Read the technical overview for the architecture, correctness guarantees, and evaluation method. Measured results live on the blog.

PRODUCT · ORCHESTRATOR

The control plane for inference fleets.

Where Kubernetes schedules containers and Slurm schedules batch jobs, the Orchestrator runs models: it places compiled plowrt plans across GPUs, nodes, and vendors, routes traffic to them, and meters every request.

01

One plane for mixed fleets

Enroll NVIDIA and AMD nodes into clusters from one console, with live per-GPU telemetry from Plow SMI.

02

Deployments, not containers

Versioned model deployments with replicas, serving configs, and rollout history — models are the unit you manage.

03

OpenAI-compatible routing

One endpoint in front of every deployment: weighted routing, capacity limits, circuit breakers, and automatic failover.

04

Keys, usage, and access

Scoped API keys, per-key usage metering, role-based access, and single sign-on with Google or passkeys.

05

Jobs and reservations

Reserve GPUs, launch jobs from templates, and watch job health, with SSH access to the nodes you hold.

06

Evals in the loop

Run evaluation suites against your deployed models on the same fleet that serves them.

See the Orchestrator for how placement and telemetry fit together, or request a demo on your own fleet.

COMPANY

An AI inference company,
built on systems research.

We study where serving time actually goes — host orchestration, synchronization, startup, multi-model hand-offs — and build the systems that remove it: a compiler that plans the model before deployment, a small runtime that executes it on NVIDIA and AMD, and a control plane that runs the fleet. Every result is correctness-gated, measured by workload phase, and published with the conditions behind it.

Building production AI inference?
See it on your workload.

A 30-minute walkthrough of plowrt and the Orchestrator, on the models and accelerators you run.

request a demo email us →