Skip to content

Multihull

Deploy hot GPU containers to many providers. One URL. Automatic failover.

One spec

Describe image, GPU class, scaling and targets once in multihull.yaml. Each provider gets its native payload.

Always warm

At least two providers stay hot by default. Fallbacks never scale to zero. Health beats price.

Per-request failover

A stateless Rust router retries across providers before the first byte is committed. Callers see one URL.


multihull.yaml
apiVersion: multihull/v1
name: llama-8b
container:
image: ghcr.io/acme/vllm-llama:1.4.0
port: 8000
health: {path: /health}
resources:
gpu: [L4, A10G]
targets:
- provider: gke-prod
type: kubernetes
priority: 1
- provider: modal-main
type: modal
priority: 2
route:
hostname: llama.api.acme.com
protocol: openai
Target State Replicas GPU
gke-prod Ready 2/2 L4
modal-main Degraded 1/1 A10G
runpod-eu Serving failover 1/1 L4

Start with the Quickstart, or read how targets and failover work.