Multihull
Deploy hot GPU containers to many providers. One URL. Automatic failover.
One spec
Describe image, GPU class, scaling and targets once in multihull.yaml. Each provider gets its native payload.
Always warm
At least two providers stay hot by default. Fallbacks never scale to zero. Health beats price.
Per-request failover
A stateless Rust router retries across providers before the first byte is committed. Callers see one URL.
A minimal spec
Section titled “A minimal spec”apiVersion: multihull/v1name: llama-8b
container: image: ghcr.io/acme/vllm-llama:1.4.0 port: 8000 health: {path: /health}
resources: gpu: [L4, A10G]
targets: - provider: gke-prod type: kubernetes priority: 1 - provider: modal-main type: modal priority: 2
route: hostname: llama.api.acme.com protocol: openaiWhat hull status shows
Section titled “What hull status shows”| Target | State | Replicas | GPU |
|---|---|---|---|
| gke-prod | Ready | 2/2 | L4 |
| modal-main | Degraded | 1/1 | A10G |
| runpod-eu | Serving failover | 1/1 | L4 |
Start with the Quickstart, or read how targets and failover work.