Skip to content

Reliability principle

Multihull exists to keep models serving. Defaults spend money to buy redundancy.

reliability:
minWarmProviders: 2
overprovision: 1.4
fallbackScaleToZero: false
spot: false
  • At least two providers warm.
  • Fallbacks never scale to zero.
  • On-demand before spot.
  • Capacity over-provisioned by 1.4x, the Envoy priority-spillover factor.

Cheapest-GPU placement, spot, and scale-to-zero fallbacks exist but must be enabled, and none can lower redundancy below the configured floor.

hull doctor and hull plan surface credential, quota and GPU-availability problems up front. At runtime, failover is invisible to callers.

Files in git, SQLite or object storage for state, gRPC or a JSON file between control plane and router. No message bus, no database cluster required.