Skip to content

The spec

One multihull.yaml per service. A pydantic v2 model validates it; hull schema exports JSON Schema for editor completion.

Section Purpose
container Image or build, command, port, health probe, env, secret names
resources GPU classes (any-of, in preference order), GPU count, memory
scaling Concurrency per replica, min and max replicas
reliability Redundancy floor: minWarmProviders, overprovision, fallbackScaleToZero, spot
targets One entry per provider deployment with priority and provider-specific block
route Hostname, protocol, failover policy, auth, optional sticky routing

A normalised enum: L4, A10G, A100-40, A100-80, H100, H200, B200. Each translator maps classes to provider SKUs and reports which it can supply through gpu_inventory, which feeds hull doctor.

Never in the file. Each provider block resolves credentials from its native location by default and accepts credentials: env:NAME or credentials: file:PATH as overrides. Service secrets are listed by name only and mirrored into each provider’s native secret store.

See examples/llama-8b/multihull.yaml in the repository or the architecture plan.