The spec
One multihull.yaml per service. A pydantic v2 model validates it; hull schema exports JSON Schema for editor completion.
Sections
Section titled “Sections”| Section | Purpose |
|---|---|
container |
Image or build, command, port, health probe, env, secret names |
resources |
GPU classes (any-of, in preference order), GPU count, memory |
scaling |
Concurrency per replica, min and max replicas |
reliability |
Redundancy floor: minWarmProviders, overprovision, fallbackScaleToZero, spot |
targets |
One entry per provider deployment with priority and provider-specific block |
route |
Hostname, protocol, failover policy, auth, optional sticky routing |
GPU classes
Section titled “GPU classes”A normalised enum: L4, A10G, A100-40, A100-80, H100, H200, B200. Each translator maps classes to provider SKUs and reports which it can supply through gpu_inventory, which feeds hull doctor.
Credentials
Section titled “Credentials”Never in the file. Each provider block resolves credentials from its native location by default and accepts credentials: env:NAME or credentials: file:PATH as overrides. Service secrets are listed by name only and mirrored into each provider’s native secret store.
Full example
Section titled “Full example”See examples/llama-8b/multihull.yaml in the repository or the architecture plan.