Targets and failover
A target is one deployment of the service on one provider. Targets carry a priority; lower is preferred.
targets: - provider: gke-prod type: kubernetes priority: 1 replicas: {min: 2} - provider: modal-main type: modal priority: 2 - provider: runpod-eu type: runpod priority: 3Routing between tiers
Section titled “Routing between tiers”Default scoring is priority_spillover: power-of-two-choices least-outstanding inside a tier, with gradual spill to the next tier as the current one loses healthy capacity (Envoy priority levels, 1.4 overprovision factor). Alternatives: weighted, ewma_latency on time-to-first-token, locality.
Failover per request
Section titled “Failover per request”route: failover: {policy: priority, retryOn: [5xx, timeout, capacity], maxRetries: 2}The router retries only when all of these hold:
- zero bytes committed to the client,
- request body fully buffered,
- the method is idempotent, or an
Idempotency-Keyis present, or the outcome is Capacity or a connect failure.
Retries go to a different provider, at most two per request, within a retry budget of 20% of live volume per service.
Streaming
Section titled “Streaming”Once the first SSE byte is forwarded the stream is committed. On mid-stream failure the router emits a final upstream_disconnected, retryable: true event; the SDK retries with the same idempotency key. Mid-stream failover is never attempted.
States you will see
Section titled “States you will see”| Pill | Meaning |
|---|---|
| Ready | Replicas ready, circuit closed |
| Degraded | Serving, but warm checks or TTFT are slow; limiter lowered |
| Failover | Circuit open; traffic spilled to the next tier |