Home › Guides › Failover and reliability

Designing for host outages with MiniMax video

Updated 2026-10-02

MiniMax H3 is sold by many hosts, which is useful for reliability only if your integration actually uses that breadth. This guide covers what the platform does on its own, what you can control, and where a naive retry loop quietly doubles your bill.

What automatic failover does

With no provider field, a request goes to the cheapest host that is currently healthy. If that host rejects the job, the platform retries once on the same host, then walks the other hosts serving the same checkpoint, cheapest first, until one accepts. This is free and on by default. It covers the submission call only: it handles a host that refuses the job before an id exists. A job that was accepted and then stalls or fails mid-generation is a different problem, covered below.

The response tells you what happened: provider names the host that took the job, and fallback_used and fallbacks_tried describe the walk. Log them. They are the fastest way to see that one of your hosts has been degraded for a week.

Floating versus pinning

ApproachHowYou getYou give up
Float"model": "minimax/h3"Cheapest healthy host, full failoverRun-to-run consistency in host
Prefer a host"model": "minimax/h3/fal"That host first, rest of the pool as fallbackNothing on availability; price may be higher than the cheapest
Rank differently"provider": {"sort": "reliability"} or "policy": "most_reliable"Order by measured 24h success rate instead of priceLowest price
Allow-list with failover"provider": {"only": ["fal", "atlascloud"]}Failover among hosts you reviewedHosts outside the list
Hard pin"provider": {"only": ["fal"], "allow_fallbacks": false}Exactly one attempt on that hostAll failover

Note that the model/host suffix is a soft preference, not a hard pin. If you need a guarantee that a request never leaves a host (for data-handling reasons, say), use the allow-list form with allow_fallbacks: false. For most projects the allow-list with fallbacks is the sweet spot: you choose the hosts you trust, and you keep failover among them. You can also filter by measured performance with provider.preferences, for example max_p95_latency_ms and min_success_rate, which drops hosts below a threshold. An over-strict threshold degrades to "ignore this filter" instead of failing the request.

Stalled jobs: timeout hedging and what it costs

Video jobs take minutes, so retrying from your side after a short wait leaves the first job running too. The platform offers an opt-in hedge:

{
  "model": "minimax/h3/fal",
  "prompt": "a paper airplane gliding over a city",
  "failover": {"on_timeout_sec": 60, "max_attempts": 2}
}

If the job has not finished by on_timeout_sec, it is resubmitted to the next-cheapest untried host without cancelling the original. Both attempts are billed if both land. A background sweep checks overdue jobs about every 30 seconds, max_attempts defaults to 2 and is capped at 3, and the creation response includes hedge_armed: true once a hedge exists. Polling either job id returns whichever finishes first, tagged pricing_mode: "timeout_failover_duplicate_billed" once a hedge is in play.

Treat it as paying for a second chance at finishing sooner, not as a free retry. Set the timeout well above your normal completion time so that you hedge only on real stalls. Use it for deadline-bound finals, not bulk drafts.

Retry design that does not double-bill

Separate three kinds of failure, because they deserve different handling:

  1. Request rejected (HTTP 4xx/5xx at creation). 429 means wait for the Retry-After header. 400 means fix the request. 402 means funds or cap. upstream_error (5xx) means every candidate failed and nothing was billed, so a delayed retry is reasonable.
  2. Job failed. Read the error. Content-policy failures should not be retried unchanged.
  3. Job slow. It is still running and still billed. Do not resubmit from your side. Either keep waiting or use the hedge deliberately.
import random, time, requests

API = "https://videorouter.sh/api/v1"
H = {"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"}

def create(payload, max_tries=4):
    for attempt in range(max_tries):
        r = requests.post(f"{API}/videos", headers=H, json=payload, timeout=30)
        if r.status_code == 429:
            time.sleep(int(r.headers.get("Retry-After", 5)))
            continue
        if r.status_code in (500, 502, 503, 504):
            time.sleep(min(60, 2 ** attempt + random.random()))
            continue
        r.raise_for_status()      # 400/401/402/403: do not retry
        return r.json()
    raise RuntimeError("could not create job")

def wait(job_id, every=5, max_wait=1200):
    deadline = time.time() + max_wait
    while time.time() < deadline:
        j = requests.get(f"{API}/videos/{job_id}", headers=H, timeout=30).json()
        if j["status"] in ("completed", "failed"):
            return j
        time.sleep(every)
    return None                    # still running: decide explicitly, do not resubmit blindly

def generate(payload):
    job = wait(create(payload)["id"])
    if job is None:
        raise TimeoutError("still running; check the job id later")
    if job["status"] == "failed":
        raise RuntimeError(job.get("error"))
    return job["data"][0]["url"]

The one deliberate gap is the None branch: a job that outlives your wait window is returned to the caller as a question, not retried automatically. Store the job id and check again later, which is free. Add an idempotency layer on your side (a key per user action) so a double click cannot create two jobs.

Cross-model fallback

If all hosts of a model are down, a top-level models array can move on to a different checkpoint, such as "models": ["minimax/h3-max"]. Only use it where a different model's output is acceptable, and record which model served the job. See the decision framework for how to choose a backup family.

A short checklist

See the quickstart for the minimal call and videorouter.sh/signup to get a key.

Frequently asked questions

What happens if a MiniMax host is down?

Unpinned requests retry once on the same host, then walk the other hosts serving the same model, cheapest first, until one accepts the job. This is free and on by default.

Is minimax/h3/fal a hard pin?

No, it is a soft preference: that host is tried first and the rest of the pool is still available. For a hard pin use provider.only with allow_fallbacks: false.

What does failover.on_timeout_sec do?

If an accepted job has not finished by the deadline, it is resubmitted to the next-cheapest untried host without cancelling the original. Both attempts are billed if both complete.

Are failed jobs billed?

Jobs that fail upstream are not billed. A job that succeeds is billed even if you discard the result, and a slow job that is still running is still billed.

Keep reading

Using MiniMax is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →