Designing for host outages with MiniMax video
Updated 2026-10-02
MiniMax H3 is sold by many hosts, which is useful for reliability only if your integration actually uses that breadth. This guide covers what the platform does on its own, what you can control, and where a naive retry loop quietly doubles your bill.
What automatic failover does
With no provider field, a request goes to the cheapest host that is currently healthy. If that host rejects the job, the platform retries once on the same host, then walks the other hosts serving the same checkpoint, cheapest first, until one accepts. This is free and on by default. It covers the submission call only: it handles a host that refuses the job before an id exists. A job that was accepted and then stalls or fails mid-generation is a different problem, covered below.
The response tells you what happened: provider names the host that took the job, and fallback_used and fallbacks_tried describe the walk. Log them. They are the fastest way to see that one of your hosts has been degraded for a week.
Floating versus pinning
| Approach | How | You get | You give up |
|---|---|---|---|
| Float | "model": "minimax/h3" | Cheapest healthy host, full failover | Run-to-run consistency in host |
| Prefer a host | "model": "minimax/h3/fal" | That host first, rest of the pool as fallback | Nothing on availability; price may be higher than the cheapest |
| Rank differently | "provider": {"sort": "reliability"} or "policy": "most_reliable" | Order by measured 24h success rate instead of price | Lowest price |
| Allow-list with failover | "provider": {"only": ["fal", "atlascloud"]} | Failover among hosts you reviewed | Hosts outside the list |
| Hard pin | "provider": {"only": ["fal"], "allow_fallbacks": false} | Exactly one attempt on that host | All failover |
Note that the model/host suffix is a soft preference, not a hard pin. If you need a guarantee that a request never leaves a host (for data-handling reasons, say), use the allow-list form with allow_fallbacks: false. For most projects the allow-list with fallbacks is the sweet spot: you choose the hosts you trust, and you keep failover among them. You can also filter by measured performance with provider.preferences, for example max_p95_latency_ms and min_success_rate, which drops hosts below a threshold. An over-strict threshold degrades to "ignore this filter" instead of failing the request.
Stalled jobs: timeout hedging and what it costs
Video jobs take minutes, so retrying from your side after a short wait leaves the first job running too. The platform offers an opt-in hedge:
{
"model": "minimax/h3/fal",
"prompt": "a paper airplane gliding over a city",
"failover": {"on_timeout_sec": 60, "max_attempts": 2}
}
If the job has not finished by on_timeout_sec, it is resubmitted to the next-cheapest untried host without cancelling the original. Both attempts are billed if both land. A background sweep checks overdue jobs about every 30 seconds, max_attempts defaults to 2 and is capped at 3, and the creation response includes hedge_armed: true once a hedge exists. Polling either job id returns whichever finishes first, tagged pricing_mode: "timeout_failover_duplicate_billed" once a hedge is in play.
Treat it as paying for a second chance at finishing sooner, not as a free retry. Set the timeout well above your normal completion time so that you hedge only on real stalls. Use it for deadline-bound finals, not bulk drafts.
Retry design that does not double-bill
Separate three kinds of failure, because they deserve different handling:
- Request rejected (HTTP 4xx/5xx at creation). 429 means wait for the
Retry-Afterheader. 400 means fix the request. 402 means funds or cap.upstream_error(5xx) means every candidate failed and nothing was billed, so a delayed retry is reasonable. - Job
failed. Read the error. Content-policy failures should not be retried unchanged. - Job slow. It is still running and still billed. Do not resubmit from your side. Either keep waiting or use the hedge deliberately.
import random, time, requests
API = "https://videorouter.sh/api/v1"
H = {"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"}
def create(payload, max_tries=4):
for attempt in range(max_tries):
r = requests.post(f"{API}/videos", headers=H, json=payload, timeout=30)
if r.status_code == 429:
time.sleep(int(r.headers.get("Retry-After", 5)))
continue
if r.status_code in (500, 502, 503, 504):
time.sleep(min(60, 2 ** attempt + random.random()))
continue
r.raise_for_status() # 400/401/402/403: do not retry
return r.json()
raise RuntimeError("could not create job")
def wait(job_id, every=5, max_wait=1200):
deadline = time.time() + max_wait
while time.time() < deadline:
j = requests.get(f"{API}/videos/{job_id}", headers=H, timeout=30).json()
if j["status"] in ("completed", "failed"):
return j
time.sleep(every)
return None # still running: decide explicitly, do not resubmit blindly
def generate(payload):
job = wait(create(payload)["id"])
if job is None:
raise TimeoutError("still running; check the job id later")
if job["status"] == "failed":
raise RuntimeError(job.get("error"))
return job["data"][0]["url"]
The one deliberate gap is the None branch: a job that outlives your wait window is returned to the caller as a question, not retried automatically. Store the job id and check again later, which is free. Add an idempotency layer on your side (a key per user action) so a double click cannot create two jobs.
Cross-model fallback
If all hosts of a model are down, a top-level models array can move on to a different checkpoint, such as "models": ["minimax/h3-max"]. Only use it where a different model's output is acceptable, and record which model served the job. See the decision framework for how to choose a backup family.
A short checklist
- Float by default; use an allow-list when you have compliance or consistency needs.
- Log
provider,fallback_usedandfallbacks_triedon every job. - Never resubmit a slow job unless you intend to pay for both.
- Use
failover.on_timeout_seconly for deadline-bound work. - Remember failed-upstream jobs are not billed; accepted and discarded ones are.
See the quickstart for the minimal call and videorouter.sh/signup to get a key.
Frequently asked questions
What happens if a MiniMax host is down?
Unpinned requests retry once on the same host, then walk the other hosts serving the same model, cheapest first, until one accepts the job. This is free and on by default.
Is minimax/h3/fal a hard pin?
No, it is a soft preference: that host is tried first and the rest of the pool is still available. For a hard pin use provider.only with allow_fallbacks: false.
What does failover.on_timeout_sec do?
If an accepted job has not finished by the deadline, it is resubmitted to the next-cheapest untried host without cancelling the original. Both attempts are billed if both complete.
Are failed jobs billed?
Jobs that fail upstream are not billed. A job that succeeds is billed even if you discard the result, and a slow job that is still running is still billed.
Keep reading
- MiniMax H3 vs H3 Max vs H3 Max Turbo — Which Variant for Which Job
- MiniMax Video API Use Cases: What to Build with the H3 Family
- MiniMax vs Kling vs Seedance API: A Decision Framework
- How to Call the MiniMax Video API: curl and Python Guide
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →