Home › Guides › Draft pipeline

A draft-then-final pipeline across the three MiniMax H3 ids

Updated 2026-10-02

The idea of drafting on a cheaper tier and finishing on a better one is easy to state and easy to implement badly. The usual failure is that nothing stops drafts from multiplying, and nothing records why a shot was promoted. This page builds the pipeline as a small state machine with an explicit gate, using minimax/h3-max-turbo for drafts and minimax/h3 or minimax/h3-max for finals. All three share one request shape, so the stages differ only in config. Cost is handled symbolically: the rates come from the live price table, never from constants in your code.

Set expectations: a draft is a prompt filter, not a preview

Different tiers are different models. A draft on Max Turbo tells you whether a prompt and composition are worth pursuing; it does not promise that the same prompt will produce the same clip on another tier. Treat promotion as "this prompt earned a final attempt", and expect small prompt edits at the next stage. If your acceptance rule needs the final to match the draft closely, measure how often that holds on your own prompts before relying on the pipeline to save money.

The stages

  1. Draft. Short duration, the draft id, the lowest resolution tier you can judge a composition on.
  2. Gate. Decides promote, revise or drop. Cheap automatic checks first, then a human or a scoring step.
  3. Final. The delivery id and delivery resolution, one attempt, with a hard cap on re-renders.

Which resolution tiers each id offers varies by host, so put them in config and read them from the model page rather than assuming.

Config, not constants

STAGES = {
    "draft": {"model": "minimax/h3-max-turbo", "duration_secs": 5, "resolution": "480p"},
    "final": {"model": "minimax/h3-max",       "duration_secs": 5, "resolution": "768p"},
}
MAX_DRAFTS_PER_SHOT = 4
MAX_FINALS_PER_SHOT = 2
BUDGET_UNITS = 500          # an abstract ledger unit you map to money elsewhere

If a tier name is not offered by the host that serves the request, the docs say an unsupported resolution is ignored and the model default is used, not rejected. So the stage config is only a request. Log the size and seconds you actually got back.

The state machine

import time, requests

API = "https://videorouter.sh/api/v1"
H = {"Authorization": "Bearer llmr_sk_live_...", "Content-Type": "application/json"}

def submit(stage, prompt):
    cfg = STAGES[stage]
    r = requests.post(f"{API}/videos", headers=H, timeout=30, json={"prompt": prompt, **cfg})
    r.raise_for_status()
    return r.json()["id"]

def wait(job_id, every=5, limit=1200):
    end = time.time() + limit
    while time.time() < end:
        j = requests.get(f"{API}/videos/{job_id}", headers=H, timeout=30).json()
        if j["status"] in ("completed", "failed"):
            return j
        time.sleep(every)
    return None

def auto_checks(job):
    # cheap, deterministic rejects before anyone looks at the clip
    return job is not None and job["status"] == "completed" and bool(job.get("data"))

def run_shot(shot, ledger, review):
    prompt, drafts, finals = shot["prompt"], 0, 0
    while drafts < MAX_DRAFTS_PER_SHOT and ledger.can_spend("draft"):
        job = wait(submit("draft", prompt))
        drafts += 1
        ledger.charge("draft", shot["id"])
        if not auto_checks(job):
            continue                               # failed upstream: not billed, just try again
        verdict, prompt = review(shot, job, prompt)  # ("promote"|"revise"|"drop", new_prompt)
        if verdict == "drop":
            return "dropped"
        if verdict == "revise":
            continue
        while finals < MAX_FINALS_PER_SHOT and ledger.can_spend("final"):
            fjob = wait(submit("final", prompt))
            finals += 1
            ledger.charge("final", shot["id"])
            if auto_checks(fjob) and review(shot, fjob, prompt)[0] == "promote":
                return fjob["data"][0]["url"]
        return "final_rejected"
    return "out_of_attempts_or_budget"

Note where the loops stop. Each shot has a draft cap and a final cap, so no single stubborn prompt can drain the budget. A failed job is not billed, so the loop treats it as a free retry, but it still counts against the attempt cap to avoid an endless loop on a prompt that always fails.

The ledger: budgeting symbolically

class Ledger:
    def __init__(self, budget, unit_cost):
        self.left, self.unit_cost = budget, unit_cost   # unit_cost[stage] from your rate sheet
    def can_spend(self, stage):
        return self.left >= self.unit_cost[stage]
    def charge(self, stage, shot_id):
        self.left -= self.unit_cost[stage]
        # append (shot_id, stage, unit_cost) to a durable log here

Build unit_cost from the live table: stage cost is seconds * rate(model, tier, host) * (1 + fee), with the platform fee a flat 2% on image and video. Refresh the rate sheet on a schedule. Hard-coding a number in the file is how a "cost-controlled" pipeline quietly stops being one. The expected cost of a shot is then D * c_draft + p * F * c_final, where D is average drafts per shot, p the fraction that get promoted and F average finals per promoted shot. Measure D, p and F on a pilot of your own prompts before you commit to a volume. The pipeline pays off only when drafts are cheap enough and promotion selective enough that D * c_draft stays below what rendering every shot on the final tier would cost.

The gate

Keep the gate boring and recorded. A workable order: automatic rejects (job failed, no data), then a rubric score (does it show the subject, is the motion plausible, any visible artifacts), then a human only for the borderline band. Store the verdict next to the job id and prompt, so that when a final is rejected you can see whether the gate let through a bad draft or the final tier changed the result.

Running it safely

The call details are in the integration guide, and the variant guide explains where each id fits. Get a key at videorouter.sh/signup.

Frequently asked questions

Does a Max Turbo draft predict the final clip?

No. They are different models, so a draft shows whether a prompt and composition are promising, not what the final will look like. Measure how often drafts predict finals on your own prompts.

Which MiniMax H3 id should I use for finals?

Either minimax/h3 or minimax/h3-max. Run your own prompts through both and compare the per-second price in the live table against how often the output is acceptable.

Are failed draft jobs billed?

Jobs that fail upstream are not billed. Jobs that complete are billed once at creation, including drafts you reject.

How do I cap the spend of a pipeline like this?

Use per-shot attempt caps and an in-code ledger fed from the live rate sheet, plus a monthly spend cap on the API key so a bug returns a 402 instead of an overrun.

Keep reading

Using MiniMax is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →