GET /v1/models to discover which models are available to the API and what
they cost. Pass ?type=image, ?type=video, or ?type=audio to filter.
Reading the response
Multiple images per request
Some image models return more than one image from a single call:- seedream-4.5 and seedream-5.0-lite support grouped generation. Set
sequential_image_generation: "auto"andmax_images(up to 4 on seedream-4.5, up to 14 on seedream-5.0-lite). The model returns a group of related images up tomax_imagesand decides the exact count — it may return fewer (a plain prompt often yields one; a prompt that explicitly asks for a set yields several). Usemax_imagesfor this, notnum_images. - seedream-5.0-pro does not support grouped generation — it has no
sequential_image_generationparameter. Usenum_images(1–4) for multiple takes; each is generated and billed independently. It prices by output size (1K = 55, 2K = 110 credits per image) and billsreference_imagesbeyond the first at 5 credits each — see pricing & credits. - seedream-5.0-flash has no grouped generation either: use
num_images(1–4). It costs a flat 25 credits per image at1Kor2K, and itsreference_images(up to 10) are free. - flux-3-image takes
num_images(1–4); each image is generated and billed independently at itssize’s price. Itsreference_images(up to 10) are free. - nano-banana-2.1 takes
num_images(1–4); each image is generated and billed independently at itssize’s price, plus search grounding if you enabled it. Itsreference_images(up to 9) are free.
Grounding
flux-3-image takes agrounding parameter (true by default). When it is on, the model
researches the prompt with web and image search before rendering, which helps when the prompt
depends on what something real looks like. It costs nothing extra. Set grounding: false for
faster, self-contained generation.
Search grounding
nano-banana-2.1 takes two optional parameters that let the model look things up before it renders. Both arefalse by default and are priced per image when enabled:
Because image search includes a web search, the two never add up:
image_search: true costs
20 extra credits per image whether or not google_search is also set. GET /v1/models lists
both amounts under option_pricing, and POST /v1/cost
quotes the exact total for a request.
45 + 10 = 55 credits. Results generated with image_search are returned
as JPEG, whatever output_format asks for.
Transparent backgrounds
seedream-5.0-pro and seedream-5.0-flash take abackground parameter (opaque, the
default, or transparent). Transparent mode is for editing a cutout: send exactly one
reference image, a PNG that already has transparent pixels (a product cutout, sticker or
logo), and the result is a PNG that keeps the transparency. It can’t make a transparent image
from text alone, and multi-reference requests must stay opaque — those requests are rejected
with 400 before anything is charged.
3 × credits_per_generation.
When the job starts we reserve the maximum (max_images × credits_per_generation) and charge
only for the images actually produced.
Video output format
Every video comes back as an MP4. The codec inside depends on the model and the resolution you ask for:
The H.265 tiers are delivered exactly as the model produced them. We don’t re-encode
them to H.264, because that would cost quality at the resolutions where you are paying
for it. The job status says which codec you got:
video_codec is h264 or h265 on every finished video job (GET /v1/videos/{id},
/v1/video-upscales/{id}, /v1/motion-control/{id}, /v1/lip-sync/{id}).
video_codec_note appears only with h265.
Video generation modes
Some video models accept different inputs for different modes.SeeDance 2.5 (flagship)
SeeDance 2.5 (seedance-2.5) is ByteDance’s flagship multimodal model. One endpoint,
three modes — there is no separate video-edit endpoint:
- Text-to-video —
promptonly (aspect_ratioapplies here; other modes are adaptive). - Image-to-video — a start frame (
image), optionally with an end frame (end_image). - Reference-to-video — any mix of up to 15
reference_images(free), up to 5reference_videos(each 2–15s, 30s combined), and up to 5reference_audios(30s combined, free; audio-only input works). Editing or extending an existing clip = passing it as areference_videosentry. Reference media can’t be combined with start/end frames.
reference_videos additionally bill per
second of input at half the output rate (480p = 80, 720p = 160, 1080p = 375 credits per
input second) — use
POST /v1/cost for an exact quote
(it probes your clips).
1080p output is H.265 (HEVC), 10-bit. 480p and 720p are H.264. H.265 does not play in
every browser — see video output format.
SeeDance 2.0
SeeDance 2.0 (seedance-2.0) supports all of these from one endpoint:
- Text-to-video —
promptonly. - Image-to-video — a start frame (
image), optionally with an end frame (end_image) to interpolate between the two. - Reference-images-to-video — up to 9
reference_imagesto guide the result. - Video-to-video — supply a
videoto edit an existing clip. Video inputs are passed by URL (a public URL or one fromPOST /v1/uploads), not inlined as base64. You can also passreference_images(up to 6) and/or anaudiotrack alongside thevideoto guide the edit.
image/end_image) can’t be combined with reference_images or a video.
But reference_images can accompany a video — that runs a reference-guided video-edit
(up to 6 reference images in that mode, vs up to 9 on their own). Pricing is per
second, by resolution (see resolution_pricing): a 5-second 1080p clip costs 5 × 550.
The 4k resolution is available for text/image/reference modes, not video-to-video.
4K output is H.265 (HEVC), 10-bit; 480p to 1080p are H.264 — see
video output format.
Reference audio (optional): attach an audio track (MP3/WAV/OGG, by public URL or one
from POST /v1/uploads) to an image- or video-input request and
SeeDance 2.0 uses it in the generated video. Audio must accompany an image/reference_images
or video input — it can’t be the only input.
Cheaper tiers: seedance-2.0-fast and seedance-2.0-mini support the same modes at
lower per-second rates, in 480p and 720p (no 1080p/4k). Check each model’s
resolution_pricing in GET /v1/models, or POST /v1/cost
to price an exact request.
MiniMax H3
MiniMax H3 (minimax-h3) is a multimodal 2K model with a different reference surface:
- Text-to-video —
promptonly (aspect_ratioapplies here; other modes are adaptive). - Image-to-video — a start frame (
image), optionally with an end frame (end_image). - Reference-to-video — mix up to 5
reference_images, up to 3reference_videosclips, and up to 3reference_audiosclips in one request. Clips are passed by URL (public or fromPOST /v1/uploads); each clip must be 2–15s, with at most 15s combined per media type.
MiniMax H3 Max
MiniMax H3 Max (minimax-h3-max) is a post-trained H3 tuned for stronger prompt
adherence and aesthetics, with native audio. One endpoint covers every mode — the inputs
decide which:
- Text-to-video —
promptonly;aspect_ratiodefaults to16:9. An optionalaudioURL pins a soundtrack (it replaces the generated audio, trimmed to the video length). - Image-to-video —
image(first frame) and/orend_image(last frame). An end frame on its own is allowed. The output follows the frame image, soaspect_ratiois ignored.audioworks here too. - Reference-to-video — up to 5
reference_images, 3reference_videos(2–15s each, 15s combined) and 3reference_audios(2–15s each, 15s combined). Address them in the prompt by order: “Image 1”, “Video 1”, “Audio 1”.aspect_ratiodefaults toadaptive(the model picks); an explicit ratio is honored. - Extend — pass a
video(1.6–60s, ≤50MB, aspect ratio between 2:5 and 5:2) and describe what happens next.length_secondsis the new footage appended; the output is the source followed by the continuation.aspect_ratiodefaults toauto(keep the source framing); any other ratio crops.
480p, 768p (default) and 1080p; 2k is available only when
extending. Any integer duration from 5 to 15 seconds. Reference media can’t be
combined with start/end frames, and video can’t be combined with any other media input.
Wan 3.0
Wan 3.0 (wan-3.0-video) is Alibaba’s all-in-one model — one endpoint, three modes,
no separate video-edit endpoint (the SeeDance 2.5 shape):
- Text-to-video —
promptonly.aspect_ratiois honored in every mode — an explicit ratio reframes the output even with a first frame or reference media, and the defaultadaptivelets the model pick a suitable ratio from the inputs and prompt. - Image-to-video — a first frame (
image), optionally with a last frame (end_image) to interpolate between the two. - Reference-to-video — any mix of up to 10
reference_images(free), up to 5reference_videos(each 2–15s, 15s combined), and up to 5reference_audios(15s combined, free). Editing, replicating the style or camera move of, or extending an existing clip = passing it as areference_videosentry. In the prompt you can address assets by their order in each type: “Image 1 hands Video 1 the guitar…”. Reference media can’t be combined with first/last frames.
Unlike MiniMax H3 and the SeeDance families, Wan 3.0 never bills input media —
reference_videos clips are free, and you pay only for the output seconds you request.
The one constraint: with a video input, the input duration plus length_seconds must not
exceed 30 seconds.wan-3.0-video-prime is the high-speed tier: the identical API
surface and modes with significantly faster end-to-end generation, at 85/170/340
credits per output second (480p/720p/1080p). Input media is free there too.
Wan 2.7
Wan 2.7 (wan-2.7-video) serves three modes from one endpoint:
- Text-to-video —
promptonly. - Image-to-video — a start frame (
image), optionally with an end frame (end_image) to interpolate between the two. - Video editing — supply a
videoand the model applies yourpromptto that clip (restyle it, change the scene, alter a subject’s action or the camera move). The video is passed by URL (public or fromPOST /v1/uploads), not inlined as base64. Up to 4reference_imagesmay accompany it to guide the edit — for example, supplying the person or style to apply to the clip.
video can’t be combined with image/end_image. Output is 720p or 1080p; text- and
image-to-video run 5, 10, or 15 seconds.
Video editing behaves differently from the other two modes in two ways worth planning for:
the input clip must be 2–15 seconds, and the output length is taken from that clip,
so length_seconds is ignored.
P-Video
P-Video (p-video) is a fast, low-cost model with three modes:
- Text-to-video —
promptonly. - Image-to-video — a start frame (
image), optionally with an end frame (end_image) to interpolate between the two. - Audio-driven — supply an
audiotrack (MP3/WAV/FLAC) by URL and the model generates video to match it.
Retired models
Model providers occasionally withdraw a model. When that happens we take it out ofGET /v1/models and its endpoint starts answering 410 Gone with the error code
model_retired, naming the replacement:
POST /v1/cost returns the same error rather than quoting a price you could not spend.
A 410 is permanent — retrying will not help; switch the model id in your request.
Retired so far:
seedance-1.5-pro (2026-09-17, → seedance-2.0).
Parameter schemas differ between models, so check the replacement’s fields in
GET /v1/models before swapping the id — and its per-second rate, which is usually
different too.