Six major video generation families — Google Veo, Kling, Sora, LTX, Seedance, and utility models — covering T2V, I2V, lip sync, and upscaling, all priced per second of output.
Video credits scale with output duration × model rate, plus resolution and audio multipliers. Studio displays the estimated cost before you confirm every generation.
MCP video tools accept durationSeconds (typically 2–12 seconds). Studio shows the cost estimate before you run — adjust duration or model tier to control spend.
High-quality text-to-video and image-to-video with built-in audio support.
Model
Notes
Veo 3.1
Flagship; full audio support
Veo 3.1 Fast
Faster generation at the same quality tier
Veo 3.1 Lite
Efficient variant; lower per-second cost
Veo 3
Previous generation; solid quality
Best for: cinematic scenes, complex motion, and productions that need built-in audio synced to video.
Versatile motion and general video generation.
Model
Notes
Kling O3
Latest generation; strong motion control
Kling V3
General-purpose video
Best for: character animation, motion-heavy clips, and mid-tier budget productions.
Three distinct families covering cinematic, efficient, and social-ready outputs.
Model
Provider
Notes
Sora 2
OpenAI
Cinematic quality; strong scene coherence
LTX 2.3
Lightricks
Efficient generation; good for fast iteration
Seedance 2.0
ByteDance
Social-ready motion; fast turnaround
Seedance 2.0 Fast
ByteDance
Fastest variant; lower cost per second
Specialized models for editing, transformation, and lip sync.
Model
Notes
Happy Horse
Distinctive motion style
Gemini Omni Flash
Multimodal video generation tasks
Fabric
Lip sync — matches audio track to existing talking-head video
Utility transformations (inpaint, background removal, upscale, extend) are available for post-processing existing clips and have fixed-ish action costs rather than per-second pricing.
Text-to-video, image-to-video, reference-to-video, and first / last frame anchoring
Motion control
Camera motion, motion intensity controls, and motion reference inputs
Lip sync
Sync dialogue audio to an existing talking-head video (Fabric model)
Post-processing
Inpaint, background removal, upscale, and extend existing clips
Use HyperFrames when you need a multi-stage animated pipeline — compositing multiple clips, adding transitions, and layering effects — rather than a single generated clip.
How pricing works
Full breakdown of per-second rates, duration, and resolution multipliers.
Image models
Browse the image model catalog with per-generation credit costs.