Kling V3 Omni T2V
[Core Function] Kling V3 Omni T2V is a multimodal-leaning text-to-video model in the V3 family, oriented toward stronger semantic control and subject consistency in prompt-led generation. [Strengths] It targets high-fidelity cinematic clips with native audio options, flexible 3-15s duration, and better adherence when scenes demand coherent characters or multi-beat storytelling from text alone. [Best For] Highly recommended for: narrative T2V with recurring subjects, dialogue-aware scenes, brand or product continuity across beats, and premium short films where consistency matters more than raw throughput. [Limitations] Do NOT use this model if the user only needs the cheapest or fastest clip; prefer Kling V3 Turbo T2V. Do NOT use it when the workflow is image-first or needs multi-image references; use Kling V3 Omni I2V or Kling V3 I2V instead. Do NOT use it for deep physics-reasoning specialty tasks better served by Kling Video O1. [Routing] Choose Kling V3 Omni T2V when the user emphasizes Omni, consistency, multimodal quality, or complex text narratives. Prefer Kling V3 T2V as the default high-quality T2V baseline; prefer Kling V3 Turbo T2V when the user stresses speed, cost, or high-volume short-form output.
Authorizations
API Key authentication. Format: Bearer YOUR_API_KEY.
Body
Kling V3 text-to-video request aligned with official settings: audio native|off, resolution 720p|1080p|4k, multi_shot, aspect_ratio, duration. No separate negative_prompt.
Video generation prompt. May include both positive and negative descriptions in the same text (no separate negative_prompt field). For custom multi-shot storyboards, use shot n, m, words; shot n, m, words; where n is the shot index (1-6), m is shot duration in seconds (each >= 1s; sum of all shot durations must equal total duration), and words is the per-shot prompt (max 512 characters). Max length 3072 (recommended <= 2500).
1 - 3072"shot 1, 2, A lone cyclist rides through a neon-lit city street after rain; shot 2, 3, Close-up of raindrops on the helmet as neon reflections glide by;"
Whether to generate audio for the video. Official settings.audio: native | off.
native, off "off"
When true, the model automatically plans multi-shot transitions. When false, generates a single continuous shot.
true
Output video resolution.
720p, 1080p, 4k "1080p"
Video aspect ratio.
16:9, 9:16, 1:1 "16:9"
Video duration in seconds. Public API accepts integer values.
3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 6