Kling V3 T2V
[Core Function] Kling V3 T2V is the next-generation text-to-video base model. [Strengths] It natively supports generating ultra-long 15-second videos, 4K resolution, and synchronized native audio directly from text. [Best For] Highly recommended for: high-end cinematic creation, 4K video generation, and creating long-form scenes with integrated sound. [Limitations] Do NOT use this model if you need complex multi-shot narratives or deep physics reasoning; use V3 Omni or Video O1 respectively. [Routing] Use this model by default for high-quality text-to-video tasks that require up to 15 seconds, 4K resolution, or native audio without reference images.
Authorizations
API Key authentication. Format: Bearer YOUR_API_KEY.
Body
Kling V3 text-to-video request aligned with official settings: audio native|off, resolution 720p|1080p|4k, multi_shot, aspect_ratio, duration. No separate negative_prompt.
Video generation prompt. May include both positive and negative descriptions in the same text (no separate negative_prompt field). For custom multi-shot storyboards, use shot n, m, words; shot n, m, words; where n is the shot index (1-6), m is shot duration in seconds (each >= 1s; sum of all shot durations must equal total duration), and words is the per-shot prompt (max 512 characters). Max length 3072 (recommended <= 2500).
1 - 3072"shot 1, 2, A lone cyclist rides through a neon-lit city street after rain; shot 2, 3, Close-up of raindrops on the helmet as neon reflections glide by;"
Whether to generate audio for the video. Official settings.audio: native | off.
native, off "off"
When true, the model automatically plans multi-shot transitions. When false, generates a single continuous shot.
true
Output video resolution.
720p, 1080p, 4k "1080p"
Video aspect ratio.
16:9, 9:16, 1:1 "16:9"
Video duration in seconds. Public API accepts integer values.
3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 6