# Modellix ## Docs - [Set Up AI Agents on Modellix](https://docs.modellix.ai/get-started/index.md): Set up an AI agent for Modellix with an API key, the official Skill, or CLI. Submit and retrieve asynchronous image, video, and audio generation tasks. - [Modellix Unified AI Model API Overview](https://docs.modellix.ai/get-started/overview.md): Learn how Modellix provides unified API access to image, video, and speech models, with transparent pricing, async workflows, and browser playgrounds. - [Modellix API Pricing for Image, Video, and Audio](https://docs.modellix.ai/get-started/pricing.md): Understand Modellix API pricing for image, video, and audio generation. Compare per-image, per-second, and per-character costs across models. - [Fund Your Modellix Account](https://docs.modellix.ai/get-started/top-up.md): Learn how to fund your Modellix account with supported payment methods, pay-as-you-go top-ups, and available discounts for your team. - [Modellix Rate Limits and Team Entitlements](https://docs.modellix.ai/get-started/entitlements.md): Understand Modellix concurrency limits, rate limits (RPM), and team entitlements that scale with your single top-up funding tier amount. - [Choose the Right Modellix Image, Video, or Speech Model](https://docs.modellix.ai/get-started/model-select.md): Choose the right Modellix image, video, or speech model by output type, quality, speed, cost, and features such as audio, references, and editing. - [Modellix AI Model Providers and Service URLs](https://docs.modellix.ai/get-started/model-providers.md): Browse every AI model provider on Modellix, including service URLs, typical modalities, and links to the Model API reference for each vendor. - [Use the Modellix REST API for Media Generation](https://docs.modellix.ai/ways-to-use/api.md): Use the Modellix REST API to authenticate, upload media inputs, submit an asynchronous image, video, or audio task, poll its status, and retrieve output assets. - [Install the Modellix Agent Skill](https://docs.modellix.ai/ways-to-use/skill.md): Install the Modellix Agent Skill so coding agents in Cursor, Claude Code, and other hosts can discover models and generate images, videos, and speech. - [Modellix Plugin for Coding Agents](https://docs.modellix.ai/ways-to-use/plugin.md): Install the official Modellix Plugin in Claude Code, Codex, Cursor, and other agent hosts to generate images, videos, and speech via CLI or REST API. - [Use Modellix Agent Canvas in Coding Agents](https://docs.modellix.ai/ways-to-use/agent-canvas.md): Install the Modellix Agent Canvas local stdio MCP plugin in Codex, Cursor, Claude Code, or OpenCode for visual image work, HTML drafts, and presentations. - [Modellix CLI for Image, Video, and Audio Generation](https://docs.modellix.ai/ways-to-use/cli.md): Use modellix-cli to authenticate, discover models, submit async image, video, and audio tasks, wait for results, and download assets from the terminal or CI. - [Search Modellix Docs with the MCP Server](https://docs.modellix.ai/ways-to-use/mcp.md): Connect AI clients like Cursor and Claude Desktop to the Modellix Docs MCP server to search the official documentation without leaving your editor. - [Generate Images, Videos, and Audio in the Playground](https://docs.modellix.ai/ways-to-use/playground.md): Generate images, videos, and speech audio in the Modellix Playground with no code required. Discover models, tune prompts, and preview outputs in your browser. - [HappyHorse 1.0 I2V](https://docs.modellix.ai/alibaba/happyhorse-1-0-i2v.md): [Core Function] HappyHorse 1.0 I2V is a streamlined image-to-video model. [Strengths] It generates high-quality 720P/1080P video (3-15s) from an image efficiently, with native audio support. [Best For] Highly recommended for: rapid image animation and robust character motion… - [HappyHorse 1.1 I2V](https://docs.modellix.ai/alibaba/happyhorse-1-1-i2v.md): [Core Function] HappyHorse 1.1 I2V is Alibaba's latest streamlined first-frame image-to-video model. [Strengths] It turns a single image into high-quality 720P/1080P video with native audio support and 3-15 second duration; output aspect ratio follows the first frame image.… - [HappyHorse 1.0 R2V](https://docs.modellix.ai/alibaba/happyhorse-1-0-r2v.md): [Core Function] HappyHorse 1.0 R2V is a reference-to-video model. [Strengths] It excels at maintaining character consistency using up to 9 reference images while generating new video actions based on a prompt. [Best For] Highly recommended for: character-consistent storytell… - [HappyHorse 1.1 R2V](https://docs.modellix.ai/alibaba/happyhorse-1-1-r2v.md): [Core Function] HappyHorse 1.1 R2V is Alibaba's latest reference-image-to-video model. [Strengths] It uses 1-9 reference images to preserve subject or character appearance while generating new video actions, supports 720P/1080P output, 3-15 second duration, and expanded aspe… - [HappyHorse 1.0 T2V](https://docs.modellix.ai/alibaba/happyhorse-1-0-t2v.md): [Core Function] HappyHorse 1.0 T2V is a breakout, highly optimized text-to-video model. [Strengths] It provides streamlined, fast, and high-quality video generation (up to 15s at 1080p) with native audio support, acting as a highly efficient alternative to Wan 2.7. [Best For… - [HappyHorse 1.1 T2V](https://docs.modellix.ai/alibaba/happyhorse-1-1-t2v.md): [Core Function] HappyHorse 1.1 T2V is Alibaba's latest streamlined text-to-video model. [Strengths] It generates 720P/1080P video with native audio support, 3-15 second duration, and an expanded set of aspect ratios including 4:5, 5:4, 9:21, and 21:9. [Best For] Highly recom… - [HappyHorse 1.0 Video Edit](https://docs.modellix.ai/alibaba/happyhorse-1-0-video-edit.md): [Core Function] HappyHorse 1.0 Video Edit is a streamlined video editing model. [Strengths] It provides high-quality video editing capabilities (with or without reference images) within the highly optimized HappyHorse architecture. [Best For] Highly recommended for: fast, hi… - [Qwen Image](https://docs.modellix.ai/alibaba/qwen-image.md): [Core Function] Qwen Image is an older generation text-to-image model. [Strengths] Maintained primarily for backward compatibility and stable API contracts. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's quirks. [L… - [Qwen Image 2.0](https://docs.modellix.ai/alibaba/qwen-image-2-0.md): [Core Function] Qwen Image 2.0 is a fast, photorealistic image generation model. [Strengths] Balances speed and cost while delivering high-quality 2K generation and standard Qwen 2.0 text rendering. [Best For] Recommended for: cost-effective photorealism and fast concept art… - [Qwen Image 2.0 Edit](https://docs.modellix.ai/alibaba/qwen-image-2-0-edit.md): [Core Function] Qwen Image 2.0 Edit is a fast, highly capable image editing model. [Strengths] Provides strong structural adherence and editing speed. [Best For] Recommended for: cost-effective image modifications. [Limitations] Lacks the finer detail rendering of the Pro ve… - [Qwen Image 2.0 Pro](https://docs.modellix.ai/alibaba/qwen-image-2-0-pro.md): [Core Function] Qwen Image 2.0 Pro is a professional-grade, highly controllable image generation model. [Strengths] It provides exquisite photorealism, professional infographics generation, supports NEGATIVE prompts, and can generate up to 6 image variants per API call. [Bes… - [Qwen Image 2.0 Pro Edit](https://docs.modellix.ai/alibaba/qwen-image-2-0-pro-edit.md): [Core Function] Qwen Image 2.0 Pro Edit is a highly controllable professional image editing model. [Strengths] It seamlessly unifies generation and editing, supporting NEGATIVE prompts during the editing process to strictly exclude elements. [Best For] Highly recommended for… - [Qwen Image 3.0](https://docs.modellix.ai/alibaba/qwen-image-3-0.md): [Core Function] Qwen Image 3.0 is Alibaba's standard text-to-image model balancing quality and speed. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite (direct/agent modes), batch generation of 1-6 images, and… - [Qwen Image 3.0 Edit](https://docs.modellix.ai/alibaba/qwen-image-3-0-edit.md): [Core Function] Qwen Image 3.0 Edit is the standard image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (width*height), intelli… - [Qwen Image 3.0 Pro](https://docs.modellix.ai/alibaba/qwen-image-3-0-pro.md): [Core Function] Qwen Image 3.0 Pro is Alibaba's latest text-to-image model with strong prompt following and photorealism. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite, batch generation of 1-6 images, and… - [Qwen Image 3.0 Pro Edit](https://docs.modellix.ai/alibaba/qwen-image-3-0-pro-edit.md): [Core Function] Qwen Image 3.0 Pro Edit is an image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (width*height), intelligent p… - [Qwen Image Edit](https://docs.modellix.ai/alibaba/qwen-image-edit.md): [Core Function] Qwen Image Edit is an older generation image-to-image editing model. [Strengths] Maintained primarily for backward compatibility and stable API contracts. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model versio… - [Qwen Image Edit Max](https://docs.modellix.ai/alibaba/qwen-image-edit-max.md): [Core Function] Qwen Image Edit Max is an older generation image-to-image editing model. [Strengths] Historically offered higher resolution and better prompt adherence for complex requests. Maintained for backward compatibility. [Best For] Existing legacy integrations and wo… - [Qwen Image Edit Plus](https://docs.modellix.ai/alibaba/qwen-image-edit-plus.md): [Core Function] Qwen Image Edit Plus is an older generation image-to-image editing model. [Strengths] Historically offered higher resolution and better prompt adherence for complex requests. Maintained for backward compatibility. [Best For] Existing legacy integrations and w… - [Qwen Image Edit Plus 2025-10-30](https://docs.modellix.ai/alibaba/qwen-image-edit-plus-2025-10-30.md): [Core Function] Qwen Image Edit Plus 2025-10-30 is an older generation image-to-image editing model. [Strengths] Historically offered higher resolution and better prompt adherence for complex requests. Maintained for backward compatibility. [Best For] Existing legacy integra… - [Qwen Image Edit Plus 2025-12-15](https://docs.modellix.ai/alibaba/qwen-image-edit-plus-2025-12-15.md): [Core Function] Qwen Image Edit Plus 2025-12-15 is an older generation image-to-image editing model. [Strengths] Historically offered higher resolution and better prompt adherence for complex requests. Maintained for backward compatibility. [Best For] Existing legacy integra… - [Qwen Image Max](https://docs.modellix.ai/alibaba/qwen-image-max.md): [Core Function] Qwen Image Max is an older generation text-to-image model. [Strengths] Historically offered higher resolution and better prompt adherence for complex requests. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that s… - [Qwen Image Plus](https://docs.modellix.ai/alibaba/qwen-image-plus.md): [Core Function] Qwen Image Plus is an older generation text-to-image model. [Strengths] Historically offered higher resolution and better prompt adherence for complex requests. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that… - [Wan 2.1 VACE Plus](https://docs.modellix.ai/alibaba/wan-2-1-vace-plus.md): [Core Function] Wan 2.1 VACE Plus - Unified Video Editing Model is an older generation image-to-image editing model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Exist… - [Wan 2.2 Animate Mix](https://docs.modellix.ai/alibaba/wan-2-2-animate-mix.md): [Core Function] Wan 2.2 Animate Mix - Character Replacement in Video is a legacy video editing and motion transfer model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integ… - [Wan 2.2 Animate Move](https://docs.modellix.ai/alibaba/wan-2-2-animate-move.md): [Core Function] Wan 2.2 Animate Move - Motion Transfer is a legacy video editing and motion transfer model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and wo… - [Wan 2.2 I2V Flash](https://docs.modellix.ai/alibaba/wan-2-2-i2v-flash.md): [Core Function] Wan 2.2 I2V Flash is an older generation image-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cos… - [Wan 2.2 I2V Plus](https://docs.modellix.ai/alibaba/wan-2-2-i2v-plus.md): [Core Function] Wan 2.2 I2V Plus is an older generation image-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend o… - [Wan 2.2 KF2V Flash](https://docs.modellix.ai/alibaba/wan-2-2-kf2v-flash.md): [Core Function] Wan 2.2 KF2V Flash is a legacy generation model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of… - [Wan 2.2 T2I Flash](https://docs.modellix.ai/alibaba/wan-2-2-t2i-flash.md): [Core Function] Wan 2.2 T2I Flash is an older generation text-to-image model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost… - [Wan 2.2 T2I Plus](https://docs.modellix.ai/alibaba/wan-2-2-t2i-plus.md): [Core Function] Wan 2.2 T2I Plus is an older generation text-to-image model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on… - [Wan 2.2 T2V Plus](https://docs.modellix.ai/alibaba/wan-2-2-t2v-plus.md): [Core Function] Wan 2.2 T2V Plus is an older generation text-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on… - [Wan 2.5 I2I Preview](https://docs.modellix.ai/alibaba/wan-2-5-i2i-preview.md): [Core Function] Wan 2.5 I2I Preview is an older generation image-to-image editing model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strict… - [Wan 2.5 I2V Preview](https://docs.modellix.ai/alibaba/wanx-2-5-i2v-preview.md): [Core Function] Wanx 2.5 I2V Preview is an older generation image-to-video model. [Strengths] Maintained primarily for backward compatibility and stable API contracts. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model version's… - [Wan 2.5 T2I Preview](https://docs.modellix.ai/alibaba/wan-2-5-t2i-preview.md): [Core Function] Wan 2.5 T2I Preview is an older generation text-to-image model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend… - [Wan 2.5 T2V Preview](https://docs.modellix.ai/alibaba/wan-2-5-t2v-preview.md): [Core Function] Wan 2.5 T2V Preview is an older generation text-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend… - [Wan 2.6 I2V](https://docs.modellix.ai/alibaba/wan-2-6-i2v.md): [Core Function] Wan 2.6 I2V is an older generation image-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on thi… - [Wan 2.6 I2V Flash](https://docs.modellix.ai/alibaba/wan-2-6-i2v-flash.md): [Core Function] Wan 2.6 I2V Flash is an older generation image-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cos… - [Wan 2.6 Image](https://docs.modellix.ai/alibaba/wan-2-6-image.md): [Core Function] Wan 2.6 Image is an older generation text-to-image model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on th… - [Wan 2.6 R2V](https://docs.modellix.ai/alibaba/wan-2-6-r2v.md): [Core Function] Wan 2.6 Reference-to-Video is a legacy generation model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on thi… - [Wan 2.6 R2V Flash](https://docs.modellix.ai/alibaba/wan-2-6-r2v-flash.md): [Core Function] Wan 2.6 Reference-to-Video Flash is a legacy generation model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cos… - [Wan 2.6 T2I](https://docs.modellix.ai/alibaba/wan-2-6-t2i.md): [Core Function] Wan 2.6 T2I is an older generation text-to-image model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this… - [Wan 2.6 T2V](https://docs.modellix.ai/alibaba/wan-2-6-t2v.md): [Core Function] Wan 2.6 T2V is an older generation text-to-video model. [Strengths] Provided enhanced dynamic camera movements and rich lighting effects. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly depend on this… - [Wan 2.7 I2V](https://docs.modellix.ai/alibaba/wan-2-7-i2v.md): [Core Function] Wan 2.7 I2V is Alibaba's flagship multimodal image-to-video model. [Strengths] It supports multimodal input (text, image, audio, video) for first-frame, start-and-end-frame (FL2V), and video continuation tasks. [Best For] Highly recommended for: complex image… - [Wan 2.7 Image](https://docs.modellix.ai/alibaba/wan-2-7-image.md): [Core Function] Wan 2.7 Image is a fast, reasoning-enhanced image generation model. [Strengths] It includes the chain-of-thought reasoning and text rendering of the Pro version, but is optimized for speed, supporting up to 2K resolution. [Best For] Highly recommended for: fa… - [Wan 2.7 Image Edit](https://docs.modellix.ai/alibaba/wan-2-7-image-edit.md): [Core Function] Wan 2.7 Image Edit is a fast, reasoning-enhanced image editing model. [Strengths] Provides the robust editing capabilities of the Wan 2.7 architecture with faster turnaround times. [Best For] Highly recommended for: standard image modifications and style tran… - [Wan 2.7 Image Pro](https://docs.modellix.ai/alibaba/wan-2-7-image-pro.md): [Core Function] Wan 2.7 Image Pro is Alibaba's flagship reasoning-enhanced image generation model. [Strengths] It features built-in chain-of-thought reasoning (Thinking Mode), exceptional prompt accuracy, native 12-language text rendering, and generates ultra-high-resolution… - [Wan 2.7 Image Pro Edit](https://docs.modellix.ai/alibaba/wan-2-7-image-pro-edit.md): [Core Function] Wan 2.7 Image Pro Edit is Alibaba's flagship reasoning-enhanced image editing model. [Strengths] It supports interactive editing, character-consistent multi-image generation, and complex multi-reference modifications with deep reasoning. [Best For] Highly rec… - [Wan 2.7 R2V](https://docs.modellix.ai/alibaba/wan-2-7-r2v.md): [Core Function] Wan 2.7 Reference-to-Video is a highly capable character/entity reference video model. [Strengths] It natively supports entity reference, voice customization, and playbook-based video generation from a single storyboard. [Best For] Highly recommended for: cre… - [Wan 2.7 T2V](https://docs.modellix.ai/alibaba/wan-2-7-t2v.md): [Core Function] Wan 2.7 T2V is Alibaba's flagship text-to-video generation model. [Strengths] It generates high-fidelity video directly from text with support for custom aspect ratios, audio generation, and intricate semantic adherence. [Best For] Highly recommended for: hig… - [Wan 2.7 Videoedit](https://docs.modellix.ai/alibaba/wan-2-7-videoedit.md): [Core Function] Wan 2.7 Video Editing is an instruction-based video modification model. [Strengths] It supports complex video editing tasks like content replacement using reference images, and replicating actions, effects, and camera movements. [Best For] Highly recommended… - [Wan 3.0 I2V](https://docs.modellix.ai/alibaba/wan3-0-i2v.md): [Core Function] Wan 3.0 I2V is Alibaba Wan 3.0 image-to-video generation supporting first-frame, first-last-frame, and reference-image modes. [Strengths] It can strictly lock the first and last frames or fuse up to 10 reference images with optional reference audio for multim… - [Wan 3.0 T2V](https://docs.modellix.ai/alibaba/wan3-0-t2v.md): [Core Function] Wan 3.0 T2V is Alibaba Wan 3.0 text-to-video generation with optional document or webpage reference. [Strengths] It generates up to 30-second video at 480P/720P/1080P with controllable aspect ratio and optional output audio, and can ground generation on a fil… - [Wan 3.0 V2V](https://docs.modellix.ai/alibaba/wan3-0-v2v.md): [Core Function] Wan 3.0 V2V is Alibaba Wan 3.0 reference-video generation that builds new video from one or more input videos. [Strengths] It supports up to 5 reference videos with optional reference images and audio for multimodal composition and prompt-referenced subjects.… - [Wanx 2.1 I2V Plus](https://docs.modellix.ai/alibaba/wanx-2-1-i2v-plus.md): [Core Function] Wanx 2.1 I2V Plus is an older generation image-to-video model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows… - [Wanx 2.1 I2V Turbo](https://docs.modellix.ai/alibaba/wanx-2-1-i2v-turbo.md): [Core Function] Wanx 2.1 I2V Turbo is an older generation image-to-video model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations that require… - [Wanx 2.1 KF2V Plus](https://docs.modellix.ai/alibaba/wanx-2-1-kf2v-plus.md): [Core Function] Wanx 2.1 KF2V Plus is a legacy generation model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly… - [Wanx 2.1 T2I Plus](https://docs.modellix.ai/alibaba/wanx-2-1-t2i-plus.md): [Core Function] Wanx 2.1 T2I Plus is an older generation text-to-image model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows t… - [Wanx 2.1 T2I Turbo](https://docs.modellix.ai/alibaba/wanx-2-1-t2i-turbo.md): [Core Function] Wanx 2.1 T2I Turbo is an older generation text-to-image model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations that require t… - [Wanx 2.1 T2V Plus](https://docs.modellix.ai/alibaba/wanx-2-1-t2v-plus.md): [Core Function] Wanx 2.1 T2V Plus is an older generation text-to-video model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows t… - [Wanx 2.1 T2V Turbo](https://docs.modellix.ai/alibaba/wanx-2-1-t2v-turbo.md): [Core Function] Wanx 2.1 T2V Turbo is an older generation text-to-video model. [Strengths] Delivered strong Chinese-language prompt understanding and regional aesthetic preferences. Maintained for backward compatibility. [Best For] Existing legacy integrations that require t… - [Z Image Turbo](https://docs.modellix.ai/alibaba/z-image-turbo.md): [Core Function] Z-Image Turbo is an older generation text-to-image model. [Strengths] Historically provided faster generation times and lower latency compared to its standard counterparts. Maintained for backward compatibility. [Best For] Existing legacy integrations that re… - [CosyVoice V3 Plus](https://docs.modellix.ai/alibaba/cosyvoice-v3-plus.md): [Core Function] CosyVoice v3 Plus is Alibaba's high-quality text-to-speech model. [Strengths] System voices (e.g. longanyang, longanhuan), SSML and LaTeX input, hot_fix pronunciation correction, AIGC watermark, and output in mp3, pcm, wav, or opus. System-voice instruction m… - [CosyVoice V3 Flash](https://docs.modellix.ai/alibaba/cosyvoice-v3-flash.md): [Core Function] CosyVoice v3 Flash is Alibaba's low-latency text-to-speech model. [Strengths] Rich system voice catalog, fixed-format instruction on Instruct-capable system voices, SSML, hot_fix, AIGC watermark, Markdown filter (cloned voices only), and multiple audio format… - [Qwen Audio 3.0 TTS Plus](https://docs.modellix.ai/alibaba/qwen-audio-3-0-tts-plus.md): [Core Function] Qwen-Audio 3.0 TTS Plus is Alibaba's high-quality Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Natural speech synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and… - [Qwen Audio 3.0 TTS Flash](https://docs.modellix.ai/alibaba/qwen-audio-3-0-tts-flash.md): [Core Function] Qwen-Audio 3.0 TTS Flash is Alibaba's low-latency Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Fast synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC water… - [CosyVoice Clone](https://docs.modellix.ai/alibaba/cosyvoice-clone.md): [Core Function] CosyVoice Clone clones a speaker from a public reference audio URL and synthesizes new speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; language_hint applies to both cloning and synthesis; SSML, hot_fi… - [CosyVoice Design](https://docs.modellix.ai/alibaba/cosyvoice-design.md): [Core Function] CosyVoice Design creates a temporary voice from a natural-language voice_prompt and synthesizes speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; language_hint (zh/en) applies to both design and synthes… - [Fun ASR](https://docs.modellix.ai/alibaba/fun-asr.md): [Core Function] Fun-ASR transcribes a single public audio file asynchronously. [Strengths] Hot-word vocabulary, optional speaker diarization, channel selection, and language hints. [Best For] Batch transcription of recordings up to 12 hours. [Limitations] Do NOT send more th… - [Fun ASR MTL](https://docs.modellix.ai/alibaba/fun-asr-mtl.md): [Core Function] Fun-ASR MTL is the multi-language variant for async recorded speech recognition. [Strengths] Same parameters as fun-asr with multi-language tuning. [Best For] Mixed-language or international audio archives. [Limitations] Do NOT send more than one file per req… - [Seedance 1.0 Pro Fast I2V](https://docs.modellix.ai/bytedance/seedance-1-0-pro-fast-i2v.md): [Core Function] Seedance 1.0 Pro Fast I2V is an older generation image-to-video model. [Strengths] Known for its rapid generation pipeline and robust performance on standard commercial prompts. Maintained for backward compatibility. [Best For] Existing legacy integrations th… - [Seedance 1.0 Pro Fast T2V](https://docs.modellix.ai/bytedance/seedance-1-0-pro-fast-t2v.md): [Core Function] Seedance 1.0 Pro Fast T2V is an older generation text-to-video model. [Strengths] Known for its rapid generation pipeline and robust performance on standard commercial prompts. Maintained for backward compatibility. [Best For] Existing legacy integrations tha… - [Seedance 1.0 Pro I2V](https://docs.modellix.ai/bytedance/seedance-1-0-pro-i2v.md): [Core Function] Seedance 1.0 Pro I2V is an older generation image-to-video model. [Strengths] Known for its rapid generation pipeline and robust performance on standard commercial prompts. Maintained for backward compatibility. [Best For] Existing legacy integrations and wor… - [Seedance 1.0 Pro T2V](https://docs.modellix.ai/bytedance/seedance-1-0-pro-t2v.md): [Core Function] Seedance 1.0 Pro T2V is an older generation text-to-video model. [Strengths] Known for its rapid generation pipeline and robust performance on standard commercial prompts. Maintained for backward compatibility. [Best For] Existing legacy integrations and work… - [Seedance 1.5 Pro I2V](https://docs.modellix.ai/bytedance/seedance-1-5-pro-i2v.md): [Core Function] Seedance 1.5 Pro I2V is a joint audio-video image-to-video model. [Strengths] It accurately follows complex instructions to animate a single image with synchronized audio. [Best For] Recommended for: standard image animation tasks where 2.0's multi-reference… - [Seedance 1.5 Pro T2V](https://docs.modellix.ai/bytedance/seedance-1-5-pro-t2v.md): [Core Function] Seedance 1.5 Pro T2V is a joint audio-video text-to-video generation model. [Strengths] It accurately follows complex text instructions to generate high-quality video with synchronized audio. [Best For] Highly recommended for: standard text-to-video generatio… - [Seedance 2.0 Fast T2V](https://docs.modellix.ai/bytedance/seedance-2-0-fast-t2v.md): [Core Function] Seedance 2.0 Fast T2V is the faster text-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output. [Routing] Use when speed is preferred over maximum resolution. - [Seedance 2.0 Fast I2V](https://docs.modellix.ai/bytedance/seedance-2-0-fast-i2v.md): [Core Function] Seedance 2.0 Fast I2V is a high-speed multimodal video generation model. [Strengths] Fast generation with the multimodal and multi-shot capabilities of the Seedance 2.0 architecture. [Best For] Highly recommended for: rapid prototyping and quick multi-shot vi… - [Seedance 2.0 Fast V2V](https://docs.modellix.ai/bytedance/seedance-2-0-fast-v2v.md): [Core Function] Seedance 2.0 Fast V2V is a high-speed video-to-video generation model. [Strengths] It offers rapid video transformation capabilities based on the 2.0 architecture. [Best For] Highly recommended for: quick video restyling and fast iterations. [Limitations] Do… - [Seedance 2.0 Mini T2V](https://docs.modellix.ai/bytedance/seedance-2-0-mini-t2v.md): [Core Function] Seedance 2.0 Mini T2V is the lightweight text-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output. [Routing] Use for cost-efficient text-to-video. - [Seedance 2.0 Mini I2V](https://docs.modellix.ai/bytedance/seedance-2-0-mini-i2v.md): [Core Function] Seedance 2.0 Mini I2V is the lightweight multimodal image-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with optional first/last frame, reference images, and audio references. [Routing] Use for cost-efficient image-to-video. - [Seedance 2.0 Mini V2V](https://docs.modellix.ai/bytedance/seedance-2-0-mini-v2v.md): [Core Function] Seedance 2.0 Mini V2V is the lightweight video-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with required reference video and optional text/image/audio references. [Routing] Use for cost-efficient video-to-video transformations. - [Seedance 2.0 T2V](https://docs.modellix.ai/bytedance/seedance-2-0-t2v.md): [Core Function] Seedance 2.0 T2V is ByteDance Dreamina Seedance 2.0 text-to-video. [Strengths] Supports 480p/720p/1080p/4k, 24 fps, 4-15s MP4 output. Text-only input — do not pass images, video, or audio. [Routing] Use for high-fidelity text-to-video when quality or 4k outpu… - [Seedance 2.0 I2V](https://docs.modellix.ai/bytedance/seedance-2-0-i2v.md): [Core Function] Seedance 2.0 I2V is ByteDance's flagship unified multimodal video generation model. [Strengths] It supports complex mixed references (multiple images, audio clips) and generates up to 15s of multi-shot audio-video output with dual-channel audio. [Best For] Hi… - [Seedance 2.0 V2V](https://docs.modellix.ai/bytedance/seedance-2-0-v2v.md): [Core Function] Seedance 2.0 V2V is ByteDance's flagship multimodal video-to-video model. [Strengths] It allows powerful editing and stylization of input videos by supporting mixed references (text, images, video, and audio) and producing multi-shot 15s outputs. [Best For] H… - [Seedance 2.5 T2V](https://docs.modellix.ai/bytedance/seedance-2-5-t2v.md): [Core Function] Seedance 2.5 T2V is ByteDance Dreamina Seedance 2.5 text-to-video generation. [Strengths] It generates longer clips up to 30 seconds at 480p/720p with optional mp4 or mov output and native audio generation. [Best For] Highly recommended for: longer-form text-… - [Seedance 2.5 I2V](https://docs.modellix.ai/bytedance/seedance-2-5-i2v.md): [Core Function] Seedance 2.5 I2V is ByteDance Dreamina Seedance 2.5 image-to-video generation supporting first-frame, first-and-last-frame, and reference-image modes. [Strengths] It supports up to 30 reference images, optional reference audio, 4-30 second duration, and mp4 o… - [Seedance 2.5 V2V](https://docs.modellix.ai/bytedance/seedance-2-5-v2v.md): [Core Function] Seedance 2.5 V2V is ByteDance Dreamina Seedance 2.5 video-to-video generation covering multimodal reference, video editing, and video extension. [Strengths] It accepts up to 10 reference videos, 30 reference images, and 10 audio clips, with 4-30 second output… - [Seedream 4.0 I2I](https://docs.modellix.ai/bytedance/seedream-4-0-i2i.md): [Core Function] Seedream 4.0 I2I is an older generation image-to-image editing model. [Strengths] Maintained primarily for backward compatibility and stable API contracts. [Best For] Existing legacy integrations and workflows that strictly depend on this specific model versi… - [Seedream 4.0 T2I](https://docs.modellix.ai/bytedance/seedream-4-0-t2i.md): [Core Function] Seedream 4.0 T2I is an older high-definition image creation model. [Strengths] It generates high-definition images up to 4K. [Best For] Existing workflows. [Limitations] Superseded by Seedream 4.5 in fidelity and consistency. [Routing] Default to Seedream 4.5… - [Seedream 4.5 I2I](https://docs.modellix.ai/bytedance/seedream-4-5-i2i.md): [Core Function] Seedream 4.5 I2I is a high-fidelity image editing model. [Strengths] It provides high-consistency, high-resolution style transfer and image-to-image transformations. [Best For] Highly recommended for: professional aesthetic modifications and high-resolution e… - [Seedream 4.5 T2I](https://docs.modellix.ai/bytedance/seedream-4-5-t2i.md): [Core Function] Seedream 4.5 T2I is a high-fidelity visual creation model. [Strengths] It provides all-round improvements in consistency, aesthetics, and photorealism, supporting high-definition outputs up to 4K resolution. [Best For] Highly recommended for: professional vis… - [Seedream 5.0 Lite](https://docs.modellix.ai/bytedance/seedream-5-0-lite.md): [Core Function] Seedream 5.0 Lite is a smart, reasoning-enhanced image generation model with real-time web search capabilities. [Strengths] It excels at generating time-sensitive imagery, infographics, and content requiring deep world knowledge or online search, boasting sup… - [Seedream 5.0 Lite Edit](https://docs.modellix.ai/bytedance/seedream-5-0-lite-edit.md): [Core Function] Seedream 5.0 Lite Edit is a reasoning-enhanced, smart image editing model. [Strengths] It features superior cross-modal understanding and reasoning, allowing for highly accurate, interactive multi-turn image editing with real-time knowledge enhancement. [Best… - [Seedream 5.0 Pro](https://docs.modellix.ai/bytedance/seedream-5-0-pro.md): [Core Function] Seedream 5.0 Pro is ByteDance's flagship professional-grade Text-to-Image (T2I) generation model. [Strengths] It delivers top-tier image quality with enhanced precision control over positions and elements, superior prompt adherence, and improved generation co… - [Seedream 5.0 Pro Edit](https://docs.modellix.ai/bytedance/seedream-5-0-pro-edit.md): [Core Function] Seedream 5.0 Pro Edit is a professional-grade single-image editing (I2I) model. [Strengths] It supports interactive precise editing: edit locations can be specified via coordinates, selection boxes, or arrows described in the prompt, with strong element-level… - [Seedream 5.0 Pro Multi Reference](https://docs.modellix.ai/bytedance/seedream-5-0-pro-multi-reference.md): [Core Function] Seedream 5.0 Pro Multi-Reference is a professional-grade multi-reference image generation (I2I) model that creates a single image from 2-10 reference images plus a text prompt. [Strengths] It excels at reference consistency, preserving characters, styles, and… - [Gemini 3.1 Flash TTS](https://docs.modellix.ai/google/gemini-3-1-flash-tts.md): [Core Function] Gemini 3.1 Flash TTS is Google's low-latency, controllable text-to-speech model. [Strengths] Single-speaker and two-speaker dialogue, 30 prebuilt voices, 70+ languages via language_code, and expressive delivery through style prompts plus inline audio tags suc… - [Gemini Omni Flash I2V](https://docs.modellix.ai/google/gemini-omni-flash-i2v.md): [Core Function] Gemini Omni Flash I2V is a fast Image-to-Video model that animates a single input image into a short 720p video via the Interactions API. [Strengths] It uses the provided image as the opening frame and generates smooth motion with natively synchronized audio… - [Gemini Omni Flash R2V](https://docs.modellix.ai/google/gemini-omni-flash-r2v.md): [Core Function] Gemini Omni Flash R2V (Reference-to-Video) generates a short 720p video guided by up to three reference images via the Interactions API. [Strengths] It fuses the styles, subjects, or elements from multiple reference images (referred to in the text prompt) int… - [Gemini Omni Flash T2V](https://docs.modellix.ai/google/gemini-omni-flash-t2v.md): [Core Function] Gemini Omni Flash T2V is Google's fast multimodal Text-to-Video generation model built on the Interactions API. [Strengths] It quickly turns a text prompt into a short 720p video with natively synchronized audio, offering low latency and solid prompt adherenc… - [Gemini Omni Flash Video Edit](https://docs.modellix.ai/google/gemini-omni-flash-video-edit.md): [Core Function] Gemini Omni Flash Video Edit performs conversational, instruction-driven editing of an existing video via the Interactions API. [Strengths] It applies natural-language edits (changing the scene, mood, style, lighting, background, or time of day) to an input v… - [Nano Banana](https://docs.modellix.ai/google/nano-banana.md): [Core Function] Nano Banana is the original fast creative image model. [Strengths] Very fast creative generation. [Best For] Quick sketches and ideas. [Limitations] Superseded by Nano Banana 2 for general speed tasks. [Routing] Default to Nano Banana 2 unless specifically re… - [Nano Banana 2](https://docs.modellix.ai/google/nano-banana-2.md): [Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creat… - [Nano Banana 2 Lite](https://docs.modellix.ai/google/nano-banana-2-lite.md): [Core Function] Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is the lightweight, cost-efficient text-to-image model of the Nano Banana 2 family. [Strengths] It generates images even faster and more cheaply than Nano Banana 2, well suited to high-volume creative and styli… - [Nano Banana 2 Edit](https://docs.modellix.ai/google/nano-banana-2-edit.md): [Core Function] Nano Banana 2 Edit is a high-speed image editing model. [Strengths] It rapidly modifies existing images or extracts image frames from videos based on text prompts. [Best For] Highly recommended for: rapid style transfer, quick image modifications, and fast cr… - [Nano Banana 2 Lite Edit](https://docs.modellix.ai/google/nano-banana-2-lite-edit.md): [Core Function] Nano Banana 2 Lite Edit (Gemini 3.1 Flash Lite Image) is the lightweight, cost-efficient image editing model of the Nano Banana 2 family; it transforms one or more input images per a text instruction. [Strengths] It performs fast, low-cost instruction-based e… - [Nano Banana Edit](https://docs.modellix.ai/google/nano-banana-edit.md): [Core Function] Nano Banana Edit is the original fast image editing model. [Strengths] Fast basic edits. [Limitations] Superseded by Nano Banana 2 Edit. [Routing] Default to Nano Banana 2 Edit unless specifically requested. - [Nano Banana Pro](https://docs.modellix.ai/google/nano-banana-pro.md): [Core Function] Nano Banana Pro (Gemini 3 Pro Image) is a high-capability creative image model. [Strengths] It balances the creative flexibility and speed of the Nano series with higher fidelity output. [Best For] Highly recommended for: high-quality stylized art and complex… - [Nano Banana Pro Edit](https://docs.modellix.ai/google/nano-banana-pro-edit.md): [Core Function] Nano Banana Pro Edit is a high-capability creative image editing model. [Strengths] It provides high-quality creative edits, background replacements, and style transformations. [Best For] Highly recommended for: detailed creative modifications and complex sty… - [Veo 2 I2V](https://docs.modellix.ai/google/veo-2-i2v.md): [Core Function] Veo 2 I2V is an older generation image-to-video model. [Strengths] Known for its distinctive cinematic style and fluid motion priors from the Veo 2 era. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly… - [Veo 2 T2V](https://docs.modellix.ai/google/veo-2-t2v.md): [Core Function] Veo 2 T2V is an older generation text-to-video model. [Strengths] Known for its distinctive cinematic style and fluid motion priors from the Veo 2 era. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that strictly… - [Veo 3 Fast I2V](https://docs.modellix.ai/google/veo-3-fast-i2v.md): [Core Function] Veo 3 Fast I2V is an older generation image-to-video model. [Strengths] Provided strong physical consistency and 1080p generation capabilities before the 3.1 update. Maintained for backward compatibility. [Best For] Existing legacy integrations that require t… - [Veo 3 Fast T2V](https://docs.modellix.ai/google/veo-3-fast-t2v.md): [Core Function] Veo 3 Fast T2V is an older generation text-to-video model. [Strengths] Provided strong physical consistency and 1080p generation capabilities before the 3.1 update. Maintained for backward compatibility. [Best For] Existing legacy integrations that require th… - [Veo 3 I2V](https://docs.modellix.ai/google/veo-3-i2v.md): [Core Function] Veo 3 I2V is an older generation image-to-video model. [Strengths] Provided strong physical consistency and 1080p generation capabilities before the 3.1 update. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that… - [Veo 3 T2V](https://docs.modellix.ai/google/veo-3-t2v.md): [Core Function] Veo 3 T2V is an older generation text-to-video model. [Strengths] Provided strong physical consistency and 1080p generation capabilities before the 3.1 update. Maintained for backward compatibility. [Best For] Existing legacy integrations and workflows that s… - [Veo 3.1 Fast I2V](https://docs.modellix.ai/google/veo-3-1-fast-i2v.md): [Core Function] Veo 3.1 Fast I2V is a high-speed image-to-video model. [Strengths] It quickly animates starting images at 1080p, optimized for low latency. [Best For] Highly recommended for: rapid prototyping and quick social media visual iterations. [Limitations] Do NOT use… - [Veo 3.1 Fast T2V](https://docs.modellix.ai/google/veo-3-1-fast-t2v.md): [Core Function] Veo 3.1 Fast T2V is a high-speed text-to-video model. [Strengths] It is heavily optimized for fast generation, delivering video content with natively synchronized audio at 1080p quickly. [Best For] Highly recommended for: rapid prototyping, quick visual itera… - [Veo 3.1 I2V](https://docs.modellix.ai/google/veo-3-1-i2v.md): [Core Function] Veo 3.1 I2V is Google's cinematic image-to-video generation model. [Strengths] It generates high-fidelity 4K video from a starting image. It supports advanced features like first-and-last frame conditioning and referencing up to three images. [Best For] Highl… - [Veo 3.1 Lite I2V](https://docs.modellix.ai/google/veo-3-1-lite-i2v.md): [Core Function] Veo 3.1 Lite I2V is a balanced image-to-video model. [Strengths] It offers a middle ground between speed and quality for animating images. [Best For] Highly recommended for: general image animation and web-ready content. [Limitations] Do NOT use this model if… - [Veo 3.1 Lite T2V](https://docs.modellix.ai/google/veo-3-1-lite-t2v.md): [Core Function] Veo 3.1 Lite T2V is a balanced text-to-video model. [Strengths] It provides a good balance between generation speed and visual quality, still supporting the advanced architecture of the 3.1 series. [Best For] Highly recommended for: general video content crea… - [Veo 3.1 T2V](https://docs.modellix.ai/google/veo-3-1-t2v.md): [Core Function] Veo 3.1 T2V is Google's state-of-the-art cinematic text-to-video engine. [Strengths] It natively generates 4K professional-grade video output with natively synchronized audio and supports complex camera movements. [Best For] Highly recommended for: high-end c… - [Kling Video O1](https://docs.modellix.ai/kling/kling-video-o1.md): [Core Function] Kling Video O1 is the world's first reasoning-enhanced video model. [Strengths] It performs deep planning over the prompt before generation, delivering best-in-class physical consistency, complex motion logic, and strict adherence to long-form semantics. [Bes… - [Kling Video O1 T2V](https://docs.modellix.ai/kling/kling-video-o1-t2v.md): [Core Function] Kling Video O1 T2V is the text-to-video slice of Kling O1 Omni Video. [Strengths] Reasoning-enhanced prompt planning with 3-10s duration and 720p/1080p output. [Best For] Complex physical interactions and logically demanding scenes from text alone. [Limitatio… - [Kling Video O1 I2V](https://docs.modellix.ai/kling/kling-video-o1-i2v.md): [Core Function] Kling Video O1 I2V is the image-to-video slice of Kling O1 Omni Video. [Strengths] Reasoning-enhanced generation from 1-7 reference images, 720p/1080p, duration 3-10s (single image only 5 or 10). [Best For] Complex physical motion grounded in reference frames… - [Kling V3 Omni T2V](https://docs.modellix.ai/kling/kling-v3-omni-t2v.md): [Core Function] Kling V3 Omni T2V is a multimodal-leaning text-to-video model in the V3 family, oriented toward stronger semantic control and subject consistency in prompt-led generation. [Strengths] It targets high-fidelity cinematic clips with native audio options, flexibl… - [Kling V3 Omni I2V](https://docs.modellix.ai/kling/kling-v3-omni-i2v.md): [Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and pro… - [Kling V3 Omni Video](https://docs.modellix.ai/kling/kling-v3-omni-video.md): [Core Function] Kling V3 Omni Video V2V is a multimodal video-to-video endpoint that edits or restyles existing footage using prompt plus optional image and video references. [Strengths] It focuses on source fidelity and subject consistency for Omni-style edit workflows, com… - [Kling V3 Turbo T2V](https://docs.modellix.ai/kling/kling-v3-turbo-t2v.md): [Core Function] Kling V3 Turbo T2V is a speed- and cost-optimized text-to-video model in the V3 family for fast short-form generation. [Strengths] It emphasizes lower latency and efficient throughput with native audio and improved lip-sync for talking-head style clips, typic… - [Kling V3 Turbo I2V](https://docs.modellix.ai/kling/kling-v3-turbo-i2v.md): [Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or… - [Kling V3 T2V](https://docs.modellix.ai/kling/kling-v3-t2v.md): [Core Function] Kling V3 T2V is the next-generation text-to-video base model. [Strengths] It natively supports generating ultra-long 15-second videos, 4K resolution, and synchronized native audio directly from text. [Best For] Highly recommended for: high-end cinematic creat… - [Kling V3 I2V](https://docs.modellix.ai/kling/kling-v3-i2v.md): [Core Function] Kling V3 I2V is the next-generation image-to-video model. [Strengths] It transforms static images into video with support for 4K resolution, 15-second durations, and native audio, providing superior motion and character expressiveness. [Best For] Highly recom… - [Kling V3 T2I](https://docs.modellix.ai/kling/kling-v3-t2i.md): [Core Function] Kling V3 T2I is the flagship text-to-image model (POST /images/generations, model_name=kling-v3). [Strengths] High aesthetic quality, prompt adherence, 1K/2K. [Best For] Concept art and photorealistic generation without a reference image. [Limitations] No ref… - [Kling V3 I2I](https://docs.modellix.ai/kling/kling-v3-i2i.md): [Core Function] Kling V3 I2I is the flagship image-to-image editing model (POST /images/generations, model_name=kling-v3 with image). [Strengths] High-quality style transfer and editing up to 2K. [Best For] Single-reference image editing. [Limitations] Do NOT send negative_p… - [Kling V3 Omni Image](https://docs.modellix.ai/kling/kling-v3-omni-image.md): [Core Function] Kling V3 Omni Image is a unified multimodal image generation endpoint (POST /images/omni-image). [Strengths] Multi-image reference, up to 4K, and optional series generation via result_type/series_amount. Use <<>> placeholders in prompt. [Best For] Ch… - [Kling Image O1](https://docs.modellix.ai/kling/kling-image-o1.md): [Core Function] Kling Image O1 is a reasoning-enhanced multimodal image model. [Strengths] It performs deep reasoning over prompts and references to handle complex logic, spatial relationships, and intricate multi-image combinations. [Best For] Highly recommended for: comple… - [Kling Video Effects](https://docs.modellix.ai/kling/kling-video-effects.md): [Core Function] Kling Video Effects applies predefined visual effects to images. [Strengths] It automatically transforms 1 or 2 images into engaging short videos using viral/predefined effect templates. [Best For] Highly recommended for: social media trends and quick visual… - [Kling Avatar](https://docs.modellix.ai/kling/kling-avatar.md): [Core Function] Kling Avatar is a specialized portrait animation model. [Strengths] It precisely animates a portrait image to lip-sync with an audio file or TTS audio ID. [Best For] Highly recommended for: virtual presenters, talking head videos, and digital avatars. [Limita… - [Kling Image Expansion](https://docs.modellix.ai/kling/kling-image-expansion.md): [Core Function] Kling Image Expansion is an outpainting model. [Strengths] It intelligently extends the borders of an image (horizontal, vertical, or asymmetric) while matching the original style and context. [Best For] Highly recommended for: changing aspect ratios, extendi… - [Kolors Virtual Try On V1](https://docs.modellix.ai/kling/kolors-virtual-try-on-v1.md): [Core Function] Kolors Virtual Try-On V1 is a legacy AI fashion and virtual try-on model. [Strengths] Maintained primarily for backward compatibility and stable API contracts. [Best For] Existing legacy e-commerce integrations that have not yet migrated to the newer try-on p… - [Kolors Virtual Try On V1-5](https://docs.modellix.ai/kling/kolors-virtual-try-on-v1-5.md): [Core Function] Kolors Virtual Try-On V1.5 is a specialized AI fashion model. [Strengths] It highly accurately applies garments (including Top+Bottom combinations) onto a person's image, preserving fabric texture and draping. [Best For] Highly recommended for: e-commerce vir… - [MAI Image 2.5](https://docs.modellix.ai/microsoft/mai-image-2-5.md): [Core Function] MAI Image 2.5 is Microsoft's flagship text-to-image generation model. [Strengths] It excels at producing high-quality, detailed images from a text prompt with precise control over output dimensions. [Best For] Highly recommended for: concept art, marketing vi… - [MAI Image 2.5 Edit](https://docs.modellix.ai/microsoft/mai-image-2-5-edit.md): [Core Function] MAI Image 2.5 Edit is Microsoft's flagship image editing model. [Strengths] It excels at applying high-quality, prompt-guided edits and transformations to a single source image. [Best For] Highly recommended for: restyling, object/scene modification, and deta… - [MAI Image 2.5 Flash](https://docs.modellix.ai/microsoft/mai-image-2-5-flash.md): [Core Function] MAI Image 2.5 Flash is Microsoft's fast, cost-efficient text-to-image generation model. [Strengths] It excels at quickly generating solid images from a text prompt with the same dimension controls as MAI Image 2.5. [Best For] Highly recommended for: rapid pro… - [MAI Image 2.5 Flash Edit](https://docs.modellix.ai/microsoft/mai-image-2-5-flash-edit.md): [Core Function] MAI Image 2.5 Flash Edit is Microsoft's fast, cost-efficient image editing model. [Strengths] It excels at quickly applying prompt-guided edits to a single source image. [Best For] Highly recommended for: rapid edits, batch processing, and cost-sensitive work… - [MAI Transcribe 1.5](https://docs.modellix.ai/microsoft/mai-transcribe-1-5.md): [Core Function] MAI-Transcribe 1.5 transcribes a single public audio URL into text via an async task. [Strengths] Multi-lingual recognition, optional locale forcing, phrase-list biasing, and word-level timestamps. [Best For] Meeting notes, captions, and batch audio-to-text.… - [Hailuo 02 FL2V](https://docs.modellix.ai/minimax/hailuo-02-fl2v.md): [Core Function] Hailuo 02 FL2V is a First-Last frame transition video model. [Strengths] It excels at generating a logical, physically accurate video transition that bridges a provided starting frame and an ending frame. [Best For] Highly recommended for: visual morphing, be… - [Hailuo 02 I2V](https://docs.modellix.ai/minimax/hailuo-02-i2v.md): [Core Function] Hailuo 02 I2V is an image-to-video model optimized for physical realism and sustained high resolution. [Strengths] It excels at animating broad scenes, maintaining complex physics, and supporting native 1080p generation for up to 10 seconds. [Best For] Highly… - [Hailuo 02 T2V](https://docs.modellix.ai/minimax/hailuo-02-t2v.md): [Core Function] Hailuo 02 T2V is a text-to-video generation model optimized for physical realism. [Strengths] It excels at complex physics simulation, fluid dynamics, broad cinematic scenes, and natively rendering 1080p video up to 10 seconds without downscaling. [Best For]… - [Hailuo 2.3 Fast I2V](https://docs.modellix.ai/minimax/hailuo-2-3-fast-i2v.md): [Core Function] Hailuo 2.3 Fast I2V is a high-speed, cost-effective image-to-video generation model. [Strengths] It excels at generating videos from images much faster and at roughly 50% lower cost than the standard 2.3 model, while still maintaining the 2.3 architecture's s… - [Hailuo 2.3 I2V](https://docs.modellix.ai/minimax/hailuo-2-3-i2v.md): [Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling… - [Hailuo 2.3 T2V](https://docs.modellix.ai/minimax/hailuo-2-3-t2v.md): [Core Function] Hailuo 2.3 T2V is a flagship text-to-video generation model optimized for human performance and stylization. [Strengths] It excels at capturing intricate human motion, nuanced facial micro-expressions, prompt adherence, and applying highly stylized aesthetics… - [MiniMax H3 FL2V](https://docs.modellix.ai/minimax/minimax-h3-fl2v.md): [Core Function] MiniMax H3 FL2V generates video guided by a first frame, a last frame, or both. [Strengths] It supports first-only, last-only, and first-plus-last conditioning with 4-15 second duration and 768P or 2K output. Output framing follows the input frame imagery. [B… - [MiniMax H3 I2V](https://docs.modellix.ai/minimax/minimax-h3-i2v.md): [Core Function] MiniMax H3 I2V generates video from reference images plus a text prompt. [Strengths] It accepts up to 9 reference images and optional reference audios, with 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] High… - [MiniMax H3 T2V](https://docs.modellix.ai/minimax/minimax-h3-t2v.md): [Core Function] MiniMax H3 T2V is a text-to-video generation model that creates video from a text prompt only. [Strengths] It supports 4-15 second clips, 768P or 2K output, and concrete aspect ratios from cinematic ultrawide to vertical. [Best For] Highly recommended for: pr… - [MiniMax H3 V2V](https://docs.modellix.ai/minimax/minimax-h3-v2v.md): [Core Function] MiniMax H3 V2V generates video guided by one or more reference videos plus a text prompt. [Strengths] It accepts up to 3 reference videos with optional reference images and audios, 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting… - [MiniMax I2V 01](https://docs.modellix.ai/minimax/minimax-i2v-01.md): [Core Function] MiniMax I2V-01 is a legacy image-to-video model. [Strengths] Standard image-to-video generation maintained for backward compatibility. [Best For] Recommended only for: maintaining existing integrations. [Limitations] Do NOT use this model for new creations. I… - [MiniMax I2V 01 Director](https://docs.modellix.ai/minimax/minimax-i2v-01-director.md): [Core Function] MiniMax I2V-01-Director is a legacy image-to-video model with explicit camera controls. [Strengths] It allows direct manipulation of virtual camera movements (pan, zoom, tilt) applied to a starting image. [Best For] Recommended for: legacy systems requiring e… - [MiniMax I2V 01 Live](https://docs.modellix.ai/minimax/minimax-i2v-01-live.md): [Core Function] MiniMax I2V-01-Live is a legacy image-to-video model optimized for anime and stylized art. [Strengths] Originally designed for fast, stylized animation, particularly for 2D illustration and anime styles. [Best For] Recommended for: specific legacy integration… - [MiniMax Image 01 I2I](https://docs.modellix.ai/minimax/minimax-image-01-i2i.md): [Core Function] MiniMax Image-01 I2I is an image-to-image editing and variation model. [Strengths] It excels at generating new images based on a text prompt while structurally referencing one or more input images. [Best For] Highly recommended for: style transfer, generating… - [MiniMax Image 01 Live I2I](https://docs.modellix.ai/minimax/minimax-image-01-live-i2i.md): [Core Function] MiniMax Image-01-Live I2I is a fast image-to-image generation model. [Strengths] It is optimized for rapid generation and style variations, particularly suited for stylized art and anime workflows. [Best For] Highly recommended for: quick style iterations, an… - [MiniMax Image 01 T2I](https://docs.modellix.ai/minimax/minimax-image-01-t2i.md): [Core Function] MiniMax Image-01 T2I is a multimodal text-to-image generation model. [Strengths] It excels at blending high-quality image generation with visual reasoning, allowing for strong prompt adherence and structural understanding. [Best For] Highly recommended for: g… - [MiniMax S2V 01](https://docs.modellix.ai/minimax/minimax-s2v-01.md): [Core Function] MiniMax S2V-01 is a Subject-to-Video generation model. [Strengths] It excels at maintaining strict identity consistency of a specific subject (provided via reference images) while generating a video of that subject performing actions described in a text promp… - [MiniMax T2V 01](https://docs.modellix.ai/minimax/minimax-t2v-01.md): [Core Function] MiniMax T2V-01 is a legacy text-to-video model. [Strengths] Standard text-to-video generation maintained for backward compatibility. [Best For] Recommended only for: maintaining existing integrations that specifically require the T2V-01 endpoint. [Limitations… - [MiniMax T2V 01 Director](https://docs.modellix.ai/minimax/minimax-t2v-01-director.md): [Core Function] MiniMax T2V-01-Director is a legacy text-to-video model with explicit camera controls. [Strengths] It natively accepts explicit camera movement commands (pan, zoom, tilt) alongside the text prompt. [Best For] Recommended for: legacy workflows that specificall… - [MiniMax Voice Clone](https://docs.modellix.ai/minimax/minimax-voice-clone.md): [Core Function] MiniMax Voice Clone clones a speaker from a public reference audio URL and synthesizes new speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; optional language_boost for clone and synthesis; prosody cont… - [Speech 2.8 HD](https://docs.modellix.ai/minimax/speech-2-8-hd.md): [Core Function] MiniMax Speech 2.8 HD is a high-quality text-to-speech model that converts text into natural spoken audio, including expressive paralinguistic cues such as (laughs) and (sighs). [Strengths] Strong narration quality, stable prosody controls (speed, volume, pit… - [Speech 2.8 Turbo](https://docs.modellix.ai/minimax/speech-2-8-turbo.md): [Core Function] MiniMax Speech 2.8 Turbo is a lower-latency text-to-speech model with the same control surface as Speech 2.8 HD, including paralinguistic tags such as (laughs). [Strengths] Faster and more cost-efficient synthesis while retaining prosody, timbre mix, pronunci… - [Whisper 1](https://docs.modellix.ai/openai/whisper-1.md): [Core Function] OpenAI Whisper transcribes a single public audio URL into text via an async task. [Strengths] Multiple output formats (verbose_json with word/segment timestamps, plain text, SRT, VTT), optional language and prompt biasing. [Best For] Meeting notes, podcasts,… - [GPT Image 1.5](https://docs.modellix.ai/openai/gpt-image-1-5.md): [Core Function] GPT Image 1.5 is a versatile text-to-image generation model. [Strengths] It balances solid visual performance with crucial utility features, notably its native support for generating images with transparent backgrounds. [Best For] Highly recommended for: crea… - [GPT Image 1.5 Edit](https://docs.modellix.ai/openai/gpt-image-1-5-edit.md): [Core Function] GPT Image 1.5 Edit is a versatile image-to-image editing and merging model. [Strengths] It excels at complex utility editing tasks, including multi-image merging (up to 16 images), transparent background support, and precise control over how strictly the mode… - [GPT Image 2](https://docs.modellix.ai/openai/gpt-image-2.md): [Core Function] GPT Image 2 is a state-of-the-art text-to-image generation model. [Strengths] It excels at generating highly detailed, photorealistic images from text descriptions, with native support for ultra-high resolutions including 2K and 4K (up to 3840x2160). [Best Fo… - [GPT Image 2 Edit](https://docs.modellix.ai/openai/gpt-image-2-edit.md): [Core Function] GPT Image 2 Edit is a high-resolution image-to-image editing model. [Strengths] It excels at making high-fidelity edits and style transformations to a single source image based on a text prompt, preserving details at up to 4K resolutions. [Best For] Highly re… - [C1 FL2V](https://docs.modellix.ai/pixverse/c1-fl2v.md): [Core Function] PixVerse c1 first-last-frame generates a video that transitions from a start frame to an end frame, guided by a prompt. [Strengths] Controlled start/end composition with smooth interpolation. [Best For] Morphs, scene transitions, before/after motion. [Limitat… - [C1 I2V](https://docs.modellix.ai/pixverse/c1-i2v.md): [Core Function] PixVerse c1 I2V animates a single starting image into a video guided by a text prompt. [Strengths] Smooth, prompt-guided motion from one frame; optional audio. [Best For] Bringing a photo/illustration to life, product showcases, quick cinematic motion from a… - [C1 R2V](https://docs.modellix.ai/pixverse/c1-r2v.md): [Core Function] PixVerse c1 reference-to-video (fusion) generates a video from a prompt while preserving subjects from 1-7 reference images; each reference can be tagged as subject/background and named for @-reference in the prompt. [Strengths] Precise multi-subject composit… - [C1 T2V](https://docs.modellix.ai/pixverse/c1-t2v.md): [Core Function] PixVerse c1 T2V generates a video purely from a text prompt, with no input image. [Strengths] Strong prompt adherence and smooth motion; optional audio. [Best For] Turning an idea or script into video, concept visualization, story beats, social clips from tex… - [Lipsync](https://docs.modellix.ai/pixverse/lipsync.md): [Core Function] PixVerse Lip Sync drives a talking video so the subject's lips match given audio or text-to-speech. [Strengths] Accurate lip synchronization for talking-head videos; supports either an existing audio track or TTS from a chosen speaker voice. [Best For] Dubbin… - [Motion Control](https://docs.modellix.ai/pixverse/motion-control.md): [Core Function] PixVerse Motion Control (Mimic) animates a subject image so it follows the motion of a reference video. [Strengths] Transfers human/animal motion from a driving video onto a still subject. [Best For] Making a character mimic a dance or action, motion retarget… - [Upscale Video](https://docs.modellix.ai/pixverse/upscale-video.md): [Core Function] PixVerse Upscale increases the resolution and clarity of an existing video. [Strengths] Sharper detail and higher-resolution output without changing content. [Best For] Enhancing low-resolution footage, finalizing clips for delivery. [Limitations] Requires an… - [V6 FL2V](https://docs.modellix.ai/pixverse/v6-fl2v.md): [Core Function] PixVerse v6 first-last-frame generates a video that transitions from a start frame to an end frame, guided by a prompt. [Strengths] Controlled start/end composition with smooth interpolation. [Best For] Morphs, scene transitions, before/after motion. [Limitat… - [V6 I2V](https://docs.modellix.ai/pixverse/v6-i2v.md): [Core Function] PixVerse v6 I2V animates a single starting image into a video guided by a text prompt. [Strengths] Smooth, prompt-guided motion from one frame; optional audio and multi-clip. [Best For] Bringing a photo/illustration to life, product showcases, quick cinematic… - [V6 R2V](https://docs.modellix.ai/pixverse/v6-r2v.md): [Core Function] PixVerse v6 reference-to-video (fusion) generates a video from a prompt while preserving subjects from 1-7 reference images; each reference can be tagged as subject/background and named for @-reference in the prompt. [Strengths] Precise multi-subject composit… - [V6 T2V](https://docs.modellix.ai/pixverse/v6-t2v.md): [Core Function] PixVerse v6 T2V generates a video purely from a text prompt, with no input image. [Strengths] Strong prompt adherence and smooth motion; optional audio and multi-clip generation. [Best For] Turning an idea or script into video, concept visualization, story be… - [V6 Video Extend](https://docs.modellix.ai/pixverse/v6-video-extend.md): [Core Function] PixVerse v6 Extend continues an existing video, generating additional seconds guided by a text prompt. [Strengths] Seamless continuation of the existing motion and scene. [Best For] Lengthening clips, continuing an action, adding an ending to footage. [Limita… - [Video Restyle](https://docs.modellix.ai/pixverse/video-restyle.md): [Core Function] PixVerse Restyle re-renders an existing video into a new visual style. [Strengths] Consistent style transfer across all frames. [Best For] Turning footage into anime/3D/painterly looks, stylized remixes. [Limitations] Do NOT use this to change content, motion… - [SkyReels T2V](https://docs.modellix.ai/skyreels/skyreels-t2v.md): **[Core Function]** SkyReels Text-to-Video generates a video purely from a text prompt, with no input media. **[Strengths]** Strong prompt adherence and smooth motion; supports optional audio, 480p/720p/1080p output, and a fast/std quality-speed trade-off. **[Best For]** Tur… - [SkyReels I2V](https://docs.modellix.ai/skyreels/skyreels-i2v.md): **[Core Function]** SkyReels Image-to-Video animates one or more keyframe images into a video guided by a text prompt. **[Strengths]** Supports a start frame, an end frame, and tagged mid-frames for keyframe control; optional audio, 480p/720p/1080p output, and fast/std modes… - [SkyReels R2V](https://docs.modellix.ai/skyreels/skyreels-r2v.md): **[Core Function]** SkyReels Reference-to-Video (multiobject) generates a video from a prompt while preserving the subjects from 1-4 reference images. **[Strengths]** Multi-subject identity preservation, placing specific characters or objects into a newly generated scene. **… - [Single Actor Avatar](https://docs.modellix.ai/skyreels/single-actor-avatar.md): **[Core Function]** SkyReels Single-Actor Avatar (audio-to-video) drives a talking-avatar video from a single portrait image and one audio track. **[Strengths]** Lip-synced single-speaker talking-head video generated from an image plus audio. **[Best For]** Virtual presenter… - [Segmented Camera Motion](https://docs.modellix.ai/skyreels/segmented-camera-motion.md): **[Core Function]** SkyReels Segmented Camera Motion (audio-to-video) generates a talking-avatar video with directed camera movement across time segments. **[Strengths]** Combines an audio-driven avatar with per-segment camera trajectories such as push, pan, crane and rotati… - [SkyReels Omni](https://docs.modellix.ai/skyreels/skyreels-omni.md): **[Core Function]** SkyReels Omni is a reference-driven video model that generates or edits video using reference images (@image) and/or a reference video (@video), bound by tags in the prompt. **[Strengths]** A single endpoint covers motion reference, subject/background rep… - [Video Restyling](https://docs.modellix.ai/skyreels/video-restyling.md): **[Core Function]** SkyReels Restyle re-renders an existing video into a preset visual style. **[Strengths]** Consistent style transfer across all frames into a chosen named art style. **[Best For]** Turning footage into simpsons, lego, paper-cutting, amigurumi, animal-cross… - [Video Extension Single Shot](https://docs.modellix.ai/skyreels/video-extension-single-shot.md): **[Core Function]** SkyReels Single-Shot Extension continues an existing single-shot video, generating additional seconds guided by a text prompt. **[Strengths]** Seamless single-shot continuation of the existing motion and scene. **[Best For]** Lengthening clips and continu… - [Video Extension Shot Switching](https://docs.modellix.ai/skyreels/video-extension-shot-switching.md): **[Core Function]** SkyReels Shot-Switching Extension continues a video while transitioning to a new shot or camera angle. **[Strengths]** Cinematic shot transitions (cut-in, cut-out, reverse-shot, multi-angle, cut-away) when extending footage. **[Best For]** Adding a new sh… - [Sky Lipsync](https://docs.modellix.ai/skyreels/sky-lipsync.md): **[Core Function]** SkyReels Lip Sync (retalking) re-drives a talking video so the subject's lips match a given audio track. **[Strengths]** Accurate lip re-synchronization on an existing talking-head video. **[Best For]** Dubbing, re-voicing talking-head video, and localizi… - [Lip Sync](https://docs.modellix.ai/vidu/lip-sync.md): [Core Function] Vidu Lip Sync is a video-to-video audio synchronization model. [Strengths] It excels at reanimating lip movements in an existing video to precisely match a new replacement audio track, while preserving the original face identity. [Best For] Highly recommended… - [Motion Sync](https://docs.modellix.ai/vidu/motion-sync.md): [Core Function] Vidu Motion Sync is a video-to-video motion transfer model. [Strengths] It excels at accurately extracting physical motion from a source video (e.g., a dancing person) and applying it to a target character image, preserving the target's identity. [Best For] H… - [One Click AD Film](https://docs.modellix.ai/vidu/one-click-ad-film.md): [Core Function] Vidu One-Click AD-Film is an automated marketing video generation model. [Strengths] It excels at transforming 1 to 7 product or scene images into a polished, commercial-style advertisement video (10-60s) automatically. [Best For] Highly recommended for: e-co… - [One Click General Film](https://docs.modellix.ai/vidu/one-click-general-film.md): [Core Function] Vidu One-Click General Film is an automated cinematic film generation model. [Strengths] It excels at automatically stringing together 1 to 7 user-provided images into a cohesive, cinematic film (up to 180s) with appropriate transitions and pacing. [Best For]… - [One Click Trending Replicate](https://docs.modellix.ai/vidu/one-click-trending-replicate.md): [Core Function] Vidu One-Click Trending Replicate is a viral video style cloning model. [Strengths] It excels at analyzing a trending or viral reference video and recreating its specific visual style, transitions, and pacing using the user's own provided subject images. [Bes… - [Template Story](https://docs.modellix.ai/vidu/template-story.md): [Core Function] Vidu Template Story is a narrative video generation model. [Strengths] It excels at placing user-provided character images into predefined, structured narrative templates (like 'love_story' or 'monkey_king') to automatically generate a cohesive short film. [B… - [Vidu Q2 Pro Digital Human](https://docs.modellix.ai/vidu/viduq2-pro-digital-human.md): [Core Function] Vidu Q2 Pro Digital Human is a premium portrait animation model. [Strengths] It excels at generating highly realistic, expressive digital humans from a single portrait image, featuring precise lip-sync to audio and natural facial micro-expressions. [Best For]… - [Vidu Q2 Pro Multi Frame](https://docs.modellix.ai/vidu/viduq2-pro-multi-frame.md): [Core Function] Vidu Q2 Pro Multi-Frame is a sequence-based animation model. [Strengths] It excels at creating continuous, high-quality animation by interpolating through a provided sequence of keyframes (up to 9 images). [Best For] Highly recommended for: complex motion con… - [Vidu Q2 Turbo Digital Human](https://docs.modellix.ai/vidu/viduq2-turbo-digital-human.md): [Core Function] Vidu Q2 Turbo Digital Human is a fast portrait animation model. [Strengths] It excels at quickly animating a static portrait image into a speaking or moving digital human, syncing lip movements to provided audio with low latency. [Best For] Highly recommended… - [Vidu Q2 Turbo Multi Frame](https://docs.modellix.ai/vidu/viduq2-turbo-multi-frame.md): Vidu Q2 Turbo multi-frame animation model. Animates through a sequence of up to 9 frames (1 start + up to 8 key images). Both `start_image` and `key_images` are required. Supports `resolution`. - [Vidu Q3 AD](https://docs.modellix.ai/vidu/viduq3-ad.md): [Core Function] Vidu Q3 AD is a keyframe-driven short-play (short drama) Image-to-Video model that turns a film-style script plus character, scene, and prop reference images into a complete multi-shot short video, automatically planning shots and compositing them in one pass… - [Vidu Q3 Drama](https://docs.modellix.ai/vidu/viduq3-drama.md): [Core Function] Vidu Q3 Drama (Short Play) is a script-to-video model that turns a written script plus character, scene, and prop reference images into a complete multi-shot short drama, automatically planning the shots, transitions, and camera work in a single pass. [Streng… - [Vidu Q3 Mix R2V](https://docs.modellix.ai/vidu/viduq3-mix-r2v.md): [Core Function] Vidu Q3 Mix R2V is a mixed-style reference-to-video generation model. [Strengths] It excels at generating highly consistent character videos by synthesizing and blending features from multiple reference images (up to 7) based on a text prompt. [Best For] High… - [Vidu Q3 Pro Fast I2V](https://docs.modellix.ai/vidu/viduq3-pro-fast-i2v.md): [Core Function] Vidu Q3 Pro Fast I2V is a high-speed Image-to-Video generation model. [Strengths] It excels at generating smooth, physically accurate continuous motion from a single starting frame with extremely low latency. [Best For] Highly recommended for: fast prototypin… - [Vidu Q3 Pro FL2V](https://docs.modellix.ai/vidu/viduq3-pro-fl2v.md): [Core Function] Vidu Q3 Pro FL2V is a premium First-Last frame transition video model. [Strengths] It excels at generating highly detailed, cinematic, and logically consistent video transitions between a starting image and an ending image. [Best For] Highly recommended for:… - [Vidu Q3 Pro I2V](https://docs.modellix.ai/vidu/viduq3-pro-i2v.md): [Core Function] Vidu Q3 Pro I2V is a premium Image-to-Video generation model. [Strengths] It excels at transforming a single starting image into high-fidelity, cinematic video with stable character consistency, complex motion, and synchronized audio-visual capabilities. [Bes… - [Vidu Q3 Pro T2V](https://docs.modellix.ai/vidu/viduq3-pro-t2v.md): [Core Function] Vidu Q3 Pro T2V is a premium cinematic text-to-video generation model. [Strengths] It excels at generating top-tier, realistic videos from text with support for advanced multi-shot 'smart cuts', complex physics, and simultaneous audio-visual generation. [Best… - [Vidu Q3 R2V](https://docs.modellix.ai/vidu/viduq3-r2v.md): [Core Function] Vidu Q3 R2V is a high-quality reference-to-video generation model. [Strengths] It excels at generating detailed, cinematic videos that precisely follow a text prompt while highly preserving the character identity from provided reference images. [Best For] Hig… - [Vidu Q3 Turbo FL2V](https://docs.modellix.ai/vidu/viduq3-turbo-fl2v.md): [Core Function] Vidu Q3 Turbo FL2V is a fast First-Last frame transition video model. [Strengths] It excels at rapidly generating a smooth video transition bridging a specific starting image and an ending image. [Best For] Highly recommended for: quick visual morphs, before-… - [Vidu Q3 Turbo I2V](https://docs.modellix.ai/vidu/viduq3-turbo-i2v.md): [Core Function] Vidu Q3 Turbo I2V is a fast Image-to-Video generation model. [Strengths] It excels at quickly animating a starting frame into a video sequence with strong motion dynamics and low latency. [Best For] Highly recommended for: rapid social media content creation,… - [Vidu Q3 Turbo R2V](https://docs.modellix.ai/vidu/viduq3-turbo-r2v.md): [Core Function] Vidu Q3 Turbo R2V is a fast reference-to-video generation model. [Strengths] It excels at quickly generating dynamic videos based on a text prompt while preserving the identity of the subjects from provided reference images. [Best For] Highly recommended for:… - [Vidu Q3 Turbo T2V](https://docs.modellix.ai/vidu/viduq3-turbo-t2v.md): [Core Function] Vidu Q3 Turbo T2V is a fast text-to-video generation model. [Strengths] It excels at rapidly generating smooth, dynamic videos from text descriptions with very low latency. [Best For] Highly recommended for: fast prototyping, quick visual brainstorming, gener… - [Grok Imagine Image](https://docs.modellix.ai/xai/grok-imagine-image.md): [Core Function] Grok Imagine Image is xAI's standard text-to-image generation model. [Strengths] It excels at quickly generating solid, visually appealing images from a text prompt across a wide range of aspect ratios. [Best For] Highly recommended for: rapid prototyping, so… - [Grok Imagine Image Edit](https://docs.modellix.ai/xai/grok-imagine-image-edit.md): [Core Function] Grok Imagine Image Edit is xAI's standard image editing model. [Strengths] It excels at quickly applying prompt-guided edits and style changes to one or more source images. [Best For] Highly recommended for: fast restyling, quick variations, and lightweight i… - [Grok Imagine Image Quality](https://docs.modellix.ai/xai/grok-imagine-image-quality.md): [Core Function] Grok Imagine Image (Quality) is xAI's high-fidelity text-to-image generation model. [Strengths] It excels at producing richly detailed, high-quality images from a text prompt, with flexible aspect ratios and an optional 2K resolution. [Best For] Highly recomm… - [Grok Imagine Image Quality Edit](https://docs.modellix.ai/xai/grok-imagine-image-quality-edit.md): [Core Function] Grok Imagine Image Edit (Quality) is xAI's high-fidelity image editing model. [Strengths] It excels at applying detailed, prompt-guided edits and style transformations to one or more source images while preserving fine detail. [Best For] Highly recommended fo… - [Grok Imagine Video](https://docs.modellix.ai/xai/grok-imagine-video.md): [Core Function] Grok Imagine Video is xAI's text-to-video generation model. [Strengths] It excels at generating short, dynamic video clips directly from a text prompt, with controllable duration, aspect ratio, and resolution. [Best For] Highly recommended for: short social c… - [Grok Imagine Video 1.5 I2V](https://docs.modellix.ai/xai/grok-imagine-video-1-5-i2v.md): [Core Function] Grok Imagine Video 1.5 I2V animates a single starting image into a video using the Grok Imagine 1.5 generation backbone. [Strengths] It excels at producing motion from one starting frame with the improved 1.5 model. [Best For] Highly recommended for: animatin… - [Grok Imagine Video Edit](https://docs.modellix.ai/xai/grok-imagine-video-edit.md): [Core Function] Grok Imagine Video Edit applies a prompt-guided transformation to an input video. [Strengths] It excels at restyling and modifying an existing video while keeping its original duration and aspect ratio. [Best For] Highly recommended for: restyling clips, appl… - [Grok Imagine Video Extend](https://docs.modellix.ai/xai/grok-imagine-video-extend.md): [Core Function] Grok Imagine Video Extend continues an existing video, generating additional footage beyond its end. [Strengths] It excels at seamlessly extending a clip with new prompt-guided motion. [Best For] Highly recommended for: lengthening short clips, continuing a s… - [Grok Imagine Video I2V](https://docs.modellix.ai/xai/grok-imagine-video-i2v.md): [Core Function] Grok Imagine Video I2V animates a single starting image into a video. [Strengths] It excels at producing smooth motion from one starting frame, guided by a text prompt for the desired movement. [Best For] Highly recommended for: bringing a photo or illustrati… - [Grok Imagine Video R2V](https://docs.modellix.ai/xai/grok-imagine-video-r2v.md): [Core Function] Grok Imagine Video R2V generates a video from a text prompt while preserving the subjects shown in up to 7 reference images. [Strengths] It excels at keeping character/subject identity consistent across a newly generated scene driven by the prompt. [Best For]… - [Grok Voice TTS](https://docs.modellix.ai/xai/grok-voice-tts.md): [Core Function] Grok Voice TTS converts text into natural spoken audio with expressive voices and optional speech tags embedded in the text. [Strengths] Supports 20+ languages (plus auto-detect), 26 built-in voices, multiple codecs (mp3/wav/pcm/mulaw/alaw), and custom voice… - [Grok Voice ASR](https://docs.modellix.ai/xai/grok-voice-asr.md): [Core Function] Grok Voice ASR transcribes a single public audio URL into text via an async task. [Strengths] Word-level timestamps, optional speaker diarization, multichannel transcription, Inverse Text Normalization (format + language), keyterm biasing, and filler-word con… - [Query Task Result](https://docs.modellix.ai/api/get-task-result.md): Query the status and results of an async task by task_id - [STT Result Schema](https://docs.modellix.ai/media-model-api/stt-result-schema.md): Normalized STT result JSON (modellix.transcript.v1) from document resources—full text, channels, sentence and word timings, and optional speaker diarization. - [Upload Media File](https://docs.modellix.ai/api/upload-media-file.md): Upload a single media file via `multipart/form-data` (field name: `file`). Returns a `file_id` and `url` you can pass into prediction APIs. See [Upload media files](/ways-to-use/api#upload-media-files) for limits, supported formats, and the full workflow. - [List Media Files](https://docs.modellix.ai/api/list-media-files.md): List non-expired media files belonging to the authenticated team. Default page size is 100; `limit` is capped at 100. See [Upload media files](/ways-to-use/api#upload-media-files) for the full workflow. - [Delete Media File](https://docs.modellix.ai/api/delete-media-file.md): Delete a media file. After deletion, the file no longer counts toward your upload limit. Returns `404` if the file is not found. See [Upload media files](/ways-to-use/api#upload-media-files) for the full workflow. - [Validate API Key](https://docs.modellix.ai/api/validate-api-key.md): Checks whether the API Key provided in the Authorization header is valid. Invalid, missing, or malformed credentials return a successful response with is_valid set to false. - [Get Team Balance](https://docs.modellix.ai/api/get-team-balance.md): Returns the current available balance for the team associated with the provided API Key. The balance is returned in USD with 4 decimal places. - [Get Logs](https://docs.modellix.ai/api/get-logs.md): Returns paginated media model request logs for the team that owns the API Key. Requires a time window of at most 30 days. Optional `mdlx_user_id` filters by the end-user id sent as `X-Mdlx-User-Id` on async inference. Subject to the query rate limit (separate from media gene… - [List Active Models](https://docs.modellix.ai/api/list-models.md): Returns currently active (published) models with slug, type, documentation URL, description, and optional display price. Optional query `featured=true` limits results to CMS featured models. - [Modellix LLM Overview](https://docs.modellix.ai/llm/overview.md): Get started with the Modellix LLM gateway—OpenAI-compatible Chat Completions and Responses, Anthropic-compatible Messages, and guides for SDKs and coding tools. - [Modellix LLM API Guide](https://docs.modellix.ai/llm/api/api.md): Call Modellix LLM with OpenAI-compatible Chat Completions and Responses or Anthropic-compatible Messages—sync requests, streaming SSE, auth, model list, request logs, errors, and billing. - [Chat Completions](https://docs.modellix.ai/llm/chat-completions.md): OpenAI Chat Completions–compatible endpoint on the Modellix LLM gateway. Accepts a synchronous request with `model` (provider/name format) and an OpenAI-style `messages` array; returns a chat.completion JSON object by default, or SSE (`text/event-stream`) when `stream=true`.… - [Responses](https://docs.modellix.ai/llm/responses.md): OpenAI Responses–compatible endpoint on the Modellix LLM gateway. Accepts a synchronous request with `model` (provider/name format) and `input` (string or content array); returns a Responses JSON object by default, or SSE (`text/event-stream`) when `stream=true`. Use `max_ou… - [Messages](https://docs.modellix.ai/llm/messages.md): Anthropic Messages–compatible endpoint on the Modellix LLM gateway. Accepts a synchronous request with `model` (provider/name format, typically `anthropic/...`), Anthropic-style `messages`, and required `max_tokens`; optional `system` prompt is supported. Returns a Messages… - [List models](https://docs.modellix.ai/llm/list-models.md): Returns an OpenAI-compatible list of models available on the Modellix LLM gateway. Each `data[].id` uses `provider/name` form. Uses the query rate limit (separate from inference RPM). - [Get LLM logs](https://docs.modellix.ai/llm/get-llm-logs.md): Lists LLM request logs for the authenticated team within a time window. Optional `mdlx_user_id` filters by the end-user id sent as `X-Mdlx-User-Id` on inference. Uses the query rate limit (separate from inference RPM). - [Use Modellix LLM with Codex CLI](https://docs.modellix.ai/llm/agent/codex.md): Configure Codex to use Modellix by setting OPENAI_API_KEY and openai_base_url in ~/.codex/config.toml with openai/... models. - [Use Modellix LLM with Claude Code](https://docs.modellix.ai/llm/agent/claude-code.md): Configure Claude Code to call Modellix with ANTHROPIC_BASE_URL (no /v1), a Modellix API key, and anthropic/... models. - [Use Modellix LLM with OpenCode](https://docs.modellix.ai/llm/agent/opencode.md): Add Modellix as a custom OpenAI-compatible provider in OpenCode with @ai-sdk/openai-compatible, baseURL, and provider/name model IDs. - [Use Modellix LLM with OpenClaw 🦞](https://docs.modellix.ai/llm/agent/openclaw.md): Add Modellix as a custom OpenAI-compatible provider in OpenClaw with models.providers, openai-completions, and provider/name model IDs. - [Use Modellix LLM with Hermes Agent](https://docs.modellix.ai/llm/agent/hermes.md): Point Hermes Agent at the Modellix LLM gateway with a Custom Endpoint, config.yaml provider custom, and provider/name model IDs. - [Use Modellix LLM with Junie CLI](https://docs.modellix.ai/llm/agent/junie.md): Add Modellix as a Junie custom LLM profile with OpenAICompletion, a full chat completions URL, and provider/name model IDs. - [Use Modellix LLM with Pi Agent](https://docs.modellix.ai/llm/agent/pi.md): Add Modellix as a custom OpenAI-compatible provider in Pi with models.json, openai-completions, and provider/name model IDs. - [Use Modellix LLM with CodeBuddy](https://docs.modellix.ai/llm/agent/codebuddy.md): Add Modellix as a CodeBuddy custom model with an OpenAI-compatible endpoint, models.json, and provider/name model IDs. - [Use Modellix LLM with WorkBuddy](https://docs.modellix.ai/llm/agent/workbuddy.md): Add Modellix as a WorkBuddy custom OpenAI-compatible model with Endpoint, models.json, and provider/name model IDs. - [Use Modellix LLM with Qwen Code](https://docs.modellix.ai/llm/agent/qwen-code.md): Point Qwen Code at the Modellix LLM gateway with OPENAI_BASE_URL, OPENAI_API_KEY, and provider/name model IDs. - [Use Modellix LLM with Kilo Code](https://docs.modellix.ai/llm/agent/kilo-code.md): Add Modellix as a Kilo Code custom OpenAI Compatible provider with kilo.jsonc, baseURL, and provider/name model IDs. - [Use Modellix LLM with the OpenAI SDK](https://docs.modellix.ai/llm/sdk/openai-sdk.md): Point the OpenAI SDK at the Modellix LLM gateway with OPENAI_BASE_URL and your Modellix API key, then call Chat Completions or Responses with provider/name models. - [Use Modellix LLM with the Anthropic SDK](https://docs.modellix.ai/llm/sdk/anthropic-sdk.md): Point the Anthropic SDK at Modellix with ANTHROPIC_BASE_URL (no /v1) and your Modellix API key, then call Messages with anthropic/... models. - [Use Modellix LLM with the OpenAI Agents SDK](https://docs.modellix.ai/llm/sdk/openai-agents.md): Point the OpenAI Agents SDK at Modellix with AsyncOpenAI base_url, Chat Completions, and provider/name model IDs. - [Use Modellix LLM with the Claude Agent SDK](https://docs.modellix.ai/llm/sdk/claude-agent-sdk.md): Point the Claude Agent SDK at Modellix with ANTHROPIC_BASE_URL (no /v1), a Modellix API key, and anthropic/... model IDs. - [Use Modellix LLM with the Vercel AI SDK](https://docs.modellix.ai/llm/sdk/vercel-ai-sdk.md): Point the Vercel AI SDK at Modellix with createOpenAICompatible, baseURL, and provider/name model IDs. - [Use Modellix LLM with LangChain](https://docs.modellix.ai/llm/framework/langchain.md): Point LangChain ChatOpenAI at the Modellix LLM gateway with base_url, a Modellix API key, and provider/name model IDs. - [Use Modellix LLM with Mastra](https://docs.modellix.ai/llm/framework/mastra.md): Point Mastra agents at Modellix with a custom OpenAI-compatible model url, apiKey, and gateway-style provider/name IDs. - [Use Modellix LLM with CrewAI](https://docs.modellix.ai/llm/framework/crewai.md): Point CrewAI at Modellix with LLM base_url, a Modellix API key, provider/name model IDs, and custom_openai for OpenAI-compatible routing. - [Use Modellix LLM with Agno](https://docs.modellix.ai/llm/framework/agno.md): Point Agno agents at Modellix with OpenAILike base_url, a Modellix API key, and provider/name model IDs. - [Use Modellix LLM with Microsoft Agent Framework](https://docs.modellix.ai/llm/framework/microsoft-agent-framework.md): Point Microsoft Agent Framework at Modellix with OpenAIChatCompletionClient base_url, a Modellix API key, and provider/name model IDs. - [Use Modellix LLM with Cursor](https://docs.modellix.ai/llm/ide/cursor.md): Add Modellix as an OpenAI-compatible provider in Cursor using https://llm.modellix.ai/v1, your Modellix API key, and provider/name model IDs. - [Use Modellix LLM with Cline](https://docs.modellix.ai/llm/ide/cline.md): Point Cline at the Modellix LLM gateway with the OpenAI Compatible provider, Base URL, and provider/name model IDs. - [Use Modellix LLM with CC Switch](https://docs.modellix.ai/llm/tool/cc-switch.md): Add Modellix as a custom CC Switch provider for Claude Code, Codex, and OpenAI Compatible apps with the correct Base URL per protocol. - [Modellix Product Updates and Announcements](https://docs.modellix.ai/changelog/product-updates.md): Follow Modellix product updates and announcements—new API endpoints, console features, billing changes, and platform improvements released each month. - [New AI Models Added to Modellix](https://docs.modellix.ai/changelog/new-models.md): Track newly integrated AI image, video, speech, and LLM models on Modellix, including provider, capabilities, parameters, and availability as they ship. ## OpenAPI Specs - [minimax-v2v](https://docs.modellix.ai/media-model-api/minimax/minimax-v2v.json) - [minimax-t2v](https://docs.modellix.ai/media-model-api/minimax/minimax-t2v.json) - [minimax-i2v](https://docs.modellix.ai/media-model-api/minimax/minimax-i2v.json) - [alibaba-v2v](https://docs.modellix.ai/media-model-api/alibaba/alibaba-v2v.json) - [alibaba-t2v](https://docs.modellix.ai/media-model-api/alibaba/alibaba-t2v.json) - [alibaba-i2v](https://docs.modellix.ai/media-model-api/alibaba/alibaba-i2v.json) - [bytedance-v2v](https://docs.modellix.ai/media-model-api/bytedance/bytedance-v2v.json) - [bytedance-t2v](https://docs.modellix.ai/media-model-api/bytedance/bytedance-t2v.json) - [bytedance-i2v](https://docs.modellix.ai/media-model-api/bytedance/bytedance-i2v.json) - [list-active-models](https://docs.modellix.ai/else-api/list-active-models.json) - [logs](https://docs.modellix.ai/team-api/logs.json) - [llm](https://docs.modellix.ai/llm/api/llm.json) - [google-t2i](https://docs.modellix.ai/media-model-api/google/google-t2i.json) - [validate-api-key](https://docs.modellix.ai/team-api/validate-api-key.json) - [get-team-balance](https://docs.modellix.ai/team-api/get-team-balance.json) - [query-task-result](https://docs.modellix.ai/media-model-api/query-task-result.json) - [alibaba-t2i](https://docs.modellix.ai/media-model-api/alibaba/alibaba-t2i.json) - [alibaba-i2i](https://docs.modellix.ai/media-model-api/alibaba/alibaba-i2i.json) - [media-files](https://docs.modellix.ai/file-api/media-files.json) - [xai-v2v](https://docs.modellix.ai/media-model-api/xai/xai-v2v.json) - [xai-tts](https://docs.modellix.ai/media-model-api/xai/xai-tts.json) - [xai-t2v](https://docs.modellix.ai/media-model-api/xai/xai-t2v.json) - [xai-t2i](https://docs.modellix.ai/media-model-api/xai/xai-t2i.json) - [xai-s2t](https://docs.modellix.ai/media-model-api/xai/xai-s2t.json) - [xai-i2v](https://docs.modellix.ai/media-model-api/xai/xai-i2v.json) - [xai-i2i](https://docs.modellix.ai/media-model-api/xai/xai-i2i.json) - [vidu-v2v](https://docs.modellix.ai/media-model-api/vidu/vidu-v2v.json) - [vidu-t2v](https://docs.modellix.ai/media-model-api/vidu/vidu-t2v.json) - [vidu-i2v](https://docs.modellix.ai/media-model-api/vidu/vidu-i2v.json) - [skyreels-v2v](https://docs.modellix.ai/media-model-api/skywork/skyreels-v2v.json) - [skyreels-t2v](https://docs.modellix.ai/media-model-api/skywork/skyreels-t2v.json) - [skyreels-i2v](https://docs.modellix.ai/media-model-api/skywork/skyreels-i2v.json) - [reve-t2i](https://docs.modellix.ai/media-model-api/reve/reve-t2i.json) - [reve-i2i](https://docs.modellix.ai/media-model-api/reve/reve-i2i.json) - [pixverse-v2v](https://docs.modellix.ai/media-model-api/pixverse/pixverse-v2v.json) - [pixverse-t2v](https://docs.modellix.ai/media-model-api/pixverse/pixverse-t2v.json) - [pixverse-i2v](https://docs.modellix.ai/media-model-api/pixverse/pixverse-i2v.json) - [openai-t2i](https://docs.modellix.ai/media-model-api/openai/openai-t2i.json) - [openai-s2t](https://docs.modellix.ai/media-model-api/openai/openai-s2t.json) - [openai-i2i](https://docs.modellix.ai/media-model-api/openai/openai-i2i.json) - [minimax-t2s](https://docs.modellix.ai/media-model-api/minimax/minimax-t2s.json) - [minimax-t2i](https://docs.modellix.ai/media-model-api/minimax/minimax-t2i.json) - [minimax-s2s](https://docs.modellix.ai/media-model-api/minimax/minimax-s2s.json) - [minimax-i2i](https://docs.modellix.ai/media-model-api/minimax/minimax-i2i.json) - [microsoft-t2i](https://docs.modellix.ai/media-model-api/microsoft/microsoft-t2i.json) - [microsoft-s2t](https://docs.modellix.ai/media-model-api/microsoft/microsoft-s2t.json) - [microsoft-i2i](https://docs.modellix.ai/media-model-api/microsoft/microsoft-i2i.json) - [kling-v2v](https://docs.modellix.ai/media-model-api/kling/kling-v2v.json) - [kling-t2v](https://docs.modellix.ai/media-model-api/kling/kling-t2v.json) - [kling-t2i](https://docs.modellix.ai/media-model-api/kling/kling-t2i.json) - [kling-i2v](https://docs.modellix.ai/media-model-api/kling/kling-i2v.json) - [kling-i2i](https://docs.modellix.ai/media-model-api/kling/kling-i2i.json) - [google-v2v](https://docs.modellix.ai/media-model-api/google/google-v2v.json) - [google-t2v](https://docs.modellix.ai/media-model-api/google/google-t2v.json) - [google-t2s](https://docs.modellix.ai/media-model-api/google/google-t2s.json) - [google-i2v](https://docs.modellix.ai/media-model-api/google/google-i2v.json) - [google-i2i](https://docs.modellix.ai/media-model-api/google/google-i2i.json) - [bytedance-t2i](https://docs.modellix.ai/media-model-api/bytedance/bytedance-t2i.json) - [bytedance-i2i](https://docs.modellix.ai/media-model-api/bytedance/bytedance-i2i.json) - [alibaba-t2s](https://docs.modellix.ai/media-model-api/alibaba/alibaba-t2s.json) - [alibaba-s2t](https://docs.modellix.ai/media-model-api/alibaba/alibaba-s2t.json) - [alibaba-s2s](https://docs.modellix.ai/media-model-api/alibaba/alibaba-s2s.json) - [openapi](https://docs.modellix.ai/api-reference/openapi.json) ## Optional - [Support](mailto:support@modellix.ai) - [Community](https://discord.gg/N2FbcB2cZT) - [Blog](https://www.modellix.ai/blog)