# CosyVoice Clone
Source: https://docs.modellix.ai/alibaba/cosyvoice-clone
/media-model-api/alibaba/alibaba-s2s.json post /alibaba/cosyvoice-clone
[Core Function] CosyVoice Clone clones a speaker from a public reference audio URL and synthesizes new speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; language_hint applies to both cloning and synthesis; SSML, hot_fix, and prosody controls. [Best For] One-off cloned narration and demos where a lasting voice library is not needed. [Limitations] Do NOT use this to obtain a reusable voice library entry; the cloned voice is temporary and is not returned. Reference URL must be publicly accessible. model must be cosyvoice-v3.5-plus or cosyvoice-v3.5-flash. [Routing] Choose cosyvoice-v3.5-plus for higher speech quality; cosyvoice-v3.5-flash for lower latency.
# CosyVoice Design
Source: https://docs.modellix.ai/alibaba/cosyvoice-design
/media-model-api/alibaba/alibaba-t2s.json post /alibaba/cosyvoice-design
[Core Function] CosyVoice Design creates a temporary voice from a natural-language voice_prompt and synthesizes speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; language_hint (zh/en) applies to both design and synthesis; same prosody and format controls as CosyVoice TTS. [Best For] One-off designed speech without storing enrolled voices. [Limitations] Do NOT use this to obtain a reusable voice library entry; the designed voice is temporary and is not returned. model must be cosyvoice-v3.5-plus or cosyvoice-v3.5-flash; voice_prompt and text are required.
# CosyVoice V3 Flash
Source: https://docs.modellix.ai/alibaba/cosyvoice-v3-flash
/media-model-api/alibaba/alibaba-t2s.json post /alibaba/cosyvoice-v3-flash
[Core Function] CosyVoice v3 Flash is Alibaba's low-latency text-to-speech model. [Strengths] Rich system voice catalog, fixed-format instruction on Instruct-capable system voices, SSML, hot_fix, AIGC watermark, Markdown filter (cloned voices only), and multiple audio formats with faster turnaround than Plus. [Best For] Voice assistants, interactive prompts, IVR, short announcements, dialect system voices, and latency-sensitive batch TTS. [Limitations] Do NOT use when maximum speech quality or long-form audiobook fidelity is the priority (use Plus). System-voice instruction must follow CosyVoice voice-list fixed Chinese formats. [Routing] Choose Flash when speed or a richer system-voice catalog matters most. Choose Plus for premium narration quality.
# CosyVoice V3 Plus
Source: https://docs.modellix.ai/alibaba/cosyvoice-v3-plus
/media-model-api/alibaba/alibaba-t2s.json post /alibaba/cosyvoice-v3-plus
[Core Function] CosyVoice v3 Plus is Alibaba's high-quality text-to-speech model. [Strengths] System voices (e.g. longanyang, longanhuan), SSML and LaTeX input, hot_fix pronunciation correction, AIGC watermark, and output in mp3, pcm, wav, or opus. System-voice instruction must use the fixed Chinese formats in the CosyVoice voice list. [Best For] Brand voiceovers, audiobooks, high-quality narration, marketing clips, and scenarios where speech quality matters more than minimum latency. [Limitations] Do NOT use for real-time streaming or word-level timestamps. Fewer system voices than Flash. [Routing] Choose Plus when quality or narration fidelity matters most. Choose cosyvoice-v3-flash for lower latency or a richer system-voice catalog.
# Fun ASR
Source: https://docs.modellix.ai/alibaba/fun-asr
/media-model-api/alibaba/alibaba-s2t.json post /alibaba/fun-asr
[Core Function] Fun-ASR transcribes a single public audio file asynchronously. [Strengths] Hot-word vocabulary, optional speaker diarization, channel selection, and language hints. [Best For] Batch transcription of recordings up to 12 hours. [Limitations] Do NOT send more than one file per request. [Routing] Use fun-asr-mtl when multi-language optimization is preferred.
# Fun ASR MTL
Source: https://docs.modellix.ai/alibaba/fun-asr-mtl
/media-model-api/alibaba/alibaba-s2t.json post /alibaba/fun-asr-mtl
[Core Function] Fun-ASR MTL is the multi-language variant for async recorded speech recognition. [Strengths] Same parameters as fun-asr with multi-language tuning. [Best For] Mixed-language or international audio archives. [Limitations] Do NOT send more than one file per request. [Routing] Choose fun-asr for general use; fun-asr-mtl when the source audio is explicitly multi-language.
# HappyHorse 1.0 I2V
Source: https://docs.modellix.ai/alibaba/happyhorse-1-0-i2v
/media-model-api/alibaba/alibaba-i2v.json post /alibaba/happyhorse-1.0-i2v
[Core Function] HappyHorse 1.0 I2V is a streamlined image-to-video model. [Strengths] It generates high-quality 720P/1080P video (3-15s) from an image efficiently, with native audio support. [Best For] Highly recommended for: rapid image animation and robust character motion. [Limitations] Does not support complex video continuation like Wan 2.7 I2V. [Routing] Use when the user requests 'HappyHorse' or a streamlined image animation.
# HappyHorse 1.0 R2V
Source: https://docs.modellix.ai/alibaba/happyhorse-1-0-r2v
/media-model-api/alibaba/alibaba-i2v.json post /alibaba/happyhorse-1.0-r2v
[Core Function] HappyHorse 1.0 R2V is a reference-to-video model. [Strengths] It excels at maintaining character consistency using up to 9 reference images while generating new video actions based on a prompt. [Best For] Highly recommended for: character-consistent storytelling and generating multiple scenes with the same subject. [Limitations] Do NOT use if you simply want to animate a single image exactly as it is (use HappyHorse I2V instead). [Routing] Use this when the user provides reference images to dictate character/subject appearance in a newly generated action.
# HappyHorse 1.0 T2V
Source: https://docs.modellix.ai/alibaba/happyhorse-1-0-t2v
/media-model-api/alibaba/alibaba-t2v.json post /alibaba/happyhorse-1.0-t2v
[Core Function] HappyHorse 1.0 T2V is a breakout, highly optimized text-to-video model. [Strengths] It provides streamlined, fast, and high-quality video generation (up to 15s at 1080p) with native audio support, acting as a highly efficient alternative to Wan 2.7. [Best For] Highly recommended for: fast experimentation, rapid content creation, and users specifically requesting 'HappyHorse'. [Limitations] Might lack some of the deeply integrated legacy editing features found strictly within the broader Wan 2.7 ecosystem. [Routing] Route to this model when the user explicitly mentions 'HappyHorse' or desires a streamlined, high-performance alternative to Wan.
# HappyHorse 1.0 Video Edit
Source: https://docs.modellix.ai/alibaba/happyhorse-1-0-video-edit
/media-model-api/alibaba/alibaba-v2v.json post /alibaba/happyhorse-1.0-video-edit
[Core Function] HappyHorse 1.0 Video Edit is a streamlined video editing model. [Strengths] It provides high-quality video editing capabilities (with or without reference images) within the highly optimized HappyHorse architecture. [Best For] Highly recommended for: fast, high-quality video modifications, especially when the user explicitly requests HappyHorse. [Limitations] May lack the deep instruction-based logical replacement mechanics of Wan 2.7 Video Editing. [Routing] Route to this model when the user explicitly requests 'HappyHorse' for their video editing task.
# HappyHorse 1.1 I2V
Source: https://docs.modellix.ai/alibaba/happyhorse-1-1-i2v
/media-model-api/alibaba/alibaba-i2v.json post /alibaba/happyhorse-1.1-i2v
[Core Function] HappyHorse 1.1 I2V is Alibaba's latest streamlined first-frame image-to-video model. [Strengths] It turns a single image into high-quality 720P/1080P video with native audio support and 3-15 second duration; output aspect ratio follows the first frame image. [Best For] Highly recommended for: rapid image animation, product motion previews, and simple character or scene animation. [Limitations] It does not accept an explicit ratio parameter; use T2V or R2V when you need a fixed generated aspect ratio. [Routing] Prefer this model when the user provides one image and requests HappyHorse image animation.
# HappyHorse 1.1 R2V
Source: https://docs.modellix.ai/alibaba/happyhorse-1-1-r2v
/media-model-api/alibaba/alibaba-i2v.json post /alibaba/happyhorse-1.1-r2v
[Core Function] HappyHorse 1.1 R2V is Alibaba's latest reference-image-to-video model. [Strengths] It uses 1-9 reference images to preserve subject or character appearance while generating new video actions, supports 720P/1080P output, 3-15 second duration, and expanded aspect ratios including 4:5, 5:4, 9:21, and 21:9. [Best For] Highly recommended for: character-consistent storytelling, reference-based product shots, and multi-image subject composition. [Limitations] Do NOT use if the user simply wants to animate a single image exactly as provided; use HappyHorse I2V instead. [Routing] Use when the user provides one or more reference images and asks for a newly generated HappyHorse video.
# HappyHorse 1.1 T2V
Source: https://docs.modellix.ai/alibaba/happyhorse-1-1-t2v
/media-model-api/alibaba/alibaba-t2v.json post /alibaba/happyhorse-1.1-t2v
[Core Function] HappyHorse 1.1 T2V is Alibaba's latest streamlined text-to-video model. [Strengths] It generates 720P/1080P video with native audio support, 3-15 second duration, and an expanded set of aspect ratios including 4:5, 5:4, 9:21, and 21:9. [Best For] Highly recommended for: fast HappyHorse text-to-video generation, social video formats, and high-throughput content creation. [Limitations] It does not expose custom audio controls; use Wan 2.7 T2V when custom audio input is required. [Routing] Prefer this model when the user explicitly requests HappyHorse text-to-video or wants the latest HappyHorse generation quality.
# Qwen Audio 3.0 TTS Flash
Source: https://docs.modellix.ai/alibaba/qwen-audio-3-0-tts-flash
/media-model-api/alibaba/alibaba-t2s.json post /alibaba/qwen-audio-3.0-tts-flash
[Core Function] Qwen-Audio 3.0 TTS Flash is Alibaba's low-latency Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Fast synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voices include longanhuan_v3.6, longjielidou_v3.6, loongeva_v3.6, and loongjohn (see the Qwen-Audio-TTS voice list). [Best For] Voice assistants, interactive prompts, short announcements, multilingual product flows, and latency-sensitive batch TTS using Qwen-Audio voices. [Limitations] Do NOT mix Plus-only voices (e.g. longanlingxin) with Flash. Use Plus when maximum narration quality matters more than turnaround time. [Routing] Choose Flash when speed matters most. Choose qwen-audio-3.0-tts-plus for premium narration quality.
# Qwen Audio 3.0 TTS Plus
Source: https://docs.modellix.ai/alibaba/qwen-audio-3-0-tts-plus
/media-model-api/alibaba/alibaba-t2s.json post /alibaba/qwen-audio-3.0-tts-plus
[Core Function] Qwen-Audio 3.0 TTS Plus is Alibaba's high-quality Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Natural speech synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voices include longanlingxin and longanlufeng (see the Qwen-Audio-TTS voice list). [Best For] Premium narration, brand voiceovers, multilingual product audio, and quality-sensitive batch TTS when Qwen-Audio voices are preferred. [Limitations] Do NOT mix Flash-only voices (e.g. longanhuan_v3.6) with Plus. [Routing] Choose Plus when speech quality is the priority. Choose qwen-audio-3.0-tts-flash when lower latency matters more.
# Qwen Image 3.0
Source: https://docs.modellix.ai/alibaba/qwen-image-3-0
/media-model-api/alibaba/alibaba-t2i.json post /alibaba/qwen-image-3.0
[Core Function] Qwen Image 3.0 is Alibaba's standard text-to-image model balancing quality and speed. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite (direct/agent modes), batch generation of 1-6 images, and long structured prompts. [Best For] General creative stills, posters with readable text, product shots, and multi-variant exploration (n up to 6) when Pro-tier photorealism is not required. [Limitations] Do NOT use this if the user needs image editing with reference images (use Qwen Image 3.0 Edit). Keep total pixels within 512*512 to 2048*2048. Very long prompts combined with a long negative_prompt may exceed the model input capacity (about 4.5k tokens total). [Routing] Prefer Qwen Image 3.0 Pro for higher photorealism; use this for balanced quality/speed. If the user provides reference image(s) to edit, route to Qwen Image 3.0 Edit instead.
# Qwen Image 3.0 Edit
Source: https://docs.modellix.ai/alibaba/qwen-image-3-0-edit
/media-model-api/alibaba/alibaba-i2i.json post /alibaba/qwen-image-3.0-edit
[Core Function] Qwen Image 3.0 Edit is the standard image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (width*height), intelligent prompt rewrite (direct mode), and 1-6 outputs while preserving subject identity. [Best For] Background replacement, outfit or style changes, multi-image fusion, and iterative retouching when Pro-tier quality is not required. [Limitations] Do NOT use this for pure text-to-image with no reference images (use Qwen Image 3.0 instead). Keep output total pixels within 512*512 to 2048*2048. prompt_extend_mode only supports direct (agent is T2I-only). [Routing] Route here when the user provides reference image(s) and wants balanced Qwen 3.0 edit quality. Prefer Qwen Image 3.0 Pro Edit for higher quality edits.
# Qwen Image 3.0 Pro
Source: https://docs.modellix.ai/alibaba/qwen-image-3-0-pro
/media-model-api/alibaba/alibaba-t2i.json post /alibaba/qwen-image-3.0-pro
[Core Function] Qwen Image 3.0 Pro is Alibaba's latest text-to-image model with strong prompt following and photorealism. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite, batch generation of 1-6 images, and long structured prompts for complex layouts. [Best For] Highly recommended for: photorealistic stills, marketing posters with readable text, detailed scene compositions, multi-panel layouts, product hero shots, and multi-variant creative exploration (n up to 6). [Limitations] Do NOT use this if the user needs native 4K output, thinking-mode reasoning, or image editing with reference images (use Qwen Image 3.0 Pro Edit for edits). Keep total pixels within 512*512 to 2048*2048. Very long prompts combined with a long negative_prompt may exceed the model input capacity (about 4.5k tokens total). [Routing] Prefer this over Qwen Image 2.0 Pro for new Qwen Image text-to-image work. If the user provides reference image(s) to edit, route to Qwen Image 3.0 Pro Edit instead.
# Qwen Image 3.0 Pro Edit
Source: https://docs.modellix.ai/alibaba/qwen-image-3-0-pro-edit
/media-model-api/alibaba/alibaba-i2i.json post /alibaba/qwen-image-3.0-pro-edit
[Core Function] Qwen Image 3.0 Pro Edit is an image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (width*height), intelligent prompt rewrite, 1-6 outputs while preserving subject identity, and long structured edit instructions. [Best For] Highly recommended for: background replacement, outfit or style changes, multi-image fusion, identity-preserving portrait edits, and iterative creative retouching. [Limitations] Do NOT use this for pure text-to-image with no reference images (use Qwen Image 3.0 Pro instead), native 4K output, or thinking-mode reasoning. Keep output total pixels within 512*512 to 2048*2048; input images should follow supported formats and size guidance. Do NOT combine very long prompts with multiple reference images and a long negative_prompt if the request may exceed the model input capacity (about 4.5k tokens total across text and images). [Routing] Route here when the user provides reference image(s) and wants Qwen 3.0 edit quality. For text-only generation without images, use Qwen Image 3.0 Pro.
# Wan 2.7 I2V
Source: https://docs.modellix.ai/alibaba/wan-2-7-i2v
/media-model-api/alibaba/alibaba-i2v.json post /alibaba/wan2.7-i2v
[Core Function] Wan 2.7 I2V is Alibaba's flagship multimodal image-to-video model. [Strengths] It supports multimodal input (text, image, audio, video) for first-frame, start-and-end-frame (FL2V), and video continuation tasks. [Best For] Highly recommended for: complex image animation, cinematic transitions, and video extension workflows. [Limitations] Do NOT use this model if you only need a quick, simple animation where HappyHorse might be faster. [Routing] Use this model by default for complex image-to-video or video continuation tasks.
# Wan 2.7 Image
Source: https://docs.modellix.ai/alibaba/wan-2-7-image
/media-model-api/alibaba/alibaba-t2i.json post /alibaba/wan2.7-image
[Core Function] Wan 2.7 Image is a fast, reasoning-enhanced image generation model. [Strengths] It includes the chain-of-thought reasoning and text rendering of the Pro version, but is optimized for speed, supporting up to 2K resolution. [Best For] Highly recommended for: fast iterations, conceptual design, and generating accurate images with text at standard resolutions. [Limitations] Do NOT use this model if you require 4K print-ready resolution. [Routing] Use this for standard, everyday high-quality image generation requests.
# Wan 2.7 Image Edit
Source: https://docs.modellix.ai/alibaba/wan-2-7-image-edit
/media-model-api/alibaba/alibaba-i2i.json post /alibaba/wan2.7-image-edit
[Core Function] Wan 2.7 Image Edit is a fast, reasoning-enhanced image editing model. [Strengths] Provides the robust editing capabilities of the Wan 2.7 architecture with faster turnaround times. [Best For] Highly recommended for: standard image modifications and style transfers. [Limitations] Do NOT use if you need absolute maximum fidelity or negative prompt support. [Routing] Use for standard, fast image editing tasks.
# Wan 2.7 Image Pro
Source: https://docs.modellix.ai/alibaba/wan-2-7-image-pro
/media-model-api/alibaba/alibaba-t2i.json post /alibaba/wan2.7-image-pro
[Core Function] Wan 2.7 Image Pro is Alibaba's flagship reasoning-enhanced image generation model. [Strengths] It features built-in chain-of-thought reasoning (Thinking Mode), exceptional prompt accuracy, native 12-language text rendering, and generates ultra-high-resolution 4K images. [Best For] Highly recommended for: print-ready large-format posters, complex logical prompts, and generating images containing specific text/typography. [Limitations] Do NOT use this model if you need to generate batch images rapidly (use Wan 2.7 Image instead) or if you specifically need negative prompts (use Qwen Image 2.0 Pro). [Routing] Use this model by default for high-end, 4K, or text-heavy image generation tasks.
# Wan 2.7 Image Pro Edit
Source: https://docs.modellix.ai/alibaba/wan-2-7-image-pro-edit
/media-model-api/alibaba/alibaba-i2i.json post /alibaba/wan2.7-image-pro-edit
[Core Function] Wan 2.7 Image Pro Edit is Alibaba's flagship reasoning-enhanced image editing model. [Strengths] It supports interactive editing, character-consistent multi-image generation, and complex multi-reference modifications with deep reasoning. [Best For] Highly recommended for: professional image retouching, consistent character sheets, and complex structural edits. [Limitations] Do NOT use this model if you specifically need to use negative prompts to exclude elements during editing (use Qwen Image 2.0 Pro Edit instead). [Routing] Use this by default for high-end image editing and multi-reference consistent character generation.
# Wan 2.7 R2V
Source: https://docs.modellix.ai/alibaba/wan-2-7-r2v
/media-model-api/alibaba/alibaba-v2v.json post /alibaba/wan2.7-r2v
[Core Function] Wan 2.7 Reference-to-Video is a highly capable character/entity reference video model. [Strengths] It natively supports entity reference, voice customization, and playbook-based video generation from a single storyboard. [Best For] Highly recommended for: creating consistent video series, brand mascot animation, and storyboard-driven storytelling. [Limitations] Do NOT use this model for simple, single-image direct animation (use I2V instead). [Routing] Use this by default for complex character consistency and storyboard generation tasks on Alibaba.
# Wan 2.7 T2V
Source: https://docs.modellix.ai/alibaba/wan-2-7-t2v
/media-model-api/alibaba/alibaba-t2v.json post /alibaba/wan2.7-t2v
[Core Function] Wan 2.7 T2V is Alibaba's flagship text-to-video generation model. [Strengths] It generates high-fidelity video directly from text with support for custom aspect ratios, audio generation, and intricate semantic adherence. [Best For] Highly recommended for: high-quality commercial video generation, professional storytelling, and dynamic cinematic sequences. [Limitations] Do NOT use this model if the user specifically requests the streamlined 'HappyHorse' workflow. [Routing] Use this model by default for high-end text-to-video requests on the Alibaba platform.
# Wan 2.7 Videoedit
Source: https://docs.modellix.ai/alibaba/wan-2-7-videoedit
/media-model-api/alibaba/alibaba-v2v.json post /alibaba/wan2.7-videoedit
[Core Function] Wan 2.7 Video Editing is an instruction-based video modification model. [Strengths] It supports complex video editing tasks like content replacement using reference images, and replicating actions, effects, and camera movements. [Best For] Highly recommended for: modifying existing video footage, style transfer on videos, and targeted element replacement. [Limitations] Do NOT use this model to generate a brand new video from scratch; it requires an input video. [Routing] Use this model by default whenever a user wants to edit, alter, or restyle an existing video.
# Wan 3.0 I2V
Source: https://docs.modellix.ai/alibaba/wan3-0-i2v
/media-model-api/alibaba/alibaba-i2v.json post /alibaba/wan3.0-i2v
[Core Function] Wan 3.0 I2V is Alibaba Wan 3.0 image-to-video generation supporting first-frame, first-last-frame, and reference-image modes. [Strengths] It can strictly lock the first and last frames or fuse up to 10 reference images with optional reference audio for multimodal guidance. [Best For] Highly recommended for: animating a single keyframe, cinematic first-to-last transitions, multi-image character or product consistency, and image-led storytelling. [Limitations] Do NOT use this model for text-only generation, document/webpage reference, or when a reference video is required. Do NOT mix first_frame/last_frame with reference_images/audio_urls. [Routing] Prefer this model for Wan 3.0 image-driven video. Use Wan 3.0 T2V for prompt/file/link inputs and Wan 3.0 V2V when video_urls are provided.
# Wan 3.0 T2V
Source: https://docs.modellix.ai/alibaba/wan3-0-t2v
/media-model-api/alibaba/alibaba-t2v.json post /alibaba/wan3.0-t2v
[Core Function] Wan 3.0 T2V is Alibaba Wan 3.0 text-to-video generation with optional document or webpage reference. [Strengths] It generates up to 30-second video at 480P/720P/1080P with controllable aspect ratio and optional output audio, and can ground generation on a file or public webpage. [Best For] Highly recommended for: pure text-to-video storytelling, product ads driven by pptx/pdf briefs, turning public articles into short videos, and square or vertical social clips. [Limitations] At least one of prompt, file_url, or link_url is required; an empty body is rejected. Do NOT pass file_url and link_url together. Do NOT use this model if the user provides images or videos as primary media; route those to Wan 3.0 I2V or V2V. [Routing] Prefer this model for Wan 3.0 text-only or file/link-to-video. Use Wan 3.0 I2V for first-frame or reference-image workflows, and Wan 3.0 V2V when reference video is required.
# Wan 3.0 V2V
Source: https://docs.modellix.ai/alibaba/wan3-0-v2v
/media-model-api/alibaba/alibaba-v2v.json post /alibaba/wan3.0-v2v
[Core Function] Wan 3.0 V2V is Alibaba Wan 3.0 reference-video generation that builds new video from one or more input videos. [Strengths] It supports up to 5 reference videos with optional reference images and audio for multimodal composition and prompt-referenced subjects. [Best For] Highly recommended for: video-to-video transformation, multi-subject scenes that cite video1/image1 in the prompt, and extending creative edits from existing clips. [Limitations] Do NOT use this model for text-only, file/link, or first-frame-only workflows. video_urls is required. [Routing] Prefer this model when the user supplies reference video. Use Wan 3.0 T2V for prompt/file/link and Wan 3.0 I2V for image-first generation.
# Z Image Turbo
Source: https://docs.modellix.ai/alibaba/z-image-turbo
/media-model-api/alibaba/alibaba-t2i.json post /alibaba/z-image-turbo
[Core Function] Z-Image Turbo is an older generation text-to-image model. [Strengths] Historically provided faster generation times and lower latency compared to its standard counterparts. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/cost tradeoff of this older model. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Wan 2.7 T2I.
# Delete Media File
Source: https://docs.modellix.ai/api/delete-media-file
/file-api/media-files.json delete /api/v1/media/files/{file_id}
Delete a media file. After deletion, the file no longer counts toward your upload limit. Returns `404` if the file is not found. See [Upload media files](/ways-to-use/api#upload-media-files) for the full workflow.
# Get Logs
Source: https://docs.modellix.ai/api/get-logs
/media-model-api/logs.json get /api/v1/logs
Returns paginated media model request logs for the team that owns the API Key. Requires a time window of at most 30 days. Optional `mdlx_user_id` filters by the end-user id sent as `X-Mdlx-User-Id` on async inference. Subject to the query rate limit (separate from media generation RPM); responses may include `X-RateLimit-*` headers.
# Get Schema
Source: https://docs.modellix.ai/api/get-schema
/media-model-api/get-schema.json get /models/{model_slug}/api_schema
Returns the request and response schema for a Media Model. Pass the model slug in the path (for example alibaba/qwen-image-3.0-pro). This endpoint is public and does not require an API key. The response includes servers (inference base URL) and post (OpenAPI-style operation: description, requestBody JSON Schema, examples, and async task responses). Use servers[0].url as the async generate URL; that call still requires a Bearer API key on https://api.modellix.ai. Resolve slugs from List Active Models.
# Query Task Result
Source: https://docs.modellix.ai/api/get-task-result
/media-model-api/query-task-result.json get /api/v1/tasks/{task_id}
Query the status and results of an async task by task_id
# Get Team Balance
Source: https://docs.modellix.ai/api/get-team-balance
/team-api/get-team-balance.json get /api/v1/team/balance
Returns the current available balance for the team associated with the provided API Key. The balance is returned in USD with 4 decimal places.
# Get Tool Logs
Source: https://docs.modellix.ai/api/get-tool-logs
/tools/logs.json get /v1/logs
Lists Web Search and Web Fetch request logs for the authenticated team within a time window. Optional `tool` filters by web-search or web-fetch. Uses the query rate limit (separate from tool call RPM).
# List Media Files
Source: https://docs.modellix.ai/api/list-media-files
/file-api/media-files.json get /api/v1/media/files
List non-expired media files belonging to the authenticated team. Default page size is 100; `limit` is capped at 100. See [Upload media files](/ways-to-use/api#upload-media-files) for the full workflow.
# List Active Models
Source: https://docs.modellix.ai/api/list-models
/else-api/list-active-models.json get /api/v1/models
Returns currently active (published) models with slug, type, documentation URL, description, and optional display price. Optional query `featured=true` limits results to CMS featured models.
# Upload Media File
Source: https://docs.modellix.ai/api/upload-media-file
/file-api/media-files.json post /api/v1/media/files
Upload a single media file via `multipart/form-data` (field name: `file`). Returns a `file_id` and `url` you can pass into prediction APIs. See [Upload media files](/ways-to-use/api#upload-media-files) for limits, supported formats, and the full workflow.
# Validate API Key
Source: https://docs.modellix.ai/api/validate-api-key
/team-api/validate-api-key.json get /api/v1/apikey/validate
Checks whether the API Key provided in the Authorization header is valid. Invalid, missing, or malformed credentials return a successful response with is_valid set to false.
# Web Fetch
Source: https://docs.modellix.ai/api/web-fetch
/tools/web-tools.json post /v1/web-fetch
Extract readable content from public URLs.
# Web Search
Source: https://docs.modellix.ai/api/web-search
/tools/web-tools.json post /v1/web-search
Search the public web and return ranked results.
# Seedance 2.0 Fast I2V
Source: https://docs.modellix.ai/bytedance/seedance-2-0-fast-i2v
/media-model-api/bytedance/bytedance-i2v.json post /bytedance/seedance-2.0-fast-i2v
[Core Function] Seedance 2.0 Fast I2V is a high-speed multimodal video generation model. [Strengths] Fast generation with the multimodal and multi-shot capabilities of the Seedance 2.0 architecture. [Best For] Highly recommended for: rapid prototyping and quick multi-shot video creation. [Limitations] Do NOT use this model for the absolute highest cinematic fidelity (use the standard 2.0 model instead). [Routing] Choose this model when the user emphasizes 'fast' or 'quick' generation.
# Seedance 2.0 Fast T2V
Source: https://docs.modellix.ai/bytedance/seedance-2-0-fast-t2v
/media-model-api/bytedance/bytedance-t2v.json post /bytedance/seedance-2.0-fast-t2v
[Core Function] Seedance 2.0 Fast T2V is the faster text-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output. [Routing] Use when speed is preferred over maximum resolution.
# Seedance 2.0 Fast V2V
Source: https://docs.modellix.ai/bytedance/seedance-2-0-fast-v2v
/media-model-api/bytedance/bytedance-v2v.json post /bytedance/seedance-2.0-fast-v2v
[Core Function] Seedance 2.0 Fast V2V is a high-speed video-to-video generation model. [Strengths] It offers rapid video transformation capabilities based on the 2.0 architecture. [Best For] Highly recommended for: quick video restyling and fast iterations. [Limitations] Do NOT use this model when absolute maximum visual quality is required (use standard 2.0). [Routing] Use this when speed is the primary concern for video transformations.
# Seedance 2.0 I2V
Source: https://docs.modellix.ai/bytedance/seedance-2-0-i2v
/media-model-api/bytedance/bytedance-i2v.json post /bytedance/seedance-2.0-i2v
[Core Function] Seedance 2.0 I2V is ByteDance's flagship unified multimodal video generation model. [Strengths] It supports complex mixed references (multiple images, audio clips) and generates up to 15s of multi-shot audio-video output with dual-channel audio. [Best For] Highly recommended for: high-end complex video generation, multi-shot narratives, and mixed-reference cinematic production. [Limitations] Do NOT use this model if you only need a very basic legacy generation without complex references. [Routing] Use this model by default for any complex, multi-reference, or high-fidelity image-to-video tasks.
# Seedance 2.0 Mini I2V
Source: https://docs.modellix.ai/bytedance/seedance-2-0-mini-i2v
/media-model-api/bytedance/bytedance-i2v.json post /bytedance/seedance-2.0-mini-i2v
[Core Function] Seedance 2.0 Mini I2V is the lightweight multimodal image-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with optional first/last frame, reference images, and audio references. [Routing] Use for cost-efficient image-to-video.
# Seedance 2.0 Mini T2V
Source: https://docs.modellix.ai/bytedance/seedance-2-0-mini-t2v
/media-model-api/bytedance/bytedance-t2v.json post /bytedance/seedance-2.0-mini-t2v
[Core Function] Seedance 2.0 Mini T2V is the lightweight text-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output. [Routing] Use for cost-efficient text-to-video.
# Seedance 2.0 Mini V2V
Source: https://docs.modellix.ai/bytedance/seedance-2-0-mini-v2v
/media-model-api/bytedance/bytedance-v2v.json post /bytedance/seedance-2.0-mini-v2v
[Core Function] Seedance 2.0 Mini V2V is the lightweight video-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with required reference video and optional text/image/audio references. [Routing] Use for cost-efficient video-to-video transformations.
# Seedance 2.0 T2V
Source: https://docs.modellix.ai/bytedance/seedance-2-0-t2v
/media-model-api/bytedance/bytedance-t2v.json post /bytedance/seedance-2.0-t2v
[Core Function] Seedance 2.0 T2V is ByteDance Dreamina Seedance 2.0 text-to-video. [Strengths] Supports 480p/720p/1080p/4k, 24 fps, 4-15s MP4 output. Text-only input — do not pass images, video, or audio. [Routing] Use for high-fidelity text-to-video when quality or 4k output is requested.
# Seedance 2.0 V2V
Source: https://docs.modellix.ai/bytedance/seedance-2-0-v2v
/media-model-api/bytedance/bytedance-v2v.json post /bytedance/seedance-2.0-v2v
[Core Function] Seedance 2.0 V2V is ByteDance's flagship multimodal video-to-video model. [Strengths] It allows powerful editing and stylization of input videos by supporting mixed references (text, images, video, and audio) and producing multi-shot 15s outputs. [Best For] Highly recommended for: complex video-to-video transformations, restyling existing footage, and creating dynamic multi-shot edits. [Limitations] Do NOT use this model for simple still-image generation (use Seedream instead). [Routing] Use this model by default for any video editing or video-to-video generation tasks.
# Seedance 2.5 I2V
Source: https://docs.modellix.ai/bytedance/seedance-2-5-i2v
/media-model-api/bytedance/bytedance-i2v.json post /bytedance/seedance-2.5-i2v
[Core Function] Seedance 2.5 I2V is ByteDance Dreamina Seedance 2.5 image-to-video generation supporting first-frame, first-and-last-frame, and reference-image modes. [Strengths] It supports up to 30 reference images, optional reference audio, 480p/720p/1080p, 4-30 second duration, and mp4 or mov output. [Best For] Highly recommended for: animating a keyframe, first-to-last transitions, multi-image character consistency, and image-led storytelling on Seedance 2.5. [Limitations] Do NOT mix first_frame_image or last_frame_image with reference_images. Do NOT send video_urls on I2V. Do NOT use audio_urls alone. Do NOT use this model when the user requires 4k output. [Routing] Prefer this model for Seedance 2.5 image-driven generation. Use Seedance 2.5 T2V for text-only requests and Seedance 2.5 V2V when reference video is required.
# Seedance 2.5 T2V
Source: https://docs.modellix.ai/bytedance/seedance-2-5-t2v
/media-model-api/bytedance/bytedance-t2v.json post /bytedance/seedance-2.5-t2v
[Core Function] Seedance 2.5 T2V is ByteDance Dreamina Seedance 2.5 text-to-video generation. [Strengths] It generates longer clips up to 30 seconds at 480p/720p/1080p with optional mp4 or mov output and native audio generation. [Best For] Highly recommended for: longer-form text-to-video storytelling, social clips beyond 15 seconds, and Seedance workflows that need mov output. [Limitations] Do NOT use this model if the user provides images, video, or audio as inputs. Do NOT use this model when the user requires 4k output. [Routing] Prefer Seedance 2.5 T2V when the user needs more than 15 seconds of text-to-video. Use Seedance 2.0 T2V when 4k is required. Use Seedance 2.5 I2V or V2V when media inputs are provided.
# Seedance 2.5 V2V
Source: https://docs.modellix.ai/bytedance/seedance-2-5-v2v
/media-model-api/bytedance/bytedance-v2v.json post /bytedance/seedance-2.5-v2v
[Core Function] Seedance 2.5 V2V is ByteDance Dreamina Seedance 2.5 video-to-video generation covering multimodal reference, video editing, and video extension. [Strengths] It accepts up to 10 reference videos, 30 reference images, and 10 audio clips, with 480p/720p/1080p, 4-30 second output, and mp4 or mov containers. [Best For] Highly recommended for: editing existing clips, extending motion from a base video, multi-reference restyling, and prompt-driven composition that cites @Video1 or @Image1. [Limitations] Do NOT use this model for text-only or first-frame-only workflows. video_urls is required. Do NOT send first_frame_image or last_frame_image. Do NOT use this model when the user requires 4k output. Do NOT use this model for audio-only input. [Routing] Prefer this model for Seedance 2.5 edit, extension, and video-reference jobs. Use Seedance 2.5 T2V for text-only and Seedance 2.5 I2V for image-first generation.
# Seedream 5.0 Lite
Source: https://docs.modellix.ai/bytedance/seedream-5-0-lite
/media-model-api/bytedance/bytedance-t2i.json post /bytedance/seedream-5.0-lite
[Core Function] Seedream 5.0 Lite is a smart, reasoning-enhanced image generation model with real-time web search capabilities. [Strengths] It excels at generating time-sensitive imagery, infographics, and content requiring deep world knowledge or online search, boasting superior prompt understanding and reasoning. [Best For] Highly recommended for: current-event posters, text-heavy designs, and concept art requiring complex logical reasoning. [Limitations] As a 'Lite' model, its absolute photorealistic aesthetic ceiling might be slightly lower than the specialized 4.5 model. [Routing] Use this model by default when the user needs real-time information, deep reasoning, or complex intent understanding in their image.
# Seedream 5.0 Lite Edit
Source: https://docs.modellix.ai/bytedance/seedream-5-0-lite-edit
/media-model-api/bytedance/bytedance-i2i.json post /bytedance/seedream-5.0-lite-edit
[Core Function] Seedream 5.0 Lite Edit is a reasoning-enhanced, smart image editing model. [Strengths] It features superior cross-modal understanding and reasoning, allowing for highly accurate, interactive multi-turn image editing with real-time knowledge enhancement. [Best For] Highly recommended for: complex image editing tasks, structural modifications, and edits requiring deep semantic understanding. [Limitations] As a 'Lite' model, raw visual rendering might not match the 4.5 tier. [Routing] Use this model by default for complex, reasoning-based image editing tasks.
# Seedream 5.0 Pro
Source: https://docs.modellix.ai/bytedance/seedream-5-0-pro
/media-model-api/bytedance/bytedance-t2i.json post /bytedance/seedream-5.0-pro
[Core Function] Seedream 5.0 Pro is ByteDance's flagship professional-grade Text-to-Image (T2I) generation model. [Strengths] It delivers top-tier image quality with enhanced precision control over positions and elements, superior prompt adherence, and improved generation consistency for professional scenarios. [Best For] Highly recommended for: professional design assets, high-fidelity photorealistic imagery, precisely controlled compositions, and brand or commercial visuals where quality matters most. [Limitations] Do NOT use this model for batch image generation or streaming output; it generates exactly one image per request and supports up to 2K resolution (no 3K/4K). [Routing] Choose this Pro model when the user emphasizes ultimate quality or precision. Choose Seedream 5.0 Lite when real-time web knowledge, batch generation, or 3K resolution is needed. To edit an existing image use Seedream 5.0 Pro Edit; to blend multiple reference images use Seedream 5.0 Pro Multi-Reference.
# Seedream 5.0 Pro Edit
Source: https://docs.modellix.ai/bytedance/seedream-5-0-pro-edit
/media-model-api/bytedance/bytedance-i2i.json post /bytedance/seedream-5.0-pro-edit
[Core Function] Seedream 5.0 Pro Edit is a professional-grade single-image editing (I2I) model. [Strengths] It supports interactive precise editing: edit locations can be specified via coordinates, selection boxes, or arrows described in the prompt, with strong element-level control and subject consistency. [Best For] Highly recommended for: precise local retouching, adding/removing/replacing objects at exact positions, style transfer of a single photo, and professional post-editing workflows. [Limitations] Do NOT use this model for text-to-image generation (an input image is required) or for blending multiple reference images; it accepts exactly one input image and outputs exactly one image (no batch or streaming). [Routing] For 2-10 reference images use Seedream 5.0 Pro Multi-Reference; for pure text-to-image use Seedream 5.0 Pro; choose Seedream 5.0 Lite Edit when batch outputs or more than 10 input images are required.
# Seedream 5.0 Pro Multi Reference
Source: https://docs.modellix.ai/bytedance/seedream-5-0-pro-multi-reference
/media-model-api/bytedance/bytedance-i2i.json post /bytedance/seedream-5.0-pro-multi-reference
[Core Function] Seedream 5.0 Pro Multi-Reference is a professional-grade multi-reference image generation (I2I) model that creates a single image from 2-10 reference images plus a text prompt. [Strengths] It excels at reference consistency, preserving characters, styles, and objects across multiple input images while following complex blending instructions with professional-grade quality. [Best For] Highly recommended for: keeping character or style consistency across references, combining subjects from different images into one scene, placing products into reference scenes, and IP-consistent content creation. [Limitations] Do NOT use this model with fewer than 2 or more than 10 reference images, and do NOT use it for batch generation or streaming; it outputs exactly one image per request. [Routing] For single-image editing use Seedream 5.0 Pro Edit; for text-only generation use Seedream 5.0 Pro; choose Seedream 5.0 Lite Edit when up to 14 reference images or batch outputs are needed.
# Deprecated Models on Modellix
Source: https://docs.modellix.ai/changelog/deprecated-models
See which AI image, video, and other media models are no longer available on Modellix, including the provider, model ID, and the date they were retired.
These models are no longer available on Modellix. Do not send new requests to them. Choose a current model from the [API Reference](/alibaba/happyhorse-1-0-i2v) or [New Models](/changelog/new-models).
## Alibaba
### Qwen Image
* `alibaba/qwen-image`: Qwen Image
* `alibaba/qwen-image-2.0`: Qwen Image 2.0
* `alibaba/qwen-image-2.0-edit`: Qwen Image 2.0 Edit
* `alibaba/qwen-image-2.0-pro`: Qwen Image 2.0 Pro
* `alibaba/qwen-image-2.0-pro-edit`: Qwen Image 2.0 Pro Edit
* `alibaba/qwen-image-edit`: Qwen Image Edit
* `alibaba/qwen-image-edit-max`: Qwen Image Edit Max
* `alibaba/qwen-image-edit-plus`: Qwen Image Edit Plus
* `alibaba/qwen-image-edit-plus-2025-10-30`: Qwen Image Edit Plus 2025-10-30
* `alibaba/qwen-image-edit-plus-2025-12-15`: Qwen Image Edit Plus 2025-12-15
* `alibaba/qwen-image-max`: Qwen Image Max
* `alibaba/qwen-image-plus`: Qwen Image Plus
### Wan 2.1
* `alibaba/wan2.1-vace-plus`: Wan 2.1 VACE Plus
### Wan 2.2
* `alibaba/wan2.2-animate-mix`: Wan 2.2 Animate Mix
* `alibaba/wan2.2-animate-move`: Wan 2.2 Animate Move
* `alibaba/wan2.2-i2v-flash`: Wan 2.2 I2V Flash
* `alibaba/wan2.2-i2v-plus`: Wan 2.2 I2V Plus
* `alibaba/wan2.2-kf2v-flash`: Wan 2.2 KF2V Flash
* `alibaba/wan2.2-t2i-flash`: Wan 2.2 T2I Flash
* `alibaba/wan2.2-t2i-plus`: Wan 2.2 T2I Plus
* `alibaba/wan2.2-t2v-plus`: Wan 2.2 T2V Plus
### Wan 2.5
* `alibaba/wan2.5-i2i-preview`: Wan 2.5 I2I Preview
* `alibaba/wan2.5-i2v-preview`: Wan 2.5 I2V Preview
* `alibaba/wan2.5-t2i-preview`: Wan 2.5 T2I Preview
* `alibaba/wan2.5-t2v-preview`: Wan 2.5 T2V Preview
### Wan 2.6
* `alibaba/wan2.6-i2v`: Wan 2.6 I2V
* `alibaba/wan2.6-i2v-flash`: Wan 2.6 I2V Flash
* `alibaba/wan2.6-image`: Wan 2.6 Image
* `alibaba/wan2.6-r2v`: Wan 2.6 R2V
* `alibaba/wan2.6-r2v-flash`: Wan 2.6 R2V Flash
* `alibaba/wan2.6-t2i`: Wan 2.6 T2I
* `alibaba/wan2.6-t2v`: Wan 2.6 T2V
### Wanx 2.1
* `alibaba/wanx2.1-i2v-plus`: Wanx 2.1 I2V Plus
* `alibaba/wanx2.1-i2v-turbo`: Wanx 2.1 I2V Turbo
* `alibaba/wanx2.1-kf2v-plus`: Wanx 2.1 KF2V Plus
* `alibaba/wanx2.1-t2i-plus`: Wanx 2.1 T2I Plus
* `alibaba/wanx2.1-t2i-turbo`: Wanx 2.1 T2I Turbo
* `alibaba/wanx2.1-t2v-plus`: Wanx 2.1 T2V Plus
* `alibaba/wanx2.1-t2v-turbo`: Wanx 2.1 T2V Turbo
## ByteDance
### Seedance 1.0
* `bytedance/seedance-1.0-pro-fast-i2v`: Seedance 1.0 Pro Fast I2V
* `bytedance/seedance-1.0-pro-fast-t2v`: Seedance 1.0 Pro Fast T2V
* `bytedance/seedance-1.0-pro-i2v`: Seedance 1.0 Pro I2V
* `bytedance/seedance-1.0-pro-t2v`: Seedance 1.0 Pro T2V
### Seedance 1.5
* `bytedance/seedance-1.5-pro-i2v`: Seedance 1.5 Pro I2V
* `bytedance/seedance-1.5-pro-t2v`: Seedance 1.5 Pro T2V
### Seedream 4.0
* `bytedance/seedream-4.0-i2i`: Seedream 4.0 I2I
* `bytedance/seedream-4.0-t2i`: Seedream 4.0 T2I
### Seedream 4.5
* `bytedance/seedream-4.5-i2i`: Seedream 4.5 I2I
* `bytedance/seedream-4.5-t2i`: Seedream 4.5 T2I
## Google
### Imagen 4.0
* `google/imagen-4.0-fast-generate-001`: Imagen 4.0 Fast Generate 001
* `google/imagen-4.0-generate-001`: Imagen 4.0 Generate 001
* `google/imagen-4.0-ultra-generate-001`: Imagen 4.0 Ultra Generate 001
### Veo 2
* `google/veo-2-i2v`: Veo 2 I2V
* `google/veo-2-t2v`: Veo 2 T2V
### Veo 3
* `google/veo-3-fast-i2v`: Veo 3 Fast I2V
* `google/veo-3-fast-t2v`: Veo 3 Fast T2V
* `google/veo-3-i2v`: Veo 3 I2V
* `google/veo-3-t2v`: Veo 3 T2V
## Kling
### Image Recognize
* `kling/kling-image-recognize`: Kling Image Recognize
### V1
* `kling/kling-v1-i2i`: Kling V1 I2I
* `kling/kling-v1-i2v`: Kling V1 I2V
* `kling/kling-v1-t2i`: Kling V1 T2I
* `kling/kling-v1-t2v`: Kling V1 T2V
### V1.5
* `kling/kling-v1.5-i2i`: Kling V1.5 I2I
* `kling/kling-v1.5-i2v`: Kling V1.5 I2V
* `kling/kling-v1.5-t2i`: Kling V1.5 T2I
### V1.6
* `kling/kling-v1.6-i2v`: Kling V1.6 I2V
* `kling/kling-v1.6-mi2v`: Kling V1.6 MI2V
* `kling/kling-v1.6-t2v`: Kling V1.6 T2V
### V2
* `kling/kling-v2-i2i`: Kling V2 I2I
* `kling/kling-v2-master-i2v`: Kling V2 Master I2V
* `kling/kling-v2-master-t2v`: Kling V2 Master T2V
* `kling/kling-v2-mi2i`: Kling V2 MI2I
* `kling/kling-v2-new-i2i`: Kling V2 New I2I
### V2.1
* `kling/kling-v2.1-i2i`: Kling V2.1 I2I
* `kling/kling-v2.1-i2v`: Kling V2.1 I2V
* `kling/kling-v2.1-master-i2v`: Kling V2.1 Master I2V
* `kling/kling-v2.1-master-t2v`: Kling V2.1 Master T2V
* `kling/kling-v2.1-mi2i`: Kling V2.1 MI2I
* `kling/kling-v2.1-t2i`: Kling V2.1 T2I
### V2.5 Turbo
* `kling/kling-v2.5-turbo-i2v`: Kling V2.5 Turbo I2V
* `kling/kling-v2.5-turbo-t2v`: Kling V2.5 Turbo T2V
### V2.6
* `kling/kling-v2.6-i2v`: Kling V2.6 I2V
* `kling/kling-v2.6-t2v`: Kling V2.6 T2V
## MiniMax
### Image 01
* `minimax/minimax-image-01-i2i`: MiniMax Image 01 I2I
* `minimax/minimax-image-01-live-i2i`: MiniMax Image 01 Live I2I
* `minimax/minimax-image-01-t2i`: MiniMax Image 01 T2I
### Video 01
* `minimax/minimax-i2v-01`: MiniMax I2V 01
* `minimax/minimax-i2v-01-director`: MiniMax I2V 01 Director
* `minimax/minimax-i2v-01-live`: MiniMax I2V 01 Live
* `minimax/minimax-s2v-01`: MiniMax S2V 01
* `minimax/minimax-t2v-01`: MiniMax T2V 01
* `minimax/minimax-t2v-01-director`: MiniMax T2V 01 Director
## Reve
Reve models are no longer available on Modellix.
* `reve/reve-create`: Reve Create
* `reve/reve-edit`: Reve Edit
* `reve/reve-remix`: Reve Remix
# New AI Models Added to Modellix
Source: https://docs.modellix.ai/changelog/new-models
Track newly integrated AI image, video, speech, and LLM models on Modellix, including provider, capabilities, parameters, and availability as they ship.
## Microsoft
### MAI Image 2.6
* `microsoft/mai-image-2.6`: Microsoft MAI Image 2.6 text-to-image model. It improves text rendering, portraits, 3D imagery, and commercial photorealism versus 2.5, with 768–1365px width/height (product ≤ 1,048,576; default 1024×1024 PNG) and optional auto aspect ratio and web grounding.
* `microsoft/mai-image-2.6-edit`: Microsoft MAI Image 2.6 image editing model. It applies prompt-guided edits to exactly one public JPEG or PNG image URL, outputting PNG with optional auto aspect ratio and web grounding.
* `microsoft/mai-image-2.6-flash`: Microsoft MAI Image 2.6 Flash text-to-image model. It shares the 2.6 request schema (768–1365px, PNG, optional auto aspect ratio and web grounding) with lower latency and cost.
* `microsoft/mai-image-2.6-flash-edit`: Microsoft MAI Image 2.6 Flash image editing model. It shares the 2.6 Edit request schema (one JPEG/PNG URL, PNG output, optional auto aspect ratio and web grounding) with lower latency and cost.
## LLM
Expanded the [LLM gateway](/llm/overview) catalog with OpenAI GPT-6 Astra. Use the same `provider/name` ID across Chat Completions, Responses, and Messages. `~openai/gpt-latest` routes to this model.
Full rates and discounts: [Models & Pricing](/llm/overview#models-and-pricing).
### OpenAI
* `openai/gpt-6-astra`: OpenAI GPT-6 Astra multimodal LLM. Text and image input with text output, 1M context, for software engineering, computer use, and professional work.
## LLM
Expanded the [LLM gateway](/llm/overview) catalog with Google Gemini 3.8 Flash. Use the same `provider/name` ID across Chat Completions, Responses, and Messages. `~google/gemini-flash-latest` routes to this model.
Full rates and discounts: [Models & Pricing](/llm/overview#models-and-pricing).
### Google
* `google/gemini-3.8-flash`: Google Gemini 3.8 Flash multimodal LLM. Text, image, audio, and video input with text output, 1M context, for software engineering, agentic tasks, and multi-step reasoning.
## LLM
Expanded the [LLM gateway](/llm/overview) catalog with Anthropic Claude Fable 5.1. Use the same `provider/name` ID across Chat Completions, Responses, and Messages. `~anthropic/claude-fable-latest` routes to this model.
Full rates and discounts: [Models & Pricing](/llm/overview#models-and-pricing).
### Anthropic
* `anthropic/claude-fable-5.1`: Anthropic Claude Fable 5.1 multimodal LLM. Text and image input with text output, 1M context, for long-horizon agentic coding, knowledge work, and research.
## LLM
Expanded the [LLM gateway](/llm/overview) catalog with ZAI GLM 5.3 Flash and Qwen 3.8 Flash. Use the same `provider/name` IDs across Chat Completions, Responses, and Messages.
Full rates and discounts: [Models & Pricing](/llm/overview#models-and-pricing).
### ZAI
* `zai/glm-5.3-flash`: ZAI GLM 5.3 Flash multimodal LLM. Native text, image, and video input with text output for coding, agent, and office workloads.
### Qwen
* `qwen/qwen3.8-flash`: Qwen 3.8 Flash multimodal LLM. Text, image, and video understanding with text output for coding, agent, and visual workloads.
## LLM
Added Latest Model IDs on the [LLM gateway](/llm/overview). Pass a `~provider/...-latest` ID to always use the current flagship in that series — Modellix updates the target when newer versions ship, so you do not need to change the Model ID.
Current routing:
* `~zai/glm-latest` → `zai/glm-5.3`
* `~xai/grok-latest` → `xai/grok-4.6`
* `~qwen/qwen-latest` → `qwen/qwen3.8-max`
* `~moonshot/kimi-latest` → `moonshot/kimi-k3`
* `~openai/gpt-latest` → `openai/gpt-5.6-sol`
* `~google/gemini-pro-latest` → `google/gemini-3.1-pro`
* `~google/gemini-flash-latest` → `google/gemini-3.7-flash`
* `~anthropic/haiku-latest` → `anthropic/claude-haiku-4.5`
* `~anthropic/sonnet-latest` → `anthropic/claude-sonnet-5`
* `~anthropic/claude-opus-latest` → `anthropic/claude-opus-5`
## LLM
Expanded the [LLM gateway](/llm/overview) catalog with Modellix Free LLM. Use the same `provider/name` ID across Chat Completions, Responses, and Messages.
Full rates and discounts: [Models & Pricing](/llm/overview#models-and-pricing).
### Modellix
* `modellix-ai/free-llm`: Free text LLM on Modellix. Requests route to SOTA open-source free models behind the scenes.
## LLM
Expanded the [LLM gateway](/llm/overview) catalog with ZAI GLM 5.3. Use the same `provider/name` ID across Chat Completions, Responses, and Messages.
Full rates and discounts: [Models & Pricing](/llm/overview#models-and-pricing).
### ZAI
* `zai/glm-5.3`: ZAI GLM 5.3 text LLM. General-purpose Chinese and multilingual chat, coding, and reasoning on the Modellix LLM gateway.
## LLM
Expanded the [LLM gateway](/llm/overview) catalog with Google Gemini 3.7 Flash, xAI Grok 4.6, and ZAI GLM 4.7 Flash. Use the same `provider/name` IDs across Chat Completions, Responses, and Messages.
Full rates and discounts: [Models & Pricing](/llm/overview#models-and-pricing).
### Google
* `google/gemini-3.7-flash`: Google Gemini 3.7 Flash multimodal LLM. Text, image, audio, and video input with text output for coding and agent workloads.
### xAI
* `xai/grok-4.6`: xAI Grok 4.6 multimodal LLM. Text and image input with text output for coding, agent, and knowledge-work tasks.
### ZAI
* `zai/glm-4.7-flash`: ZAI GLM 4.7 Flash text LLM. Fast Chinese and multilingual chat, coding, and reasoning; free on Modellix.
## MiniMax
### MiniMax H3
* `minimax/minimax-h3-t2v`: MiniMax H3 text-to-video model. It generates 4–15s clips at 768P/2K from a prompt only, with required aspect ratios from cinematic ultrawide (21:9) to vertical (9:16).
* `minimax/minimax-h3-i2v`: MiniMax H3 image-to-video model. It generates 4–15s clips at 768P/2K from 1–9 reference images plus a prompt, with optional reference audio (up to 3) and aspect ratio defaulting to 16:9.
* `minimax/minimax-h3-fl2v`: MiniMax H3 first-last-frame video model. It supports first-only, last-only, and first-and-last-frame conditioning for 4–15s clips at 768P/2K; output framing follows the input frame imagery.
* `minimax/minimax-h3-v2v`: MiniMax H3 video-to-video model. It remixes 1–3 reference videos with a prompt, with optional reference images (up to 9) and audio (up to 3), producing 4–15s clips at 768P/2K.
## Alibaba
### Wan 3.0
* `alibaba/wan3.0-t2v`: Alibaba Wan 3.0 text-to-video model. It generates 2–30s clips at 480P/720P/1080P from a prompt, a document (`file_url`), or a public webpage (`link_url`), with optional native audio output.
* `alibaba/wan3.0-i2v`: Alibaba Wan 3.0 image-to-video model. It supports first-frame, first-and-last-frame, and multi-reference image modes (up to 10 images) with optional reference audio, 2–30s duration, and 480P/720P/1080P output.
* `alibaba/wan3.0-v2v`: Alibaba Wan 3.0 video-to-video model. It builds new video from 1–5 reference clips (combined input up to 15s), with optional reference images and audio, producing 2–30s clips at 480P/720P/1080P.
## LLM
Expanded the [LLM gateway](/llm/overview) catalog with Moonshot Kimi, DeepSeek, ZAI GLM, and Qwen models. Use the same `provider/name` IDs across Chat Completions, Responses, and Messages.
Full rates and discounts: [Models & Pricing](/llm/overview#models-and-pricing).
### Moonshot
* `moonshot/kimi-k2.7`: Moonshot Kimi K2.7 multimodal LLM. Text and image input with text output for agent, coding, and general chat workloads.
* `moonshot/kimi-k3`: Moonshot Kimi K3 multimodal LLM. Higher-capability text and image understanding for complex reasoning and long-context tasks.
### DeepSeek
* `deepseek/deepseek-v4-flash`: DeepSeek V4 Flash text LLM. Cost-efficient reasoning and coding for high-volume, latency-sensitive chat and tool use.
* `deepseek/deepseek-v4-pro`: DeepSeek V4 Pro text LLM. Stronger reasoning and coding quality for complex analysis and agent workflows.
### ZAI
* `zai/glm-5.2`: ZAI GLM 5.2 text LLM. General-purpose Chinese and multilingual chat, coding, and reasoning on the Modellix LLM gateway.
### Qwen
* `qwen/qwen3.7-plus`: Qwen 3.7 Plus multimodal LLM. Text and image input with balanced quality and cost for everyday chat and coding.
* `qwen/qwen3.7-max`: Qwen 3.7 Max text LLM. Higher-capability text reasoning and coding for demanding agent and analysis workloads.
* `qwen/qwen3.8-max`: Qwen 3.8 Max multimodal LLM. Text, image, and video understanding with text output for richer multimodal applications.
## ByteDance
### Seedance 2.5
* `bytedance/seedance-2.5-t2v`: ByteDance Dreamina Seedance 2.5 text-to-video model. It generates longer clips up to 30 seconds at 480p/720p with optional mp4 or mov output and native audio generation, suited to longer-form storytelling and social clips beyond 15 seconds.
* `bytedance/seedance-2.5-i2v`: ByteDance Dreamina Seedance 2.5 image-to-video model. It supports first-frame, first-and-last-frame, and multi-reference image modes (up to 30 images) with optional reference audio, 4–30s duration, and mp4 or mov output.
* `bytedance/seedance-2.5-v2v`: ByteDance Dreamina Seedance 2.5 video-to-video model. It covers multimodal reference, video editing, and video extension with up to 10 reference videos, 30 reference images, and 10 audio clips, producing 4–30s mp4 or mov output.
## LLM
Modellix now offers an **LLM gateway** at `https://llm.modellix.ai` for synchronous chat and coding workloads (optional SSE streaming). Use one Modellix API key with `provider/name` model IDs. Media generation remains on the async media API at `https://api.modellix.ai`.
See [LLM Overview](/llm/overview) for pricing, protocols, and client setup.
### Protocols
Three mainstream protocol surfaces:
* **OpenAI-compatible Chat Completions** — [`POST /v1/chat/completions`](/llm/chat-completions)
* **OpenAI-compatible Responses** — [`POST /v1/responses`](/llm/responses)
* **Anthropic-compatible Messages** — [`POST /v1/messages`](/llm/messages)
### Client Integrations
Drop-in guides for common SDKs and coding tools:
* [OpenAI SDK](/llm/sdk/openai-sdk)
* [Anthropic SDK](/llm/sdk/anthropic-sdk)
* [Codex](/llm/agent/codex)
* [Claude Code](/llm/agent/claude-code)
* [Cursor](/llm/ide/cursor)
* [OpenCode](/llm/agent/opencode)
### Supported LLMs
Initial catalog (Model IDs use `provider/name`):
#### OpenAI
* `openai/gpt-5.6-sol`
* `openai/gpt-5.6-terra`
* `openai/gpt-5.6-luna`
* `openai/gpt-5.5`
#### Anthropic
* `anthropic/claude-opus-5`
* `anthropic/claude-sonnet-5`
* `anthropic/claude-haiku-4.5`
#### Google
* `google/gemini-3.6-flash`
* `google/gemini-3.5-flash`
* `google/gemini-3.1-pro`
#### xAI
* `xai/grok-4.5`
* `xai/grok-4.3`
Full rates and discounts: [Models & Pricing](/llm/overview#models-and-pricing).
## Alibaba
### Qwen Image 3.0 Pro
* `alibaba/qwen-image-3.0-pro`: Alibaba's latest Qwen Image 3.0 Pro text-to-image model. It delivers strong prompt following and photorealism, supports free-form output size, optional negative prompts, intelligent prompt rewrite, and batch generation of up to 6 images (within 512×512–2048×2048 total pixels).
* `alibaba/qwen-image-3.0-pro-edit`: Alibaba's Qwen Image 3.0 Pro image editing model. It applies instruction-based edits and multi-image fusion from 1–3 reference images, with optional negative prompts, free-form output size, and up to 6 outputs while preserving subject identity.
## Kling
### Kling Video O1
* `kling/kling-video-o1-t2v`: Kling Video O1 text-to-video model. Reasoning-enhanced prompt planning for complex physical interactions and logically demanding scenes from text alone, with 3–10s duration and 720p/1080p output.
* `kling/kling-video-o1-i2v`: Kling Video O1 image-to-video model. Animates from 1–7 reference images with reasoning-enhanced motion planning for complex physics grounded in reference frames, at 720p/1080p and 3–10s duration.
### Kling V3 Omni
* `kling/kling-v3-omni-t2v`: Kling V3 Omni text-to-video model. Multimodal-leaning T2V with stronger semantic control and subject consistency, native audio options, flexible 3–15s duration, and up to 4K cinematic output for narrative clips.
* `kling/kling-v3-omni-i2v`: Kling V3 Omni image-to-video model. Reference-led animation from one or more images that preserves identity, wardrobe, and product look, with flexible duration and optional native audio.
### Kling V3 Turbo
* `kling/kling-v3-turbo-t2v`: Kling V3 Turbo text-to-video model. Speed- and cost-optimized short-form generation with native audio and improved lip-sync, targeting practical 720p/1080p clips for rapid prototyping and high-volume pipelines.
* `kling/kling-v3-turbo-i2v`: Kling V3 Turbo image-to-video model. Fast, cost-efficient single-keyframe animation with optional native audio and strong lip-sync for social ads, talking-head starters, and bulk I2V jobs.
## Alibaba
### Text-to-Speech
* `alibaba/qwen-audio-3.0-tts-plus`: Alibaba's high-quality Qwen-Audio 3.0 text-to-speech model. It delivers natural speech with Qwen-Audio system voices, prosody and format controls, SSML, instruction, and language hints for premium narration, brand voiceovers, and multilingual product audio.
* `alibaba/qwen-audio-3.0-tts-flash`: Alibaba's low-latency Qwen-Audio 3.0 text-to-speech model. It shares the same control surface as Plus with faster turnaround, suited to voice assistants, interactive prompts, short announcements, and latency-sensitive batch TTS using Qwen-Audio voices.
## ByteDance
### Seedream 5.0 Pro
* `bytedance/seedream-5.0-pro`: ByteDance's flagship professional-grade text-to-image model. It delivers top-tier image quality with enhanced precision control over positions and elements, superior prompt adherence, and improved generation consistency for professional design assets and commercial visuals (up to 2K).
* `bytedance/seedream-5.0-pro-edit`: ByteDance's professional-grade single-image editing model. It supports interactive precise edits via coordinates, selection boxes, or arrows described in the prompt, with strong element-level control and subject consistency.
* `bytedance/seedream-5.0-pro-multi-reference`: ByteDance's professional-grade multi-reference image generation model. It creates a single image from 2–10 reference images plus a text prompt, preserving characters, styles, and objects across references for IP-consistent content creation.
## Alibaba
### Text-to-Speech
* `alibaba/cosyvoice-v3-plus`: Alibaba's high-quality CosyVoice text-to-speech model. It supports system voices, SSML and LaTeX input, hot-word pronunciation correction, and mp3/pcm/wav/opus output for brand voiceovers, audiobooks, and premium narration.
* `alibaba/cosyvoice-v3-flash`: Alibaba's low-latency CosyVoice text-to-speech model. It offers a richer system-voice catalog and faster turnaround for voice assistants, interactive prompts, dialect voices, and latency-sensitive batch TTS.
* `alibaba/cosyvoice-design`: Alibaba's one-shot voice design plus TTS model. It creates a temporary voice from a natural-language voice prompt and synthesizes speech in a single async request without storing an enrolled voice library entry.
### Speech-to-Speech
* `alibaba/cosyvoice-clone`: Alibaba's one-shot CosyVoice clone plus TTS model. It clones a speaker from a public reference audio URL and synthesizes new speech in one request, with language hints, SSML, hot words, and prosody controls.
### Speech-to-Text
* `alibaba/fun-asr`: Alibaba's Fun-ASR speech recognition model. It asynchronously transcribes a single public audio file with hot-word vocabulary, optional speaker diarization, channel selection, and language hints for batch recordings up to 12 hours.
* `alibaba/fun-asr-mtl`: Alibaba's multi-language Fun-ASR variant. It uses the same async transcription controls as Fun-ASR with multi-language tuning for mixed-language or international audio archives.
## Google
### Text-to-Speech
* `google/gemini-3.1-flash-tts`: Google's low-latency Gemini 3.1 Flash TTS model. It supports single-speaker and two-speaker dialogue, 30 prebuilt voices, 70+ languages, and expressive delivery via style prompts and inline audio tags, outputting WAV (24 kHz mono PCM).
## MiniMax
### Text-to-Speech
* `minimax/speech-2.8-hd`: MiniMax Speech 2.8 HD text-to-speech model. It converts text into natural spoken audio with expressive paralinguistic cues, stable prosody controls, optional timbre mixing, pronunciation overrides, and subtitles for premium narration and brand voiceovers.
* `minimax/speech-2.8-turbo`: MiniMax Speech 2.8 Turbo text-to-speech model. It shares the Speech 2.8 control surface with lower latency and cost, suited to interactive assistants, high-volume batches, and quick narration drafts.
### Speech-to-Speech
* `minimax/minimax-voice-clone`: MiniMax's one-shot voice clone plus TTS model. It clones a speaker from a public reference audio URL and synthesizes new speech in one request, with language boost, prosody, pronunciation overrides, and flexible audio formats.
## Microsoft
### Speech-to-Text
* `microsoft/mai-transcribe-1.5`: Microsoft's MAI-Transcribe 1.5 speech recognition model. It asynchronously transcribes a single public audio URL with multi-lingual recognition, optional locale forcing, phrase-list biasing, and word-level timestamps.
## OpenAI
### Speech-to-Text
* `openai/whisper-1`: OpenAI Whisper speech recognition model. It asynchronously transcribes a single public audio URL with multiple output formats (verbose JSON with timestamps, plain text, SRT, VTT) and optional language or prompt biasing.
## xAI
### Text-to-Speech
* `xai/grok-voice-tts`: xAI's Grok Voice text-to-speech model. It converts text into natural spoken audio with 26 built-in voices, 20+ languages, optional speech tags, and multiple codecs (mp3/wav/pcm/mulaw/alaw) for product voiceovers and telephony.
### Speech-to-Text
* `xai/grok-voice-asr`: xAI's Grok Voice ASR model. It asynchronously transcribes a single public audio URL with word-level timestamps, optional speaker diarization, multichannel transcription, and Inverse Text Normalization.
## Google
### Gemini Omni Flash
* `google/gemini-omni-flash-t2v`: Google's fast multimodal Text-to-Video model built on the Interactions API. It turns a text prompt into a short 720p video with natively synchronized audio, optimized for low-latency prototyping and short social clips.
* `google/gemini-omni-flash-i2v`: Google's fast Image-to-Video model. It animates a single input image as the opening frame into a short 720p video with synchronized audio.
* `google/gemini-omni-flash-r2v`: Google's fast Reference-to-Video model. It fuses up to three reference images guided by a text prompt into a coherent short 720p clip with synchronized audio.
* `google/gemini-omni-flash-video-edit`: Google's instruction-driven Video Edit model. It applies natural-language edits to an existing video while preserving the source length and aspect ratio, with synchronized audio.
## Google
### Nano Banana 2 Lite
* `google/nano-banana-2-lite`: Google's lightweight, cost-efficient text-to-image model in the Nano Banana 2 family. Built on Gemini 3.1 Flash Lite Image, it delivers rapid image generation at a fraction of the cost, making it ideal for high-volume creative iteration and rapid prototyping.
* `google/nano-banana-2-lite-edit`: Google's lightweight, cost-efficient image-to-image editing model. It excels at fast, low-cost instruction-based editing across up to 14 input images.
## Skywork
### Text-to-Video
* `skyreels/skyreels-t2v`: SkyReels Text-to-Video model. Generates highly realistic videos purely from text prompts with smooth motion dynamics, strong prompt compliance, optional audio, and adjustable speed/quality modes.
### Image-to-Video & Talking Avatar
* `skyreels/skyreels-i2v`: SkyReels Image-to-Video model. Animates static images with advanced first-frame, end-frame, or mid-frame keyframe control, supporting custom motion guidance and optional audio.
* `skyreels/skyreels-r2v`: SkyReels Reference-to-Video model. Generates a video from a text prompt while preserving the unique identities of subjects from 1-4 reference images.
* `skyreels/single-actor-avatar`: SkyReels Single-Actor Talking Avatar model. It drives a portrait image using a single audio track, creating a lifelike virtual presenter with perfect lip synchronization.
* `skyreels/segmented-camera-motion`: SkyReels Segmented Camera Motion model. Combines audio-driven portrait animation with precise, multi-segment directed camera trajectories (such as pan, push, crane, and rotation).
### Video-to-Video & Omni Editing
* `skyreels/skyreels-omni`: SkyReels Omni Reference Video model. A unified endpoint for advanced motion reference, subject/background swap, object insertion/removal, and multi-image driven video synthesis.
* `skyreels/video-restyling`: SkyReels Video Restyle model. Consistent video style transfer that re-renders an existing clip into specific creative art styles (Lego, Simpsons, Van Gogh, and more).
* `skyreels/video-extension-single-shot`: SkyReels Video Extension (Single Shot) model. Seamlessly extends the footage and motion of an existing single-shot video.
* `skyreels/video-extension-shot-switching`: SkyReels Video Extension (Shot Switching) model. Appends additional video footage while transitioning cinematic shots and angles (cut-in, reverse-shot, etc.).
* `skyreels/sky-lipsync`: SkyReels Lip Sync model. Re-synchronizes talking lip movements on any existing video to match a target audio file.
## Reve
### Image Generation
* `reve/reve-create`: Reve's flagship text-to-image model. It features exceptional prompt adherence and superior in-image text and typography rendering, producing clean, beautifully composed professional layouts.
### Image Editing & Remixing
* `reve/reve-edit`: Reve's single-image editing model. It performs highly targeted, localized edits on an input image based on natural-language instructions, while preserving the surrounding details.
* `reve/reve-remix`: Reve's multi-image composition model. It blends 1 to 6 reference images guided by a prompt, supporting precise in-prompt references to specific images using `
` tags.
## Vidu
### Q3 AD (Keyframe Short Play)
* `viduq3-ad`: Vidu Q3 AD keyframe-driven short-play Image-to-Video model. It transforms a film-style script plus character, scene, and prop reference images into a complete multi-shot short video, automatically planning shots and compositing them in one pass with high narrative coherence.
## PixVerse
### Text-to-Video
* `v6-t2v`: V6 Text-to-Video model. Generates videos purely from a text prompt with strong prompt adherence, smooth motion, and support for optional audio and multi-clip generation.
* `c1-t2v`: C1 Text-to-Video model. Generates videos purely from a text prompt with strong prompt adherence and smooth motion, optimized for standard single-clip generation.
### Image-to-Video & Transitions
* `v6-i2v`: V6 Image-to-Video model. Animates a single starting image into a video guided by a text prompt, supporting smooth motion, optional audio, and multi-clip generation.
* `c1-i2v`: C1 Image-to-Video model. Animates a single starting image into a video guided by a text prompt, optimized for standard single-clip generation.
* `v6-fl2v`: V6 First-Last-Frame to Video model. Generates a video transitioning smoothly from a starting frame to an ending frame, guided by a prompt.
* `c1-fl2v`: C1 First-Last-Frame to Video model. Generates a video transitioning smoothly from a starting frame to an ending frame, guided by a prompt.
### Character & Reference Video (Fusion)
* `v6-r2v`: V6 Reference-to-Video (Fusion) model. Generates a video from a prompt while preserving subjects from 1-7 reference images, allowing role tagging (subject/background) and named references.
* `c1-r2v`: C1 Reference-to-Video (Fusion) model. Generates a video from a prompt while preserving subjects from 1-7 reference images, allowing role tagging (subject/background) and named references.
### Video-to-Video & Editing
* `lipsync`: Lip Sync model. Re-times the subject's lips in a talking-head video to match a given audio URL or text-to-speech content.
* `video-restyle`: Restyle Video model. Re-renders an existing video into a new visual style using preset style IDs or custom text prompts.
* `v6-video-extend`: V6 Extend Video model. Continues an existing video for 1-15 additional seconds guided by a text prompt, ensuring seamless motion and scene continuation.
* `motion-control`: Motion Control (Mimic) model. Animates a subject image to follow the body movements and poses of a reference video.
* `upscale-video`: Upscale Video model. Enhances the resolution and clarity of an existing video without changing its content, style, or motion.
## ByteDance
### Seedance 2.0 (Standard)
* `seedance-2.0-t2v`: ByteDance Dreamina Seedance 2.0 text-to-video model. It excels at high-fidelity text-to-video generation, supporting resolutions up to 4K, 24 fps, 4-15s MP4 outputs, and optional image or audio references.
### Seedance 2.0 Fast
* `seedance-2.0-fast-t2v`: A speed-optimized variant of Seedance 2.0 text-to-video model, offering rapid generation at 480p/720p resolutions and 24 fps.
### Seedance 2.0 Mini
* `seedance-2.0-mini-t2v`: A lightweight, cost-efficient text-to-video model in the Seedance 2.0 family, supporting 480p/720p resolutions and 24 fps.
* `seedance-2.0-mini-i2v`: A lightweight, cost-efficient image-to-video model that animates static images into 24 fps videos, supporting optional first/last frame, reference images, and audio references.
* `seedance-2.0-mini-v2v`: A lightweight, cost-efficient video-to-video model for transforming existing video clips with text, image, or audio references.
## Microsoft
### Image Generation
* `mai-image-2.5`: Microsoft's flagship text-to-image generation model. It excels at producing high-quality, detailed images from a text prompt with precise control over output dimensions.
* `mai-image-2.5-flash`: Microsoft's fast, cost-efficient text-to-image generation model. It excels at quickly generating solid images from a text prompt with the same dimension controls.
### Image Editing
* `mai-image-2.5-edit`: Microsoft's flagship image editing model. It excels at applying high-quality, prompt-guided edits and transformations to a single source image.
* `mai-image-2.5-flash-edit`: Microsoft's fast, cost-efficient image editing model. It excels at quickly applying prompt-guided edits to a single source image.
## xAI
### Image Generation
* `grok-imagine-image`: xAI's standard text-to-image generation model. It excels at quickly generating solid, visually appealing images from a text prompt across a wide range of aspect ratios.
* `grok-imagine-image-quality`: xAI's high-fidelity text-to-image generation model. It excels at producing richly detailed, high-quality images from a text prompt, with flexible aspect ratios and an optional 2K resolution.
### Image Editing
* `grok-imagine-image-edit`: xAI's standard image editing model. It excels at quickly applying prompt-guided edits and style changes to one or more source images (up to 3).
* `grok-imagine-image-quality-edit`: xAI's high-fidelity image editing model. It excels at applying detailed, prompt-guided edits and style transformations to one or more source images (up to 3) while preserving fine detail.
### Video Generation
* `grok-imagine-video`: xAI's text-to-video generation model. It excels at generating short, dynamic video clips directly from a text prompt, with controllable duration, aspect ratio, and resolution.
* `grok-imagine-video-i2v`: xAI's standard image-to-video model. It animates a single starting image into a video, producing smooth motion guided by a text prompt.
* `grok-imagine-video-1.5-i2v`: xAI's image-to-video model using the Grok Imagine 1.5 generation backbone. It animates a single starting image into a video with the improved 1.5 model.
* `grok-imagine-video-r2v`: xAI's reference-to-video model. It generates a video from a text prompt while preserving the subjects shown in up to 7 reference images.
### Video Editing & Extension
* `grok-imagine-video-edit`: xAI's video editing model. It applies prompt-guided transformations, restyling, and modifications to an input video while keeping its original duration and aspect ratio.
* `grok-imagine-video-extend`: xAI's video extension model. It continues an existing video, generating additional footage (2-10 seconds) beyond its end with new prompt-guided motion.
## Alibaba
### HappyHorse 1.1
* `happyhorse-1.1-t2v`: Alibaba's latest streamlined text-to-video model. It generates 720P/1080P video with native audio support, 3-15 second duration, and expanded aspect ratios (4:5, 5:4, 9:21, 21:9).
* `happyhorse-1.1-i2v`: Alibaba's latest streamlined first-frame image-to-video model. It turns a single image into high-quality 720P/1080P video with native audio support and 3-15 second duration.
* `happyhorse-1.1-r2v`: Alibaba's latest reference-image-to-video model. It uses 1-9 reference images to preserve subject or character appearance while generating new video actions, with 720P/1080P output and native audio.
## Vidu
### Q3 Drama
* `viduq3-drama`: Vidu Q3 Drama (Short Play) script-to-video model. It turns a written script plus character, scene, and prop reference images into a complete multi-shot short drama, automatically planning the shots, transitions, and camera work in a single pass.
## Vidu
### Text-to-Video
* `viduq3-turbo-t2v`: Vidu Q3 Turbo text-to-video model. Fast video generation from a text prompt with support for custom duration (1-16s), aspect ratio selection, and resolutions up to 1080p.
* `viduq3-pro-t2v`: Vidu Q3 Pro text-to-video model. High-quality video generation with enhanced prompt adherence, cinematic detail, realistic motion physics, and support for up to 1080p output.
### Image-to-Video & Transitions
* `viduq3-pro-fast-i2v`: Vidu Q3 Pro Fast image-to-video model. High-speed generation from a starting frame image, supporting motion guidance, custom duration (1-16s), and HD output.
* `viduq3-pro-i2v`: Vidu Q3 Pro image-to-video model. High-quality video generation from a starting frame, delivering superb motion fluidity and visual consistency.
* `viduq3-turbo-i2v`: Vidu Q3 Turbo image-to-video model. Speed-optimized animation from a starting frame, offering rapid generation with high-fidelity outputs.
* `viduq3-turbo-fl2v`: Vidu Q3 Turbo first-last frame model. Generates a smooth, fluid dynamic transition video between a user-defined start and end frame.
* `viduq3-pro-fl2v`: Vidu Q3 Pro first-last frame model. High-quality video transition model that bridges opening and closing frames with physical realism and precise scene progression.
### Character & Reference Video (R2V)
* `viduq3-mix-r2v`: Vidu Q3 Mix reference-to-video model. Generates rich cinematic videos with character consistency from 1-7 reference images using mixed-style synthesis.
* `viduq3-turbo-r2v`: Vidu Q3 Turbo reference-to-video model. Fast generation of character-consistent video clips utilizing up to 7 reference images and detailed prompts.
* `viduq3-r2v`: Vidu Q3 reference-to-video model. Balanced reference-to-video model designed to maintain strong character likeness and background consistency across generations.
### Digital Human & Multi-Frame Animation
* `viduq2-turbo-digital-human`: Vidu Q2 Turbo digital human model. Animates a portrait image into a speaking digital human, supporting motion prompts and audio inputs for precise lip synchronization.
* `viduq2-pro-digital-human`: Vidu Q2 Pro digital human model. High-quality portrait animation for lifelike speaking videos with expressive details and natural lip sync tied to an audio track.
* `viduq2-turbo-multi-frame`: Vidu Q2 Turbo multi-frame animation model. Directly animates through a custom sequence of up to 9 keyframes (1 start frame plus up to 8 key images).
* `viduq2-pro-multi-frame`: Vidu Q2 Pro multi-frame animation model. High-quality sequence animation that maintains strong identity and style consistency across up to 9 provided frames.
### One-Click Film & Story Creation
* `template-story`: Vidu template story model. Generates a narrative video by placing custom characters from reference images into predefined templates (e.g., love\_story, workday\_feels, monkey\_king).
* `one-click-general-film`: Vidu one-click general film model. Automatically strings together 1-7 scene or subject images with narrative prompts into cohesive cinematic films ranging from 10 to 180 seconds.
* `one-click-ad-film`: Vidu one-click ad film model. Automatically generates marketing or promotional video ads (10-60s) from 1-7 product or scene images. Support scripts in both English and Chinese.
### Video-to-Video & Editing
* `motion-sync`: Vidu motion sync model. Seamlessly extracts complex body movements and physical poses from a reference video and transfers them onto a target static character image.
* `lip-sync`: Vidu lip sync model. Reanimates lip movements in a source video to align with a new replacement audio file while keeping original characters intact.
* `one-click-trending-replicate`: Vidu one-click trending replicate model. Automatically clones a trending reference video's specific template style, timing, and transition effects over your own subject images.
## Alibaba
### Wan 2.7
* `wan2.7-t2v`: The latest generation Wan text-to-video model that generates videos from a single sentence with enhanced cinematic quality, improved motion stability, and advanced multi-shot narrative capabilities. Supports automatic dubbing and custom audio.
* `wan2.7-i2v`: The latest generation Wan image-to-video model that generates videos using prompts and image references with superior prompt adherence, cinematic quality, and advanced multi-shot narrative capabilities. Supports automatic dubbing and custom audio.
## Google
### Veo 3.1 Lite
* `Veo 3.1 Lite T2V`: Google Veo 3.1 Lite text-to-video model with person generation control. Supports resolutions up to 1080p. Duration: 4/6/8 seconds.
* `Veo 3.1 Lite I2V`: Google Veo 3.1 Lite image-to-video model with person generation control. Supports resolutions up to 1080p. Duration: 4/6/8 seconds.
## OpenAI
### Image Generation
* `gpt-image-2`: OpenAI's state-of-the-art image generation model for fast, high-quality output, supporting resolutions up to 4K (3840x2160).
* `gpt-image-1.5`: A balanced image generation model offering solid performance, transparent background support, and resolutions up to 1536x1024.
### Image Editing
* `gpt-image-2-edit`: OpenAI's state-of-the-art image editing model, supporting high-resolution outputs including 2K and 4K.
* `gpt-image-1.5-edit`: A balanced image editing model featuring transparent background support, input fidelity control, and multi-image editing capabilities (up to 16 images).
## Alibaba
### HappyHorse
* `happyhorse-1.0-t2v`: HappyHorse text-to-video model generates physically realistic and smoothly animated video content from text prompts. The model focuses on physical realism and motion fluidity, supporting various resolution and aspect ratio combinations with 3-15 seconds duration.
* `happyhorse-1.0-i2v`: HappyHorse image-to-video model generates physically realistic and smoothly animated video content from a first frame image. The model can optionally use text prompts for guidance, supporting 720P/1080P resolution and 3-15 seconds duration.
* `happyhorse-1.0-r2v`: HappyHorse reference-to-video model generates fluid videos by fusing characters from multiple reference images (1-9 images) through text prompts with character references. Supports 720P/1080P resolution, multiple aspect ratios, and 3-15 seconds duration.
* `happyhorse-1.0-video-edit`: HappyHorse video editing model supports style transformation and local replacement by combining input video with reference images (0-5) and text instructions. Input video duration: 3-60 seconds. Output video duration: 3-15 seconds.
## Kling
### Kling V3
* `Kling V3 T2I`: Kling V3 text-to-image model with improved prompt adherence and 1K/2K output support for higher-fidelity creative generation.
* `Kling V3 I2I`: Kling V3 image-to-image model for higher-fidelity editing and restyling with 1K/2K output support.
* `Kling V3 Omni Image`: Kling V3 Omni image model supporting single-image and series generation with image references, element references, and up to 4K output.
* `Kling V3 T2V`: Kling V3 text-to-video model supporting single-shot and storyboard-based multi-shot generation with 3-15 second output duration.
* `Kling V3 I2V`: Kling V3 image-to-video model supporting prompt-driven animation, storyboard workflows, element references, and 3-15 second output duration.
* `Kling V3 Omni Video`: Kling V3 Omni video model for advanced video-to-video generation, combining prompts, storyboard segments, image references, element references, and optional video inputs.
### Kling Video O1
* `Kling Video O1`: Kling Video O1 omni video model for storyboard-driven video-to-video generation, supporting reference videos, optional first/end frames, and element-guided editing.
## Alibaba
### Image Generation
* `qwen-image-2.0`: Accelerated text-to-image model balancing quality and speed, supporting resolutions up to 2688\*1536 and batch generation of 1-6 images.
* `qwen-image-2.0-pro`: The most capable Qwen-Image 2.0 model with stronger text rendering, realistic textures, and semantic adherence, supporting batch generation of 1-6 images.
* `qwen-image-max`: High-realism text-to-image model with reduced AI artifacts, fixed resolution options, and 1 image output per request.
* `wan2.7-image`: Standard text-to-image model with faster generation speed, up to 2K resolution, thinking mode, sequential generation, and custom color themes.
* `wan2.7-image-pro`: Professional text-to-image model supporting up to 4K resolution, thinking mode, sequential multi-image generation, and custom color themes.
* `z-image-turbo`: Lightweight text-to-image model optimized for speed with Chinese and English text rendering support, outputting 1 PNG image per request.
### Image Editing
* `qwen-image-2.0-edit`: Accelerated image editing model balancing quality and speed, supporting 1-3 input images and 1-6 outputs with customizable resolution.
* `qwen-image-2.0-pro-edit`: Professional image editing model with stronger text rendering, realistic quality, and semantic following, supporting 1-3 input images and 1-6 outputs.
* `qwen-image-edit-max`: Advanced image editing model focused on industrial design, geometry reasoning, and character consistency, supporting 1-3 input images and 1-6 outputs.
* `wan2.7-image-edit`: Standard image editing model with faster generation speed, supporting multi-image reference, interactive editing, sequential generation, and max 2K output.
* `wan2.7-image-pro-edit`: Professional image editing model supporting multi-image reference, interactive bounding-box editing, sequential generation, 1-9 input images, and max 2K output.
### Video Generation & Editing
* `wan2.1-vace-plus`: Unified video editing model supporting five functions: multi-image reference, video repainting, local editing, video extension, and video outpainting.
* `wan2.2-animate-mix`: Video character replacement model that swaps the main character with a reference image while preserving scene, lighting, and color tone.
* `wan2.2-animate-move`: Motion transfer model that applies movement and expressions from a reference video to a character in a static image.
* `wan2.6-r2v`: Reference-to-video model that generates from reference images or videos with multi-character interaction and role-playing, producing silent video by default.
* `wan2.6-r2v-flash`: Faster reference-to-video model with audio or silent output switching, optimized for quick previews and cost-effective generation.
* `wan2.7-r2v`: Advanced reference-to-video model using reference images/videos with prompt guidance, supporting multi-subject references, storyboard generation, and custom audio voice cloning.
* `wan2.7-videoedit`: Multi-modal video editing model for style modification and edits with text/image/video inputs, with typical processing time of 1-5 minutes.
## ByteDance
### Seedream 5.0
* `seedream-5.0-lite`: ByteDance Seedream 5.0 Lite text-to-image model with 2K/3K custom resolutions and configurable output format.
* `seedream-5.0-lite-edit`: ByteDance Seedream 5.0 Lite Edit image-to-image model supporting single-image editing, multi-image fusion, and configurable output format.
### Seedance 2.0
* `seedance-2.0-i2v`: Dreamina Seedance 2.0 image-to-video—generate from a text prompt with optional frame, image, and audio references.
* `seedance-2.0-fast-i2v`: Faster Seedance 2.0 image-to-video with the same request parameters.
* `seedance-2.0-v2v`: Dreamina Seedance 2.0 video-to-video—transform 1–3 reference clips with a text prompt and optional image or audio references.
* `seedance-2.0-fast-v2v`: Faster Seedance 2.0 video-to-video with the same request parameters.
## Google
### Nano Banana
* `Nano Banana`: Nano Banana image generation model. Generates images from text prompts with support for 10 aspect ratios, delivering fast and cost-effective results.
* `Nano Banana Pro`: Nano Banana Pro image generation model with higher quality output. Supports aspect ratio selection and output resolutions up to 4K.
* `Nano Banana 2`: Nano Banana 2 multimodal image model supporting both text-to-image and image-to-image workflows, with 14 aspect ratios and resolutions from 512 to 4K.
* `Nano Banana Edit`: Nano Banana image editing model. Transforms existing images based on prompt instructions with support for 10 aspect ratios.
* `Nano Banana Pro Edit`: Nano Banana Pro image editing model with superior detail preservation and enhanced prompt adherence. Supports output resolutions up to 4K.
* `Nano Banana 2 Edit`: Nano Banana 2 multimodal model in image-to-image editing mode. Requires a base64 data URI image input with support for multiple aspect ratios and resolutions from 512 to 4K.
### Imagen 4.0
* `Imagen 4.0`: Google Imagen 4.0 standard text-to-image model delivering high-quality photorealistic images. Supports batch generation (up to 4 images), person generation control, and output resolutions up to 2K.
* `Imagen 4.0 Ultra`: Google Imagen 4.0 Ultra text-to-image model with the highest quality output. Optimized for maximum detail and photorealism with batch generation and up to 2K resolution.
* `Imagen 4.0 Fast`: Google Imagen 4.0 Fast text-to-image model optimized for speed. Supports batch generation (up to 4 images), multiple aspect ratios, and person generation control.
### Veo 3.1
* `Veo 3.1 T2V`: Google's flagship text-to-video model supporting resolutions up to 4K and optional reference images (up to 3) for style or character consistency across generations. Duration: 4/6/8 seconds.
* `Veo 3.1 Fast T2V`: A faster variant of Veo 3.1 with the same capabilities, including 4K resolution and reference image support. Duration: 4/6/8 seconds.
* `Veo 3.1 I2V`: Google's flagship image-to-video model supporting resolutions up to 4K with person generation control. Duration: 4/6/8 seconds.
* `Veo 3.1 Fast I2V`: A faster variant of Veo 3.1 I2V with the same capabilities, including 4K resolution support. Duration: 4/6/8 seconds.
### Veo 3
* `Veo 3 T2V`: Google Veo 3.0 stable text-to-video model with resolutions up to 1080p and person generation control. Duration: 4/6/8 seconds.
* `Veo 3 Fast T2V`: A faster variant of Veo 3 with the same parameter set, supporting resolutions up to 1080p. Duration: 4/6/8 seconds.
* `Veo 3 I2V`: Google Veo 3.0 stable image-to-video model with resolutions up to 1080p and person generation control. Duration: 4/6/8 seconds.
* `Veo 3 Fast I2V`: A faster variant of Veo 3 I2V with the same parameter set, supporting resolutions up to 1080p. Duration: 4/6/8 seconds.
### Veo 2
* `Veo 2 T2V`: Google Veo 2.0 classic text-to-video model with flexible person generation policies (allow all, adult only, or disallow). Duration: 5/6/8 seconds.
* `Veo 2 I2V`: Google Veo 2.0 classic image-to-video model with flexible person generation policies (allow adult or disallow). Duration: 5/6/8 seconds.
## Kling
### Kling V1
* `Kling V1 T2I`: Kuaishou's foundational text-to-image model offering fast, cost-effective 1K image generation with strong prompt adherence and multiple aspect ratios.
* `Kling V1 I2I`: Kuaishou's first-generation AI image model using a diffusion transformer architecture, capable of generating 1K-resolution images with strong prompt adherence and realistic detail.
* `Kling V1 T2V`: Kuaishou's first-generation text-to-video model generating 5s or 10s clips with camera motion presets (pan, tilt, zoom) and adjustable prompt relevance.
* `Kling V1 I2V`: Kuaishou's first-generation image-to-video model that animates static images into 5s or 10s videos with motion brush support and adjustable prompt relevance.
### Kling V1.5
* `Kling V1.5 T2I`: An enhanced text-to-image model with improved realism and subject/face reference support for generating consistent character images at 1K resolution.
* `Kling V1.5 I2I`: An upgraded image-to-image model with improved realism, better prompt interpretation, and subject/face reference modes for precise character control.
* `Kling V1.5 I2V`: The most feature-complete V1.x image-to-video model, adding simple camera motion control alongside motion brush and cfg\_scale for precise video generation.
### Kling V1.6
* `Kling V1.6 T2V`: An improved text-to-video model with significantly better prompt adherence and visual quality over V1.5, supporting dual standard/professional generation modes.
* `Kling V1.6 I2V`: An improved image-to-video model with significantly better prompt adherence and visual quality over V1.5, supporting first-and-last frame control for smooth transitions.
* `Kling V1.6 MI2V`: Transforms up to 4 reference images into a cohesive video sequence with multi-element fusion, enabling character interaction and complex visual narratives.
### Kling V2
* `Kling V2 T2I`: A next-generation text-to-image model with significantly improved detail and visual fidelity, supporting both 1K and 2K resolutions for professional output.
* `Kling V2 New T2I`: A refined variant of V2 with updated model weights for sharper details, better consistency, and improved prompt-to-image alignment at up to 2K resolution.
* `Kling V2 I2I`: A major generational leap in image quality and creativity, featuring enhanced style diversity and significantly improved visual fidelity over V1.5.
* `Kling V2 New I2I`: A refined variant of Kling V2 with updated model weights for improved consistency, sharper details, and better prompt-to-image alignment.
* `Kling V2 MI2I`: Combines up to 4 subject images with optional scene and style references into a single cohesive output, supporting subject fusion, scene replacement, and style transfer.
* `Kling V2 Master T2V`: The V2-generation base text-to-video model producing cinematic-quality clips with superior motion realism and temporal coherence.
* `Kling V2 Master I2V`: The V2-generation base image-to-video model delivering cinematic-quality animations with superior temporal coherence and smoother motion transitions.
### Kling V2.1
* `Kling V2.1 T2I`: The latest and highest-quality text-to-image model in the Kling family, delivering state-of-the-art results at up to 2K resolution.
* `Kling V2.1 I2I`: The latest cost-efficient image-to-image model offering studio-grade quality with faster rendering and excellent prompt adherence.
* `Kling V2.1 MI2I`: The latest and highest-quality multi-image composition model, delivering superior subject fusion, scene replacement, and style transfer with up to 4 subject images.
* `Kling V2.1 Master T2V`: The V2.1-generation text-to-video model with enhanced rendering quality, improved frame consistency, and studio-grade 1080p output.
* `Kling V2.1 I2V`: A cost-efficient image-to-video model with advanced frame control and up to 1080p output, suitable for professional content creation.
* `Kling V2.1 Master I2V`: The recommended high-quality image-to-video model in the V2.1 series, producing studio-grade 1080p videos with precise start and end frame control.
### Kling V2.5
* `Kling V2.5 Turbo T2V`: A speed-optimized text-to-video model delivering cinematic 1080p videos with physics-accurate motion at \~30% lower cost than previous versions.
* `Kling V2.5 Turbo I2V`: A speed-optimized image-to-video model delivering cinematic 1080p videos with physics-accurate motion at \~30% lower cost than previous versions.
### Kling V2.6
* `Kling V2.6 T2V`: The first Kling text-to-video model to natively generate synchronized audio and video in one pass, including dialogue, ambient sounds, and lip-synced speech.
* `Kling V2.6 I2V`: The first Kling image-to-video model to natively generate synchronized audio and video in a single pass, supporting dialogue, sound effects, and lip-synced speech alongside motion brush.
### Kling Avatar & Effects
* `Kling Avatar`: Generates realistic talking-head videos from a reference image and audio input, with precise lip synchronization, expressive gestures, and support for multiple languages.
* `Kling Video Effects`: Applies 212 preset creative video effects -- including dance, transformation, interaction, and animation styles -- to one or two person images for instant viral content.
### Kling Image Utilities
* `Kling Image Expansion`: Intelligently extends images in any direction (up, down, left, right) with prompt-guided content generation, ideal for panorama creation, background extension, and canvas expansion.
* `Kling Image Recognize`: Detects and segments image content into 4 categories -- object, head (with hair), face (without hair), and clothing -- returning segmentation masks synchronously.
* `Kling Image O1`: A multimodal image generation model that accepts text, up to 10 reference images, and element inputs to produce 1K/2K images with precise style control and multi-reference feature extraction.
### Kolors Virtual Try-On
* `Kolors Virtual Try-On V1`: AI-powered virtual clothing try-on built on the Kolors diffusion model, generating realistic fitting results from a person photo and a single garment image (tops, bottoms, or dresses).
* `Kolors Virtual Try-On V1-5`: Enhanced virtual try-on model that supports both single garments and top+bottom outfit combinations, delivering higher-quality results with automatic clothing type detection.
## MiniMax
### MiniMax Image-01
* `MiniMax Image-01 T2I`: MiniMax's multimodal vision model that blends text-to-image generation with visual reasoning for seamless cross-modal tasks.
* `MiniMax Image-01 I2I`: MiniMax's multimodal vision model that blends text-to-image generation with visual reasoning for seamless cross-modal tasks.
* `MiniMax Image-01-Live I2I`: MiniMax's multimodal vision model that blends text-to-image generation with visual reasoning for seamless cross-modal tasks.
### Hailuo
* `Hailuo 2.3 T2V`: Hailuo 2.3 not only generates high-quality videos from text or images with exceptional instruction following, but also redefines realism through its state-of-the-art mastery of extreme physics.
* `Hailuo 2.3 I2V`: Hailuo 2.3 not only generates high-quality videos from text or images with exceptional instruction following, but also redefines realism through its state-of-the-art mastery of extreme physics.
* `Hailuo 2.3 Fast I2V`: Hailuo 2.3 Fast efficiently transforms images into dynamic videos with extreme physics mastery. It delivers exceptional value by generating high-quality, realistic motion at a reduced computational cost.
* `Hailuo 02 T2V`: Hailuo 02 masters both text-to-video and image-to-video generation with exceptional instruction following, while setting a new standard in visual realism through its extreme physics simulation.
* `Hailuo 02 I2V`: Hailuo 02 masters both text-to-video and image-to-video generation with exceptional instruction following, while setting a new standard in visual realism through its extreme physics simulation.
* `Hailuo 02 FL2V`: Hailuo 02's FL2V function provides unprecedented creative control by generating dynamic videos between a user-defined start and end frame. This feature not only masters extreme physics and complex transitions but also enables the novel capability to deduce a story leading up to a specified final image.
### MiniMax T2V-01
* `MiniMax T2V-01`: MiniMax T2V-01 is a text-to-video model that uniquely delivers professional-level camera movement control, transforming written prompts into cinematic video clips with dynamic shots.
* `MiniMax T2V-01-Director`: T2V-01-Director is a text-to-video AI model that offers precise camera control, allowing users to create professional-looking video clips with cinematic movements through a variety of lens instructions.
### MiniMax I2V-01
* `MiniMax I2V-01`: MiniMax I2V-01 is a foundational image-to-video model that converts static pictures into high-quality video sequences, delivering smooth animation especially optimized for illustrations and anime styles.
* `MiniMax I2V-01-Director`: T2V-01-Director is a text-to-video AI model that offers precise camera control, allowing users to create professional-looking video clips with cinematic movements through a variety of lens instructions.
* `MiniMax I2V-01-Live`: I2V-01-Live is an image-to-video model specifically optimized for animating 2D illustrations and cartoon styles, enhancing smoothness and vivid motion to bring static art to life with fluid character movements and natural expressions.
### MiniMax S2V-01
* `MiniMax S2V-01`: The MiniMax S2V-01 is a specialized subject reference video model designed to solve the industry challenge of character consistency. It can generate dynamic videos where the main character's identity stays highly consistent across every frame, using just a single photo as a reference and at a computational cost significantly lower than traditional solutions.
## Alibaba
### Wan
* `Wan 2.6 T2V`: The Wan text-to-video model can generate videos from a single sentence, presenting rich artistic styles and cinematic quality. Wan 2.6 introduces multi-shot narrative capabilities and supports both automatic dubbing and uploading custom audio files.
* `Wan 2.6 I2V Flash`: The Wan image-to-video model can generate videos using prompts and image references, featuring rich artistic styles and cinematic quality. Wan 2.6 introduces multi-shot narrative capabilities and supports both automatic dubbing and uploading custom audio files.
* `Wan 2.6 I2V`: The Wan image-to-video model can generate videos using prompts and image references, featuring rich artistic styles and cinematic quality. Wan 2.6 introduces multi-shot narrative capabilities and supports both automatic dubbing and uploading custom audio files.
* `Wan 2.5 T2V Preview`: The Wan text-to-video model can generate videos from a single sentence, presenting rich artistic styles and cinematic quality. Wan 2.5 supports automatic dubbing and uploading custom audio files.
* `Wan 2.5 I2V Preview`: The Wan image-to-video model can generate videos using prompts and image references, featuring rich artistic styles and cinematic quality. Wan 2.5 supports automatic dubbing and uploading custom audio files.
* `Wan 2.2 T2V Plus`: The Wan text-to-video model can generate videos from a single sentence, presenting rich artistic styles and cinematic quality. Wan 2.2 features more accurate instruction understanding, stable and smooth motion generation, and richer details.
* `Wan 2.2 I2V Flash`: The Wan image-to-video model can generate videos using prompts and image references, presenting rich artistic styles and cinematic-quality visuals. Wan 2.2 Flash features ultimate generation speed, with more accurate instruction understanding and camera control, consistent visual elements, and comprehensively improved stability and success rates.
* `Wan 2.2 I2V Plus`: The Wan image-to-video model can generate videos using prompts and image references, presenting rich artistic styles and cinematic-quality visuals. Wan 2.2 Plus features more accurate instruction understanding, controllable camera movements, consistent visual elements, and comprehensively improved stability and success rates, delivering richer generated content.
* `Wan 2.2 KF2V Flash`: The Wan First-and-Last-Frame Video Generation Model: simply provide the first and last frame images, and it can generate a smooth, fluid dynamic video based on the prompt.
### Wanx
* `Wanx 2.1 T2V Turbo`: Wan text-to-video model can generate videos with a single sentence, featuring rich artistic styles and cinematic quality. Wanx 2.1 Turbo offers high cost-effectiveness.
* `Wanx 2.1 T2V Plus`: Wan text-to-video model can generate videos from a single sentence, featuring rich artistic styles and cinematic quality. Wanx 2.1 Plus offers even more refined visuals.
* `Wanx 2.1 I2V Plus`: The Wan image-to-video model can generate videos using prompts and image references, presenting rich artistic styles and cinematic-quality visuals. Wanx 2.1 Plus offers even more refined image quality.
* `Wanx 2.1 I2V Turbo`: The Wan image-to-video model can generate videos using prompts and image references, featuring rich artistic styles and cinematic-quality visuals. Wanx 2.1 Turbo offers high cost-effectiveness.
* `Wanx 2.1 KF2V Plus`: The Wan First-and-Last-Frame Video Generation Model: simply provide the first and last frame images, and it can generate a smooth, fluid dynamic video based on the prompt.
## ByteDance
### Seedream
* `Seedream 3.0 T2I`: Seedream 3.0 is a Chinese-English bilingual image generation foundation model that supports native high resolution. Its overall capabilities are comparable to GPT-4o, ranking it among the world's top tier. Faster response speed; more accurate small text generation and enhanced text typesetting effect; strong instruction-following ability, improved aesthetics & structure, and good fidelity and detail performance.
* `Seedream 4.0 T2I`: A SOTA-level multimodal image creation model based on a leading architecture. It breaks the creative boundaries of traditional text-to-image models and natively supports text, single-image, and multi-image inputs. Users can freely fuse text and images, and in the same model, realize diverse applications like multi-image fusion creation based on subject consistency, image editing, and group image generation.
* `Seedream 4.0 I2I`: A SOTA-level multimodal image creation model based on a leading architecture. It breaks the creative boundaries of traditional text-to-image models and natively supports text, single-image, and multi-image inputs. Users can freely fuse text and images, and in the same model, realize diverse applications like multi-image fusion creation based on subject consistency, image editing, and group image generation.
* `Seedream 4.5 T2I`: Seedream 4.5 is the latest in-house image generation model developed by ByteDance. Compared with Seedream 4.0, it delivers comprehensive improvements—especially in editing consistency, including better preservation of subject details, lighting, and color tone. It also enhances portrait refinement and small-text rendering. The model's multi-image composition capabilities have been significantly strengthened.
* `Seedream 4.5 I2I`: Seedream 4.5 is the latest in-house image generation model developed by ByteDance. Compared with Seedream 4.0, it delivers comprehensive improvements—especially in editing consistency, including better preservation of subject details, lighting, and color tone. It also enhances portrait refinement and small-text rendering. The model's multi-image composition capabilities have been significantly strengthened.
### Seededit
* `Seededit 3.0 I2I`: SeedEdit 3.0 is an image editing model that supports editing images via text instructions. SeedEdit 3.0 is trained based on the text-to-image model Seedream 3.0, integrated with diverse data fusion methods and specific reward models. Its ability to preserve image subjects, backgrounds, and details has been further improved, especially in scenarios such as portrait editing, background modification, perspective and light conversion.
### Seedance
* `Seedance 1.0 Lite T2V`: ByteDance's small-parameter version of the video generation model achieves excellent video generation quality while significantly increasing generation speed, balancing both effect and efficiency.
* `Seedance 1.0 Lite I2V`: ByteDance's small-parameter version of the video generation model achieves excellent video generation quality while significantly increasing generation speed, balancing both effect and efficiency.
* `Seedance 1.0 Pro Fast T2V`: Seedance 1.0 pro fast, inheriting the core advantages of the Seedance 1.0 pro model, has a 3x faster generation speed and a 72% lower price. It is a video generation model that achieves an excellent balance among quality, speed, and cost.
* `Seedance 1.0 Pro Fast I2V`: Seedance 1.0 pro fast, inheriting the core advantages of the Seedance 1.0 pro model, has a 3x faster generation speed and a 72% lower price. It is a video generation model that achieves an excellent balance among quality, speed, and cost.
* `Seedance 1.0 Pro T2V`: Seedance 1.0 is a video generation foundation model launched by ByteDance. As the large-parameter version of this model series, Seedance 1.0 Pro has unique multi-shot narrative capabilities and performs excellently across all dimensions. It has made breakthroughs in semantic understanding and instruction-following capabilities, and can generate 1080P high-definition videos that are smooth in motion, rich in details, diverse in style, and have cinematic-level aesthetics.
* `Seedance 1.0 Pro I2V`: Seedance 1.0 is a video generation foundation model launched by ByteDance. As the large-parameter version of this model series, Seedance 1.0 Pro has unique multi-shot narrative capabilities and performs excellently across all dimensions. It has made breakthroughs in semantic understanding and instruction-following capabilities, and can generate 1080P high-definition videos that are smooth in motion, rich in details, diverse in style, and have cinematic-level aesthetics.
* `Seedance 1.5 Pro T2V`: Seedance 1.5 pro is ByteDance's new professional-grade audio-visual co-generation model. It builds on multi-shot narrative and HD generation capabilities, supporting integrated audio and video output for a unified creation experience (visuals, human voice, music, and sound effects). The model includes a start/end frame feature, allowing creators to lock the video's style, composition, and characters by setting the first and last frames.
* `Seedance 1.5 Pro I2V`: Seedance 1.5 pro is ByteDance's new professional-grade audio-visual co-generation model. It builds on multi-shot narrative and HD generation capabilities, supporting integrated audio and video output for a unified creation experience (visuals, human voice, music, and sound effects). The model includes a start/end frame feature, allowing creators to lock the video's style, composition, and characters by setting the first and last frames.
# Alibaba
### Wan
### Wanx
* `Wanx 2.1 T2I Turbo`: The Wan text-to-image model generates beautiful images from text. Supports multiple styles and generates quickly.
* `Wanx 2.1 T2I Plus`: The Wan text-to-image model generates beautiful images from text. Supports multiple styles and generates images with rich details.
* `Wanx 2.1 Image Edit`: Can achieve diverse image editing through simple instructions, suitable for scenarios such as image expansion, watermark removal, style transfer, image restoration, and image enhancement.
* `Wanx 2.0 T2I Turbo`: The Wan text-to-image model excels in textured portraits and creative design, offering great value for money.
* `Wanx Style Repaint V1`: Can perform various stylized redraws on input portrait images, allowing the newly generated images to maintain the original facial features while presenting different artistic painting effects.
* `Wanx Sketch to Image Lite`: Based on input hand-drawn sketches and text descriptions, exquisite doodle artworks can be generated.
* `Wanx Background Generation V2`: Can expand and generate background information based on input foreground image materials, achieving natural light and shadow fusion effects, as well as delicate and realistic image generation.
### Qwen
* `Qwen Image Plus`: The qwen-image excels in text rendering, particularly for Chinese text. Currently more cost-effective than qwen-image.
* `Qwen Image`: The qwen-image excels in text rendering, particularly for Chinese text.
* `Qwen Image Edit Plus`: Supports precise bilingual Chinese-English text editing, color adjustment, detail enhancement, style transfer, object addition and removal, and other operations, enabling complex image and text editing.
* `Qwen Image Edit Plus 2025-12-15`: Supports precise bilingual Chinese-English text editing, color adjustment, detail enhancement, style transfer, object addition and removal, and other operations, enabling complex image and text editing.
* `Qwen Image Edit Plus 2025-10-30`: Supports precise bilingual Chinese-English text editing, color adjustment, detail enhancement, style transfer, object addition and removal, and other operations, enabling complex image and text editing.
* `Qwen Image Edit`: Supports precise bilingual Chinese-English text editing, color adjustment, detail enhancement, style transfer, object addition and removal, and other operations, enabling complex image and text editing.
* `Qwen MT Image`: Supports translating text from images in 11 languages into Chinese or English, accurately preserving original layout and content information, and provides custom features such as terminology definitions, sensitive word filtering, and image subject detection.
### WordArt
* `WordArt Semantic`: Can creatively deform the edge contours of input text based on prompt content, achieving more creative uses of a font, and returns a black-background white mask image containing the text.
* `Wordart Texture`: Can perform creative design on input text content or text images, adding materials and textures to the text based on prompt content to achieve effects such as 3D prominence or scene integration.
### AI Try-On
* `AI Try-On`: A virtual try-on image generation model that generates try-on images based on portrait photos and clothing images.
* `AI Try-On Plus`: Compared to the AI Try-On, there are improvements in image clarity, clothing texture details, and logo restoration effects, but the generation time is longer.
* `AI Try-On Parsing V1`: Supports segmentation of model images and clothing images, and can be used for pre-processing and post-processing of AI fitting room images.
* `AI Try-On Refiner`: Perform secondary generation on the effect images created by AI virtual try-on, outputting finely polished virtual try-on effect images with higher fidelity.
### Image Utilities
* `Image Outpainting`: Allows for free image extension, supporting image rotation and expansion through both expansion coefficient and pixel count methods.
# Modellix Product Updates and Announcements
Source: https://docs.modellix.ai/changelog/product-updates
Follow Modellix product updates and announcements—new API endpoints, console features, billing changes, and platform improvements released each month.
## Modellix CLI Get Schema
[`modellix-cli`](https://www.npmjs.com/package/modellix-cli) now includes **`model get-schema`**.
Use the exact `provider/model` slug from `model list`. The command calls the public [Get Schema](/api/get-schema) endpoint and does not require an API key. JSON is the default; `--output human` summarizes the contract; `--quiet` prints the inference URL.
```bash theme={null}
npm install --global modellix-cli@latest
modellix-cli model get-schema alibaba/qwen-image-3.0-pro
```
For full usage details, see the [CLI documentation](/ways-to-use/cli).
## Skill and Plugin Schema Lookup
The [Modellix Skill](/ways-to-use/skill) and [Plugin](/ways-to-use/plugin) now read request-body schemas with `model get-schema` instead of guessing fields from documentation pages.
The agent fetches the schema when you name a non-default model, when the body needs extra fields, or after HTTP `400`. It skips the call for documented default examples or a complete body you already supplied. If the CLI is unavailable, it falls back to Docs MCP, `docs_url` from `model describe`, or [llms.txt](https://docs.modellix.ai/llms.txt).
## DeepSeek Harness Plugin
**[dsh-modellix](/ways-to-use/deepseek-harness)** is now available for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness).
Install the plugin in the Harness Web profile, then connect one Modellix API key. The plugin loads the live LLM catalog into the model selector, runs chat-first media generation with a session result panel, and lets the Agent call Web Search and Web Fetch when a question needs the public web.
Source: [github.com/Modellix/dsh-modellix](https://github.com/Modellix/dsh-modellix)
## Signup Credit Ended
New registrations no longer receive a complimentary **\$1** credit.
To request trial credit, contact the Modellix team at [support@modellix.ai](mailto:support@modellix.ai).
## Web Search and Web Fetch
Modellix now provides **Web Search** and **Web Fetch** at `https://tool.modellix.ai`. Authenticate with your Modellix API Key.
### Web Search
**[Web Search](/api/web-search)** (`POST /v1/web-search`) searches the public web and returns ranked results. Choose depth (`lite`, `standard`, or `rich`) and optionally filter by domain, time range, topic, or country.
### Web Fetch
**[Web Fetch](/api/web-fetch)** (`POST /v1/web-fetch`) extracts readable content from public HTTP or HTTPS URLs (up to 20 per request). Failed URLs are returned separately.
Query Web Search and Web Fetch history with **[Get Tool Logs](/api/get-tool-logs)** (`GET /v1/logs`).
## Modellix Agent Canvas Plugin
We launched **[Modellix Agent Canvas](/ways-to-use/agent-canvas)**, a local workspace-bound `stdio` MCP plugin for visual AI work.
Use an Excalidraw infinite canvas with Modellix image generation and editing, paid-operation confirmation, durable task recovery, HTML drafts, and presentations—without a deployed Canvas service. Install from the host plugin marketplace or MCP config for Codex, Cursor, Claude Code, OpenCode, and other local MCP hosts.
Source: [github.com/Modellix/modellix-agent-canvas](https://github.com/Modellix/modellix-agent-canvas)
## Japanese UI Language Support
The Modellix website and console UI now support **Japanese**.
Open the Japanese locale at [modellix.ai/ja\_JP](https://www.modellix.ai/ja_JP) to browse product pages and the console in Japanese.
ようこそ、日本のユーザーの皆さま。Modellix を日本語でお使いいただけます。ぜひご利用ください。
## Online Invoice Export
You can now export invoices online from the console.
On the [Orders](https://www.modellix.ai/console/billing/order) page, open any order and download its invoice directly—no manual request required.
## Modellix CLI Upgrade
[`modellix-cli`](https://www.npmjs.com/package/modellix-cli) has been upgraded with a fuller command set for terminal, scripting, and agent workflows.
```bash theme={null}
npm install --global modellix-cli
```
### What's New
* **Authentication profiles**: `init`, `auth login/status/whoami/logout`, and named profiles with `config` inspection and cleanup.
* **Environment checks**: `doctor` verifies Node.js, API key resolution, connectivity, and balance.
* **Model discovery**: `model list` and `model describe` to find and inspect model slugs.
* **Model execution**: `model run` (with `model invoke` as a compatible alias), `--wait`, stdin bodies, and `model batch` for JSONL submissions.
* **Task tooling**: `task get`, `task wait`, `task download`, and local `task history`.
* **Automation-ready output**: `--json`, `--quiet`, stable exit codes, and CI-friendly flags.
For full usage details, see the [CLI documentation](/ways-to-use/cli).
## File API
We are excited to launch the **File API**, a free, API-Key-authenticated way to upload media assets and feed them directly into Modellix prediction APIs — no need to host the files yourself.
### What's New
* **[Upload Media File](/api/upload-media-file)** (`POST /api/v1/media/files`): Upload a single image, video, or audio file via `multipart/form-data`. The response returns a `file_id`, media `type`, and a `download_url` you can pass into prediction API input fields such as `image_url`.
* **[List Media Files](/api/list-media-files)** (`GET /api/v1/media/files`): List the non-expired files belonging to your team with `limit` / `offset` pagination (max page size 100).
* **[Delete Media File](/api/delete-media-file)** (`DELETE /api/v1/media/files/{file_id}`): Delete a file owned by your team to free upload quota immediately.
### Highlights
* **Free of charge**: Uploads are not billed and are not subject to balance freeze (admission).
* **Secure**: Files are scoped to your team — there is no cross-team leakage.
* **Time-boxed retention**: Uploaded files are retained for about **7 days** by default.
* **Flexible formats**: Supports common image, video, and audio extensions (e.g., `png`, `mp4`, `mp3`), with a default max file size of **16 MB**.
To get started, see [Upload media files](/ways-to-use/api#upload-media-files) for the full workflow, limits, and supported formats.
## One-Click Media Export from Playground
We are excited to launch a powerful export feature in the Modellix Playground, allowing you to integrate generated media directly into your external storage or automated workflows.
### Supported Export Destinations
Instead of downloading files locally and manually uploading them elsewhere, you can now transfer your generated images, videos, and audio directly to third-party platforms:
* **Webhook URL**: Deliver the media payload directly to your custom Webhook endpoint to instantly trigger automated backend pipelines.
* **Google Drive**: Save assets straight to your Google Drive folders for seamless cloud storage and sharing.
* **Dropbox**: Send media assets directly to your Dropbox workspace with a single click.
To get started, generate any media in the Playground and click the **Export** button in the result panel to select your preferred destination.
## Webhooks and Notification Settings
We are excited to announce two major updates that improve integration automation and account monitoring:
### API Webhook Support
You can now configure a Webhook URL to receive task results asynchronously once your media generation tasks reach a terminal state.
* **Asynchronous Delivery**: No more polling! Include the `X-Webhook-URL` header in your prediction creation requests, and Modellix will automatically deliver the payload to your endpoint.
* **Detailed Configuration**: For full setup instructions, URL requirements, retry rules, and event payload schemas, please check the [REST API Webhooks documentation](/ways-to-use/api#webhooks).
### Console Notification Settings
You can now configure notification alerts directly in the console to stay on top of your workspace events:
* **Low Balance Alerts**: Set up email alerts to notify you when your team balance drops below your specified threshold, helping you prevent service interruptions.
* **Configure Now**: Set up your notification preferences on the [Team Notification Settings](https://www.modellix.ai/console/team/notification) page.
## Common API Updates
We have expanded our suite of utility endpoints under the **Common API** to improve key management, billing transparency, and model discovery:
### New Utility Endpoints
* **[Validate API Key](/api/validate-api-key)** (`GET /api/v1/apikey/validate`): A lightweight endpoint to programmatically verify whether an API Key is currently valid and active on the Modellix platform.
* **[Get Team Balance](/api/get-team-balance)** (`GET /api/v1/team/balance`): Retrieve the current available balance (in USD with 4 decimal places) for the team workspace associated with the API Key.
### List Active Models Upgrade
* **Model Descriptions**: The [List Active Models](/api/list-models) endpoint (`GET /api/v1/models`) now returns a `description` field for each model dynamically, providing a quick summary of the model's capabilities and use cases directly in the JSON response.
## First Top-Up Discount, Concurrency Entitlements, and Account Deletion
### First Top-Up: 10% Off Every Tier
Each top-up option now includes a **10% discount on your first purchase** of that tier. Visit the [Top Up](https://www.modellix.ai/console/billing/top-up) page in the console to see available amounts and apply the discount at checkout.
### Concurrency Entitlements
When a **single top-up** reaches a qualifying amount, your workspace unlocks higher **concurrent task** and **RPM** limits. Thresholds and current entitlements are listed on the [Team Entitlements](https://www.modellix.ai/console/team/entitlements) page.
### Self-Service Account Deletion
You can now **delete your account** from the console without contacting support. Open your account settings and follow the deletion flow when you are ready to close your Modellix account.
We hope you stick around — but if you ever need to leave, you are in full control.
## Explore Filtering and Playground Shortcuts
### Model Explore: Series and Collections Filtering
The [Model Explore](https://www.modellix.ai/explore) page now supports filtering models by **series** and **collections**, making it easier to browse related models and discover what you need faster.
### Playground Shortcuts
The Model Playground now includes a set of quick actions to streamline your workflow:
* **Copy model details**: Copy the model's Markdown documentation content or direct URL with one click.
* **Copy Modellix Skill install command**: Instantly copy the `npx skills add` command for the current model.
* **One-click Skill installation**: Install the Modellix Skill directly into Manus, Codex, Claude Cowork, or Cursor.
* **One-click AI chat**: Open a conversation about the current model in ChatGPT, Claude, Grok, and other supported assistants.
* **One-click MCP installation**: Install the Modellix MCP server into Cursor, Devin, or Windsurf.
## Pricing Page, Model Search, Dashboard, and Request Logs
### Public Pricing Page
We have launched a dedicated [pricing page](https://www.modellix.ai/pricing) where pricing for every model supported on Modellix is publicly available. You can search and look up the price of any model directly on the page.
### Improved Model Search
Model search has been upgraded to deliver faster, more accurate results when you browse and discover models on Modellix.
### Dashboard Improvements
The Dashboard now presents clearer, more detailed information and guidance. We have also added a convenient top-up entry so you can add funds without leaving your workflow.
### Request Log Input Parameters
Request logs now record input parameters for each API call. You can view these parameters directly in the log details to help with debugging and auditing.
## Simplified API Endpoints
We have updated all model API endpoints to provide a cleaner and more intuitive developer experience. The `/async` suffix has been removed from all model invocation paths.
For example, `POST /alibaba/qwen-image-edit-plus/async` is now simply `POST /alibaba/qwen-image-edit-plus`.
**Backwards Compatibility**: If you are currently using the `/async` endpoints, your integrations will not be affected. We have implemented full backwards compatibility, so existing applications will continue to function normally without requiring immediate updates.
## Models API
We have introduced a new utility endpoint to help developers programmatically discover model availability on Modellix:
* **List Active Models**: A new `GET /api/v1/models` endpoint that retrieves a rich JSON array of all currently active (published) models.
### Key Fields Explained
Each model entry in the returned payload contains:
* `slug`: The unique model identifier in `provider/model_id` format (e.g., `google/nano-banana-2-edit`).
* `type`: The functional category of the model, supporting `text-to-image`, `image-to-image`, `text-to-video`, `image-to-video`, and `video-to-video`.
* `docs_url`: Direct fully qualified URI to the model's detailed developer integration documentation page.
For integration examples and the full OpenAPI schema spec, please head over to the [List Active Models API reference](/api/list-models).
## Banner Notifications and Team Settings
* **Banner Notifications**: You can now receive important updates and notifications via a banner displayed at the top of the page.
* **Team Settings**: You can now configure basic information (like Team Name) and tax details for your workspace team, ensuring compliance and proper invoicing.
## Explore and Playground Features Added
Modellix now supports Model Explore and Playground:
* **Explore**: Users can search and query all models supported by Modellix.
* **Playground**: Users can directly experience all models supported by Modellix.
Additionally, in this update, the pricing for all models has been explicitly published.
## Modellix CLI
[`modellix-cli`](https://www.npmjs.com/package/modellix-cli) is now available -- the official command-line tool for Modellix.
```bash theme={null}
npm install -g modellix-cli
```
### What You Can Do
* **List model types**: `modellix-cli model types` to see all supported model types.
* **Create tasks**: `modellix-cli model invoke` to submit async generation tasks with inline JSON or a file body.
* **Query results**: `modellix-cli task get ` to poll task status and retrieve generation output.
### Designed for Automation
The CLI is built for scripting and AI agent workflows. All commands support `--json` output, and authentication works via the `MODELLIX_API_KEY` environment variable or the `--api-key` flag.
For full usage details, see the [CLI documentation](/ways-to-use/cli).
## Modellix Is Now Live!
[Modellix](https://modellix.ai) is officially launched! 🎉
## Console
* **Usage**: Rich visualizations with charts and lists to monitor your model API consumption on Modellix.
* **Top Up**: Securely add funds via Stripe, with support for auto-recharge.
* **Order**: A complete and clear view of every order you have placed.
* **Transaction**: A complete and clear record of every transaction, including top-ups and consumption.
Fill out [this form](https://forms.gle/VU8VAapGzYuEwiT79) to receive **\$10 – 30 in bonus credits** and become one of our first users!
Join our [Discord community](https://discord.gg/N2FbcB2cZT) to connect with us.
## Agent Skill
* Updated `SKILL.md` to improve Skill accuracy.
* Published the Skill to [ClawHub](https://clawhub.ai/Modellix/modellix).
## Console
* Registration and login, supporting Email OTP, Google and Github.
* API Key management: supports creating, deleting, and viewing API Keys, which can be used to access the Model API.
# Modellix Rate Limits and Team Entitlements
Source: https://docs.modellix.ai/get-started/entitlements
Understand Modellix concurrency limits, rate limits (RPM), and team entitlements that scale with your single top-up funding tier amount.
## Team Entitlements
Modellix provides different tiers of usage limits and entitlements based on your single top-up amount.
Check your current team entitlements and limits in the Modellix Console.
### Entitlements Table
| Single Top-up Amount | Concurrent Tasks | Rate Limit (RPM) |
| :------------------- | :--------------- | :--------------- |
| `< $10` | 2 | 100 |
| `$10` | 10 | 100 |
| `$100` | 20 | 200 |
| `$200` | 30 | 300 |
| `$500` | 50 | 500 |
| `$1,000` | 100 | 1,000 |
| `Custom` | Custom | Custom |
**RPM** stands for Requests Per Minute. **Concurrent Tasks** refers to the maximum number of asynchronous generation tasks (e.g., video or image generation) running in parallel.
### Custom Limits
If your application requires higher concurrency limits or a larger Rate Limit (RPM), please contact us via [email](mailto:support@modellix.ai) to discuss custom enterprise plans.
# Set Up AI Agents on Modellix
Source: https://docs.modellix.ai/get-started/index
Set up an AI agent with one Modellix API key for Media Models, the LLM gateway, and optional Web Tools. Use async polling for media; call LLM and Tools synchronously.
> If you are an AI agent, this document provides the essential context and tools you need to interact with Modellix.
Modellix is a MaaS platform. One Modellix API key covers **Media Models** and **LLMs**. Web Search and Web Fetch are optional Tools. Pick the host that matches the job—do not send media generation to the LLM gateway, and do not poll LLM or Tools responses.
| Job | Host | Call style |
| ---------------------------------- | -------------------------- | ------------------------------- |
| Image, video, or speech generation | `https://api.modellix.ai` | Async: submit a task, then poll |
| Chat, coding, or text generation | `https://llm.modellix.ai` | Sync (optional SSE) |
| Public web results or page content | `https://tool.modellix.ai` | Sync |
Human-oriented product map: [Platform Overview](/get-started/overview).
## Prerequisite: Create an API Key
A human must create a Modellix account and generate an API key. Create one in the [console](https://modellix.ai/console/api-key). With that key, your agent can call Media Models, LLM, and Web Tools.
The API key is displayed only once after creation. Store it securely, typically as `MODELLIX_API_KEY`.
## Agent Skills
The official Modellix Skill teaches coding agents how to discover media models, inspect request schemas, and generate images, videos, and speech. It does not replace the [LLM Overview](/llm/overview) for chat gateway setup.
Install via skills.sh:
```bash theme={null}
npx skills add https://github.com/Modellix/modellix-plugin --skill modellix
```
Target one agent:
```bash theme={null}
npx skills add https://github.com/Modellix/modellix-plugin --skill modellix --agent cursor
```
Installation for Cursor, Claude Code, GitHub Copilot, Codex, and other Agent Skills hosts.
## Modellix CLI
`modellix-cli` creates **media** generation tasks and fetches results from the terminal. Use it for async image, video, and speech workflows—not for LLM chat.
```bash theme={null}
# Install the CLI
npm install -g modellix-cli
# Export the API key
export MODELLIX_API_KEY="your_api_key"
# Create a generation task
modellix-cli model run \
--model-slug alibaba/qwen-image-edit \
--body '{"prompt":"A cute cat playing in a garden on a sunny day"}'
# Fetch the task result using the returned task_id
modellix-cli task get task-abc123
```
Authentication, `model run --wait`, schemas, and task downloads.
## Media Models: Async Two-Step
Media generation uses `https://api.modellix.ai`. Copy the prompt below for your agent, or open it in Cursor.
Modellix **Media Models** use an asynchronous two-step pattern. Do not apply this pattern to `https://llm.modellix.ai` or `https://tool.modellix.ai`.
### The Two-Step Pattern
1. **Create Task (`POST`)**: Send generation parameters to the model's endpoint on `https://api.modellix.ai`. The API responds immediately with a `task_id` and a status of `pending`.
2. **Poll Result (`GET`)**: Query `https://api.modellix.ai/api/v1/tasks/{task_id}` until `status` is `success` or `failed`.
### Error Handling
| Code | Action |
| ----------- | -------------------------------------------------------------------------------- |
| **400** | **Do not retry**. Fix parameters or request body format. |
| **401** | **Do not retry**. Verify the API key is provided and valid. |
| **402** | **Do not retry**. Account balance is insufficient. Human intervention required. |
| **404** | **Do not retry**. Verify the `task_id` or `model-slug`. |
| **429** | **Retry with exponential backoff**. You hit rate or concurrency limits. |
| **500/503** | **Retry with exponential backoff** (up to 3 times). Temporary server-side issue. |
### Minimal Polling Example (Node.js)
```javascript theme={null}
const API_KEY = process.env.MODELLIX_API_KEY;
const MODEL_URL = 'https://api.modellix.ai/api/v1/alibaba/qwen-image-plus/async';
async function generateImage(prompt) {
const createRes = await fetch(MODEL_URL, {
method: 'POST',
headers: {
'Authorization': `Bearer ${API_KEY}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({ prompt })
});
const createData = await createRes.json();
if (createData.code !== 0) throw new Error(`Task creation failed: ${createData.message}`);
const taskId = createData.data.task_id;
const pollUrl = `https://api.modellix.ai/api/v1/tasks/${taskId}`;
while (true) {
await new Promise((resolve) => setTimeout(resolve, 3000));
const pollRes = await fetch(pollUrl, {
headers: { 'Authorization': `Bearer ${API_KEY}` }
});
const pollData = await pollRes.json();
if (pollData.code !== 0) throw new Error(`Polling failed: ${pollData.message}`);
const status = pollData.data.status;
if (status === 'success') {
return pollData.data.result.resources;
} else if (status === 'failed') {
throw new Error(`Task failed: ${pollData.data.error_message || 'Unknown error'}`);
}
}
}
```
Full media workflow: [REST API](/ways-to-use/api).
## LLM: Synchronous Chat
The LLM gateway at `https://llm.modellix.ai` returns text in one response (optional SSE). Pass `model` as a `provider/name` ID. Do not poll and do not mix Chat Completions fields with Messages fields.
| Client type | Base URL | Protocol |
| ------------------------------------------------- | ------------------------------------ | ----------------------------- |
| OpenAI SDK, Codex, Cursor, OpenCode (OpenAI mode) | `https://llm.modellix.ai/v1` | Chat Completions or Responses |
| Anthropic SDK, Claude Code | `https://llm.modellix.ai` (no `/v1`) | Messages |
Modellix **LLM** is a synchronous gateway at `https://llm.modellix.ai`. Do not create media tasks on this host. Do not poll.
### Auth and Model IDs
Use a Modellix API key (Bearer or `x-api-key`), not a vendor platform key. Set `model` to a `provider/name` ID from [https://www.modellix.ai/llm](https://www.modellix.ai/llm) (for example `openai/gpt-5.6-sol` or `anthropic/claude-sonnet-5`).
### Base URLs
* OpenAI-compatible Chat Completions / Responses: `https://llm.modellix.ai/v1`
* Anthropic-compatible Messages: `https://llm.modellix.ai` (no `/v1`)
### Minimal Chat Completions Example
```bash theme={null}
curl -sS "https://llm.modellix.ai/v1/chat/completions" \
-H "Authorization: Bearer ${MODELLIX_API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-5.6-sol",
"stream": false,
"max_tokens": 256,
"messages": [{"role": "user", "content": "Introduce yourself in one sentence"}]
}'
```
### Error Handling
| Code | Action |
| ------- | -------------------------------------------------------------------- |
| **400** | **Do not retry**. Fix the request body or protocol fields. |
| **401** | **Do not retry**. Verify the Modellix API key. |
| **402** | **Do not retry**. Insufficient balance (`insufficient_quota`). |
| **404** | **Do not retry**. Unknown path or model unavailable. |
| **429** | **Retry with backoff**. Rate limit or model temporarily unavailable. |
| **5xx** | **Retry with backoff**. Temporary upstream or service error. |
Protocols, multimodal inputs, and billing: [https://docs.modellix.ai/llm/overview](https://docs.modellix.ai/llm/overview)
Protocols, SDK overrides, coding tools, and token billing.
## Web Tools
Web Search and Web Fetch at `https://tool.modellix.ai` are optional. Use them when the user needs live public web results or readable page content. Calls are synchronous. REST requires `X-Mdlx-User-Id`. MCP clients can connect to `https://tool.modellix.ai/mcp`.
Web Search, Web Fetch, headers, and per-request pricing.
Streamable HTTP MCP for Cursor, Claude Code, Codex, and other clients.
# Modellix AI Providers and Service URLs
Source: https://docs.modellix.ai/get-started/model-providers
Browse every upstream provider Modellix routes through for Media Models, LLM, and Web Tools, including official service URLs and links to the matching API docs.
Modellix aggregates production AI models and web tools from multiple upstream providers into one API, one billing account, and one Modellix API key. The service URLs below point to each provider's official site or console for reference. You still call models and tools through Modellix—you do not need a separate key for every vendor.
Async image, video, and speech at `https://api.modellix.ai`.
Sync chat gateway at `https://llm.modellix.ai`.
Sync Web Search and Web Fetch at `https://tool.modellix.ai`.
For help choosing a media model by quality, speed, or modality, see [Model Selection](/get-started/model-select). To browse media models by capability, use the [category pages on modellix.ai](https://www.modellix.ai/category/text-to-image). For LLM Model IDs and token rates, use [Modellix LLM](https://www.modellix.ai/llm).
## Media Models
Image, video, and speech generation use the async Media Model API. Submit a job, then poll [Query Task Result](/api/get-task-result).
| Provider | Typical modalities on Modellix | Upstream service URLs | Model API |
| --------- | ---------------------------------------------- | --------------------------------------------------------------------------------------------------- | --------------------------------------------- |
| Alibaba | Image, video, speech (TTS / STT / voice clone) | [aliyun.com](https://www.aliyun.com/) | [Alibaba docs](/alibaba/happyhorse-1-0-i2v) |
| ByteDance | Image, video | [byteplus.com](https://www.byteplus.com/) | [ByteDance docs](/bytedance/seedance-2-5-i2v) |
| Google | Image, video, speech | [Google AI Studio](https://aistudio.google.com/), [Google Cloud](https://console.cloud.google.com/) | [Google docs](/google/nano-banana-2-lite) |
| Kling | Image, video, video editing | [klingai.com](https://klingai.com/) | [Kling docs](/kling/kling-video-o1) |
| Microsoft | Image | [Azure](https://azure.microsoft.com/) | [Microsoft docs](/microsoft/mai-image-2-5) |
| MiniMax | Video, speech | [minimax.io](https://www.minimax.io/) | [MiniMax docs](/minimax/hailuo-02-fl2v) |
| OpenAI | Image, speech-to-text | [Azure OpenAI](https://azure.microsoft.com/) | [OpenAI docs](/openai/gpt-image-2) |
| PixVerse | Image, video, video editing | [pixverse.ai](https://pixverse.ai/) | [PixVerse docs](/pixverse/c1-fl2v) |
| Skywork | Image, video, video editing | [skywork.ai](https://skywork.ai/) | [Skywork docs](/skyreels/skyreels-t2v) |
| Vidu | Image, video, video editing | [vidu.com](https://vidu.com/) | [Vidu docs](/vidu/lip-sync) |
| xAI | Image, video | [x.ai](https://x.ai/) | [xAI docs](/xai/grok-imagine-image) |
## LLM
The LLM gateway is synchronous (optional SSE) at `https://llm.modellix.ai`. Pass `model` as a `provider/name` ID from the [LLM catalog](https://www.modellix.ai/llm)—for example `openai/gpt-5.6-sol` or `anthropic/claude-sonnet-5`. Protocols and client setup: [LLM Overview](/llm/overview).
The table lists the upstream platforms Modellix routes those calls through. Cloud hosts (Azure, AWS Bedrock, GCP, OpenRouter) are not Model ID prefixes unless the catalog lists that prefix.
| Provider | Typical Model IDs on Modellix | Upstream service URLs | LLM API |
| ----------- | ---------------------------------- | ---------------------------------------------------------------------------------------------------------------- | ----------------------------- |
| Azure | OpenAI GPT (`openai/...`) | [Azure](https://azure.microsoft.com/) | [LLM Overview](/llm/overview) |
| AWS Bedrock | Anthropic Claude (`anthropic/...`) | [Amazon Bedrock](https://aws.amazon.com/bedrock/) | [LLM Overview](/llm/overview) |
| GCP | Google Gemini (`google/...`) | [Google Cloud](https://cloud.google.com/), [Vertex AI](https://cloud.google.com/vertex-ai) | [LLM Overview](/llm/overview) |
| xAI | Grok (`xai/...`) | [x.ai](https://x.ai/) | [LLM Overview](/llm/overview) |
| ZAI | GLM (`zai/...`) | [z.ai](https://z.ai/) | [LLM Overview](/llm/overview) |
| DeepSeek | DeepSeek (`deepseek/...`) | [deepseek.com](https://www.deepseek.com/) | [LLM Overview](/llm/overview) |
| Qwen | Qwen (`qwen/...`) | [Alibaba Cloud](https://www.alibabacloud.com/), [Model Studio](https://www.alibabacloud.com/product/modelstudio) | [LLM Overview](/llm/overview) |
| Moonshot | Kimi (`moonshot/...`) | [moonshot.ai](https://www.moonshot.ai/) | [LLM Overview](/llm/overview) |
| OpenRouter | Additional routed catalog models | [openrouter.ai](https://openrouter.ai/) | [LLM Overview](/llm/overview) |
Qwen LLM traffic is sourced from Alibaba Cloud International (Model Studio), not the China-region `aliyun.com` console used for Alibaba Media Models.
## Tools
Web Search and Web Fetch are synchronous calls at `https://tool.modellix.ai`. See [Tools Overview](/tools/overview) for pricing, headers, and logs.
| Provider | Typical tools on Modellix | Upstream service URLs | Tools API |
| -------- | ------------------------- | ------------------------------------- | ---------------------------------------------------------- |
| Tavily | Web Search, Web Fetch | [tavily.com](https://www.tavily.com/) | [Web Search](/api/web-search), [Web Fetch](/api/web-fetch) |
| Exa | Web Search, Web Fetch | [exa.ai](https://exa.ai/) | [Web Search](/api/web-search), [Web Fetch](/api/web-fetch) |
## How Modellix Uses Provider Services
* **Unified API**: Authenticate with your Modellix API key. Media Models use provider paths under `https://api.modellix.ai` (for example `/alibaba/...`, `/google/...`). LLM uses `https://llm.modellix.ai` with `provider/name` Model IDs. Tools use `https://tool.modellix.ai`.
* **Request model**: Media and speech jobs are async—poll [Query Task Result](/api/get-task-result). LLM and Tools return synchronously (LLM may stream SSE).
* **Upstream URLs**: The tables list vendor marketing or console sites so you can read provider-level product context. Runtime traffic for Modellix customers goes through Modellix infrastructure, not those URLs directly.
## Availability
Modellix may add or remove providers and individual models based on product needs. Treat the [Media Models](/alibaba/happyhorse-1-0-i2v) navigation, [LLM catalog](https://www.modellix.ai/llm), [New Models](/changelog/new-models), and [Deprecated Models](/changelog/deprecated-models) as the source of truth for what is currently documented and what has been removed.
# Choose the Right Modellix Image, Video, or Speech Model
Source: https://docs.modellix.ai/get-started/model-select
Choose the right Modellix image, video, or speech model by output type, quality, speed, cost, and features such as audio, references, and editing.
Sometimes you may not know which model is best suited for your specific creative task. Follow this step-by-step guide to get tailored model recommendations using an AI assistant.
Determine what kind of media you want to generate. Consider factors such as:
* **Format**: Do you need text-to-image, image editing, text-to-video, image-to-video, video editing, text-to-speech, speech-to-text, or voice clone?
* **Quality vs. Speed**: Do you need high-fidelity cinematic outputs (e.g., Vidu Q3 Pro, Grok Imagine Quality) or fast, cost-effective generations (e.g., MAI Image 2.5 Flash, Vidu Q3 Turbo, CosyVoice Flash)?
* **Special Features**: Do you need character consistency, native video audio, specific voices or languages, or specific aspect ratios?
Submit the Modellix model directory link `https://docs.modellix.ai/llms.txt` along with your specific requirements to an AI assistant (such as [Claude](https://claude.ai/), ChatGPT, or Gemini).

For example, you can use a prompt like this:
```plaintext Prompt Example wrap theme={null}
https://docs.modellix.ai/llms.txt
I want to create cinematic, high-quality videos from a starting image. Which model would be more suitable?
# Or for speech:
# I need natural English TTS with SSML and low latency. Which Modellix speech model should I use?
```
The AI assistant will read the comprehensive model descriptions in our `llms.txt` file, analyze your requirements, and recommend the best-fitting models for your specific use case.

The `llms.txt` file is automatically kept up to date with descriptions of every model available on the [Modellix Platform](https://www.modellix.ai/explore), ensuring you always get recommendations based on the latest model capabilities.
# Modellix Platform Overview
Source: https://docs.modellix.ai/get-started/overview
Learn how Modellix provides unified API access to Media Models and LLMs, plus optional Web Tools, with one API key, public pricing, and browser playgrounds.
[Modellix](https://modellix.ai) is a MaaS (Model as a Service) platform. Use one Modellix API key and one billing account to call **Media Models** and **LLMs**. Web Search and Web Fetch are optional Tools for workflows that need live web context.
Developers call the APIs; creators generate media in the browser Playground. Both share the same account and public pricing.
## What Modellix Provides
Asynchronous image, video, and speech at `https://api.modellix.ai`. Submit a task, then poll for the result.
Synchronous chat gateway at `https://llm.modellix.ai`. OpenAI-compatible Chat Completions and Responses, plus Anthropic-compatible Messages.
**Web Tools** at `https://tool.modellix.ai` are optional. Use [Web Search](/api/web-search) and [Web Fetch](/api/web-fetch) to ground agents with public web results and page content. See [Tools Overview](/tools/overview). Use the host that matches the product—do not send media jobs to the LLM gateway or chat requests to the media API.
| | Media Models | LLM | Web Tools |
| ----------- | ------------------------------- | ----------------------------- | --------------------------------- |
| Host | `https://api.modellix.ai` | `https://llm.modellix.ai` | `https://tool.modellix.ai` |
| Call style | Async task, then poll | Sync, optional SSE | Sync |
| Billing | Per image, second, or character | Per token | Per request or successful URL |
| Typical use | Generate or edit media | Chat, coding, and agents | Live web context for agents |
| Start here | [REST API](/ways-to-use/api) | [LLM Overview](/llm/overview) | [Tools Overview](/tools/overview) |
Upstream providers for all three surfaces: [Providers](/get-started/model-providers).
## Media Models
Image, video, and speech generation use the async Media Model API. Developers submit a job and poll [Query Task Result](/api/get-task-result). Creators can skip the API and run the same models in the [Playground](/ways-to-use/playground).
Browse production-ready models by modality, compare capabilities and pricing, then try them in the Playground or call them through the API.
Generate still images from text prompts across providers such as Qwen, Seedream, Nano Banana, and GPT Image.
Edit, restyle, or compose images with reference inputs and instruction-based image-to-image models.
Create videos from text descriptions with models such as Seedance, Veo, Kling, and Wan.
Animate a still image into video while preserving subject and composition.
Restyle, extend, lip-sync, or otherwise transform an existing video with editing models.
Convert text into controllable speech audio with system voices, instructions, and format options.
Transcribe audio asynchronously and retrieve normalized transcript documents from task results.
Clone a speaker from reference audio and synthesize new speech in one workflow.
Need help picking a media model? See [Model Selection](/get-started/model-select), the [REST API guide](/ways-to-use/api), and the [Playground guide](/ways-to-use/playground).
## LLM
The LLM gateway at `https://llm.modellix.ai` returns text. Some models accept image, audio, or video as **input**—see [Multimodal Inputs](/llm/api/api#multimodal-inputs). Do not send image, video, or speech **generation** jobs to this host.
Pass `model` as a `provider/name` ID (for example `openai/gpt-5.5` or `anthropic/claude-sonnet-5`). Point OpenAI or Anthropic SDKs, and coding tools such as Codex, Claude Code, Cursor, and OpenCode, at Modellix with a base URL override. Use the same Modellix API key as Media Models.
Protocols, Quick Start, SDK and coding-tool guides.
Current Model IDs, Input Context tiers, and USD-per-1M-token rates.
## Web Tools
Web Search and Web Fetch at `https://tool.modellix.ai` are optional helpers, not a third model product. Call them when an agent needs ranked public web results or readable page content. Authenticate with the same Modellix API key.
Web Search, Web Fetch, per-request pricing, and tool logs.
## What You Get
### For Developers
One Modellix key for Media Models, LLM, and Web Tools. Flattened media parameters; `provider/name` IDs on the LLM gateway.
Published rates for every media model, LLM token tier, and tool SKU. No hidden add-on fees.
Get help by [email](mailto:support@modellix.ai) or [Discord](https://discord.gg/N2FbcB2cZT).
Track deposits and usage in the console so you can see what each call costs.
Trace requests, responses, and costs for media, LLM, and tool calls.
Backed by a NASDAQ-listed company (JG) with 15+ years of enterprise AI infrastructure experience.
### For Creators
Creators use Media Models without writing API calls. Discover a model, open its detail page, and generate in the browser.
Browse featured models at [Modellix Models](https://www.modellix.ai/models), or filter by provider and modality at [Modellix Explore](https://www.modellix.ai/explore).
Every media model's detail page includes a Playground. Set prompts and parameters, generate in the browser, and inspect the JSON payload when you are ready to automate.
## How to Start
Authenticate, submit an async media task, poll status, and retrieve output assets.
Pick a protocol, set the LLM base URL, and send a chat request.
Media units, LLM token rates, and Web Tool SKUs, with links to live catalogs.
Install the official Skill so agents construct Modellix requests correctly.
Create media tasks, poll results, and automate workflows from the terminal.
Find a media model, then generate in the Playground on its detail page.
# Modellix API Pricing for Media, LLM, and Tools
Source: https://docs.modellix.ai/get-started/pricing
Understand Modellix pay-as-you-go pricing: media generation by image, second, or character; LLM by token; optional Web Tools by request or URL.
Modellix bills from one account balance. Media Models, LLMs, and Web Tools use different units. Rates are public; confirm the live price before you ship a workload.
| Product | How You Are Billed | Live rates |
| ---------------- | --------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------- |
| **Media Models** | Per image, per second, or per character (some speech jobs bill by audio duration) | [Modellix Models](https://www.modellix.ai/models), [Pricing table](https://www.modellix.ai/pricing) |
| **LLM** | Per token (`usage` on successful responses), USD per 1M tokens | [Modellix LLM](https://www.modellix.ai/llm) |
| **Web Tools** | Web Search per request; Web Fetch per successful URL | [Tools Overview](/tools/overview) |
Fund the same balance for all three: [Top-up](/get-started/top-up). Insufficient balance returns `402`.
## Media Models
Open a model on [Modellix Models](https://www.modellix.ai/models) to see prices for each parameter combination. To compare active media models in one table, use the [Modellix Pricing Page](https://www.modellix.ai/pricing).
### Pricing Units
* `to-image`: Charged by **USD/img**. `img` is the number of output images.
* `to-video`: Charged by **USD/sec**. `sec` is the output video duration.
* `to-audio`: Charged by **USD/M chars**. `M chars` means one million input characters (text-to-speech and related audio models).
Some models change price with input parameters. For example, a video model may cost more at `1080p` than at `720p`. Some speech or transcription models bill by audio duration (**USD/sec**). Always confirm the unit on the model detail page or the pricing table.
Submit an async media task, poll status, and retrieve billed outputs.
## LLM
The LLM gateway bills **successful** responses from token `usage` at the model's input and output rates. Cache read/write tokens are billed when the response includes them. `prompt_tokens_details.video_tokens` and `audio_tokens` are already included in `prompt_tokens` and are not billed as a separate line.
Pass `model` as a `provider/name` ID. Current Model IDs, Input Context tiers, discounts, and USD-per-1M-token rates are on the [Modellix LLM](https://www.modellix.ai/llm) page. That page also estimates a call with the cost calculator.
Filter by provider and input modality; compare list prices with Modellix rates.
Protocols, `usage` fields, and billing notes for Chat Completions, Responses, and Messages.
## Web Tools
Web Tools are optional. Prices below are in **USD**. Confirm live SKUs on [Tools Overview](/tools/overview) or in the [console](https://modellix.ai/console).
### Web Search
[Web Search](/api/web-search) (`POST /v1/web-search`) is charged **once per request** according to `depth`. It is not billed per result or per URL. Default `depth` is `standard`.
| Depth | SKU | Price (USD / request) |
| ---------- | --------------------- | --------------------: |
| `lite` | `web-search.lite` | \$0.01 |
| `standard` | `web-search.standard` | \$0.015 |
| `rich` | `web-search.rich` | \$0.02 |
### Web Fetch
[Web Fetch](/api/web-fetch) (`POST /v1/web-fetch`) is charged **\$0.002 per successful URL** (`web-fetch`). Failed URLs are not billed.
Availability and prices can change. Confirm live rates in the [Modellix console](https://modellix.ai/console) and on [Tools Overview](/tools/overview).
# Fund Your Modellix Account
Source: https://docs.modellix.ai/get-started/top-up
Learn how to fund your Modellix account with supported payment methods, pay-as-you-go top-ups, and available discounts for your team.
## Account Funding
Modellix operates on a **Pay-as-you-go** billing model, allowing you to top up your account balance based on your actual usage requirements.
Fund your account balance securely through the Modellix Console.
### Payment Methods
Currently, Modellix supports **Stripe** as the primary payment gateway. You can securely pay using credit/debit cards and other payment methods supported by Stripe.
### Top-Up Discounts
Modellix currently offers a **10% discount on your first top-up** for each tier!
The discount applies to the first transaction in each of the following tiers:
* **\$10**
* **\$100**
* **\$200**
* **\$500**
* **\$1,000**
* **Custom Amount**
This means you can enjoy the 10% first-time discount up to **6 times** in total (once for each tier).
# Gemini 3.1 Flash TTS
Source: https://docs.modellix.ai/google/gemini-3-1-flash-tts
/media-model-api/google/google-t2s.json post /google/gemini-3.1-flash-tts
[Core Function] Gemini 3.1 Flash TTS is Google's low-latency, controllable text-to-speech model. [Strengths] Single-speaker and two-speaker dialogue, 30 prebuilt voices, 70+ languages via language_code, and expressive delivery through style prompts plus inline audio tags such as [whispers], [slow], [fast], and [laughs]. Output is WAV (24 kHz mono PCM). [Best For] Voiceovers, virtual presenters, audiobook narration, multilingual speech, podcast-style scripts, and two-person dialogue. [Limitations] Do NOT use for lip-syncing existing video, music or sound-effect generation, image or audio inputs, MP3/OGG export, or API-level speed, volume, encoding, or sample-rate controls. Combined prompt and text must stay within 8,000 bytes. [Routing] Route here for controllable Gemini TTS from text only; for video with native audio use Veo; for talking-head lip sync from a portrait plus audio use Kling Avatar or SkyReels avatar models.
# Gemini 3.5 Transcribe
Source: https://docs.modellix.ai/google/gemini-3-5-transcribe
/media-model-api/google/google-s2t.json post /google/gemini-3.5-transcribe
[Core Function] Gemini 3.5 Transcribe is Google's speech-to-text model for complete pre-recorded audio. [Strengths] Accurate multilingual recognition across 85+ languages with optional language hints, custom vocabulary for brand names, speaker labels (up to 8 speakers), word-level timestamps, and Smart formatting that cleans punctuation and numbers. [Best For] Meeting notes, captions and subtitles, multilingual recordings, call logs, and speaker-attributed transcripts. [Limitations] Do NOT use this if the audio is not a public HTTPS URL, longer than 15 minutes, or larger than 300 MB. Do NOT use file upload. Do NOT combine mode=smart with word timestamps or speaker labels. This endpoint does not support live or streaming transcription. [Routing] Choose this model for Google Gemini file transcription with speaker labels or Smart formatting. For Microsoft recognition use MAI-Transcribe; for OpenAI formats such as SRT or VTT use Whisper.
# Gemini Omni 1.1 Flash I2V
Source: https://docs.modellix.ai/google/gemini-omni-1-1-flash-i2v
/media-model-api/google/google-i2v.json post /google/gemini-omni-1.1-flash-i2v
[Core Function] Gemini Omni 1.1 Flash I2V animates a single image into a short video with synchronized audio. [Strengths] It uses the input image as the opening frame and supports 360p, 720p, 1080p, and 4k output, 16:9 or 9:16, and 3 to 10 second clips. [Best For] Highly recommended for: animating a still, product or character motion from one frame, and social clips that need built-in audio at 1080p or 4k. [Limitations] Do NOT use this model if you have no starting image (use Gemini Omni 1.1 Flash T2V) or want to edit an existing video (use Gemini Omni 1.1 Flash V2V). Clips are capped at 10 seconds. [Routing] Choose this when the user provides one starting frame. For text-only generation, use Gemini Omni 1.1 Flash T2V.
# Gemini Omni 1.1 Flash T2V
Source: https://docs.modellix.ai/google/gemini-omni-1-1-flash-t2v
/media-model-api/google/google-t2v.json post /google/gemini-omni-1.1-flash-t2v
[Core Function] Gemini Omni 1.1 Flash T2V generates a short video with synchronized audio from a text prompt. [Strengths] It supports 360p, 720p, 1080p, and 4k output, 16:9 or 9:16, and 3 to 10 second clips with native speech, music, and sound effects. [Best For] Highly recommended for: rapid prototyping, short social and marketing clips, concept visualization, and cases that need built-in audio at 1080p or 4k. [Limitations] Do NOT use this model if you need clips longer than 10 seconds, a starting image as the first frame (use Gemini Omni 1.1 Flash I2V), or edits to an existing video (use Gemini Omni 1.1 Flash V2V). [Routing] Choose this when the user wants a new video from text only. If they provide a starting image, use Gemini Omni 1.1 Flash I2V. For maximum cinematic control, choose Veo 3.1 T2V.
# Gemini Omni 1.1 Flash V2V
Source: https://docs.modellix.ai/google/gemini-omni-1-1-flash-v2v
/media-model-api/google/google-v2v.json post /google/gemini-omni-1.1-flash-v2v
[Core Function] Gemini Omni 1.1 Flash V2V edits an existing video from a text instruction. [Strengths] It can change scene, mood, style, lighting, or time of day while keeping the source length and aspect ratio, with optional 360p to 4k output and synchronized audio. [Best For] Highly recommended for: re-styling or re-lighting a clip, changing setting or atmosphere, and quick revisions of a short video. [Limitations] Do NOT use this model to generate a video from scratch (use Gemini Omni 1.1 Flash T2V or I2V). The source video should be 3 to 10 seconds; output length and aspect ratio follow the source. [Routing] Choose this only when the user provides an existing video to modify.
# Gemini Omni Flash I2V
Source: https://docs.modellix.ai/google/gemini-omni-flash-i2v
/media-model-api/google/google-i2v.json post /google/gemini-omni-flash-i2v
[Core Function] Gemini Omni Flash I2V is a fast Image-to-Video model that animates a single input image into a short 720p video via the Interactions API. [Strengths] It uses the provided image as the opening frame and generates smooth motion with natively synchronized audio at low latency. [Best For] Highly recommended for: bringing a still photo to life, quick product or portrait animation, and short social clips derived from a single image. [Limitations] Do NOT use this model if you need 1080p or 4K output, clips longer than 10 seconds, or the fusion of multiple reference images; it takes exactly one image and outputs 720p up to 10 seconds (16:9 or 9:16). Do NOT use it to edit an existing video. [Routing] Choose this when the user provides one image to animate. To fuse multiple reference images use Gemini Omni Flash R2V; to edit an existing video use Gemini Omni Flash Video Edit; for 4K cinematic results use Veo 3.1 I2V.
# Gemini Omni Flash R2V
Source: https://docs.modellix.ai/google/gemini-omni-flash-r2v
/media-model-api/google/google-i2v.json post /google/gemini-omni-flash-r2v
[Core Function] Gemini Omni Flash R2V (Reference-to-Video) generates a short 720p video guided by up to three reference images via the Interactions API. [Strengths] It fuses the styles, subjects, or elements from multiple reference images (referred to in the text prompt) into a single coherent animated clip with synchronized audio. [Best For] Highly recommended for: blending characters or visual styles from several images, reference-guided creative shots, and multi-subject compositions where the prompt directs how the references combine. [Limitations] Do NOT use this model if you only have a single starting frame (use I2V instead), or if you need 1080p or 4K or clips longer than 10 seconds; it accepts 1 to 3 reference images and outputs 720p up to 10 seconds (16:9 or 9:16). [Routing] Choose this when the user supplies multiple reference images to combine into one video. For single first-frame animation use Gemini Omni Flash I2V; to modify an existing video use Gemini Omni Flash Video Edit.
# Gemini Omni Flash T2V
Source: https://docs.modellix.ai/google/gemini-omni-flash-t2v
/media-model-api/google/google-t2v.json post /google/gemini-omni-flash-t2v
[Core Function] Gemini Omni Flash T2V is Google's fast multimodal Text-to-Video generation model built on the Interactions API. [Strengths] It quickly turns a text prompt into a short 720p video with natively synchronized audio, offering low latency and solid prompt adherence. [Best For] Highly recommended for: rapid text-to-video prototyping, short social and marketing clips, quick concept visualization, and cases where speed and built-in audio matter more than 4K cinematic detail. [Limitations] Do NOT use this model if you need 1080p or 4K resolution or clips longer than 10 seconds; output is fixed at 720p, capped at 10 seconds, with aspect ratio limited to 16:9 or 9:16. [Routing] Choose this model when the user emphasizes 'fast', 'quick', or short multimodal clips with sound. If the user demands maximum cinematic quality, 4K, or longer videos, choose Veo 3.1 T2V instead.
# Gemini Omni Flash Video Edit
Source: https://docs.modellix.ai/google/gemini-omni-flash-video-edit
/media-model-api/google/google-v2v.json post /google/gemini-omni-flash-video-edit
[Core Function] Gemini Omni Flash Video Edit performs conversational, instruction-driven editing of an existing video via the Interactions API. [Strengths] It applies natural-language edits (changing the scene, mood, style, lighting, background, or time of day) to an input video while preserving the source video's length and aspect ratio, with synchronized audio. [Best For] Highly recommended for: re-styling or re-lighting an existing clip, changing a video's setting or atmosphere, and quick instruction-based revisions of a short video. [Limitations] Do NOT use this model to generate a video from scratch (use T2V, I2V, or R2V), and do NOT expect to change the output resolution, aspect ratio, or duration: the output preserves the source video's aspect ratio and length, and the model does not accept aspectRatio or duration parameters. The source video should be 3 to 10 seconds. [Routing] Choose this only when the user provides an existing video to modify. To create a new video from text or images, use Gemini Omni Flash T2V, I2V, or R2V instead.
# Nano Banana
Source: https://docs.modellix.ai/google/nano-banana
/media-model-api/google/google-t2i.json post /google/nano-banana
[Core Function] Nano Banana is the original fast creative image model. [Strengths] Very fast creative generation. [Best For] Quick sketches and ideas. [Limitations] Superseded by Nano Banana 2 for general speed tasks. [Routing] Default to Nano Banana 2 unless specifically requested.
# Nano Banana 2
Source: https://docs.modellix.ai/google/nano-banana-2
/media-model-api/google/google-t2i.json post /google/nano-banana-2
[Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creative iteration, generating large batches of images quickly, and artistic/stylized graphics. [Limitations] Do NOT use this model if you require strict, high-end photorealism; prefer Nano Banana Pro for higher-fidelity creative output. [Routing] Route to this model for 'fast', 'creative', or 'stylized' high-volume requests.
# Nano Banana 2 Edit
Source: https://docs.modellix.ai/google/nano-banana-2-edit
/media-model-api/google/google-i2i.json post /google/nano-banana-2-edit
[Core Function] Nano Banana 2 Edit is a high-speed image editing model. [Strengths] It rapidly modifies existing images or extracts image frames from videos based on text prompts. [Best For] Highly recommended for: rapid style transfer, quick image modifications, and fast creative edits. [Limitations] Do NOT use this model for meticulous photorealistic retouching. [Routing] Use this model by default for fast, creative image editing tasks.
# Nano Banana 2 Lite
Source: https://docs.modellix.ai/google/nano-banana-2-lite
/media-model-api/google/google-t2i.json post /google/nano-banana-2-lite
[Core Function] Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is the lightweight, cost-efficient text-to-image model of the Nano Banana 2 family. [Strengths] It generates images even faster and more cheaply than Nano Banana 2, well suited to high-volume creative and stylized output at lower cost. [Best For] Highly recommended for: high-volume batch generation, quick drafts and thumbnails, and cost-sensitive creative iteration. [Limitations] Do NOT use this model when you need maximum detail, high-end photorealism, or the richest quality; use Nano Banana 2 or Nano Banana Pro instead. [Routing] Choose the Lite variant when cost and throughput matter more than peak quality; step up to Nano Banana 2 for richer results.
# Nano Banana 2 Lite Edit
Source: https://docs.modellix.ai/google/nano-banana-2-lite-edit
/media-model-api/google/google-i2i.json post /google/nano-banana-2-lite-edit
[Core Function] Nano Banana 2 Lite Edit (Gemini 3.1 Flash Lite Image) is the lightweight, cost-efficient image editing model of the Nano Banana 2 family; it transforms one or more input images per a text instruction. [Strengths] It performs fast, low-cost instruction-based editing across up to 14 input images. [Best For] Highly recommended for: high-volume edits, quick style transforms, and batch background or attribute changes where cost and throughput matter. [Limitations] Do NOT use this model when you need the highest edit fidelity or richest detail; use Nano Banana 2 Edit or Nano Banana Pro Edit instead. [Routing] Choose the Lite variant for cost- and throughput-sensitive edits; step up to Nano Banana 2 Edit for higher quality.
# Nano Banana Edit
Source: https://docs.modellix.ai/google/nano-banana-edit
/media-model-api/google/google-i2i.json post /google/nano-banana-edit
[Core Function] Nano Banana Edit is the original fast image editing model. [Strengths] Fast basic edits. [Limitations] Superseded by Nano Banana 2 Edit. [Routing] Default to Nano Banana 2 Edit unless specifically requested.
# Nano Banana Pro
Source: https://docs.modellix.ai/google/nano-banana-pro
/media-model-api/google/google-t2i.json post /google/nano-banana-pro
[Core Function] Nano Banana Pro (Gemini 3 Pro Image) is a high-capability creative image model. [Strengths] It balances the creative flexibility and speed of the Nano series with higher fidelity output. [Best For] Highly recommended for: high-quality stylized art and complex creative compositions. [Limitations] Do NOT use this model when you need the absolute highest photorealism or detail; step up within the Nano Banana family (Nano Banana 2 or Nano Banana Pro Edit workflows) as needed. [Routing] Use this model for high-quality, creative, non-photorealistic requests.
# Nano Banana Pro Edit
Source: https://docs.modellix.ai/google/nano-banana-pro-edit
/media-model-api/google/google-i2i.json post /google/nano-banana-pro-edit
[Core Function] Nano Banana Pro Edit is a high-capability creative image editing model. [Strengths] It provides high-quality creative edits, background replacements, and style transformations. [Best For] Highly recommended for: detailed creative modifications and complex style transfers. [Limitations] Do NOT use for strict photorealistic restoration. [Routing] Use this model for detailed, high-quality creative edits.
# Veo 3.1 Fast I2V
Source: https://docs.modellix.ai/google/veo-3-1-fast-i2v
/media-model-api/google/google-i2v.json post /google/veo-3.1-fast-i2v
[Core Function] Veo 3.1 Fast I2V is a high-speed image-to-video model. [Strengths] It quickly animates starting images at 1080p, optimized for low latency. [Best For] Highly recommended for: rapid prototyping and quick social media visual iterations. [Limitations] Do NOT use this model for the absolute highest visual fidelity or 4K output. [Routing] Choose this model when the user emphasizes 'fast' or 'quick' animation.
# Veo 3.1 Fast T2V
Source: https://docs.modellix.ai/google/veo-3-1-fast-t2v
/media-model-api/google/google-t2v.json post /google/veo-3.1-fast-t2v
[Core Function] Veo 3.1 Fast T2V is a high-speed text-to-video model. [Strengths] It is heavily optimized for fast generation, delivering video content with natively synchronized audio at 1080p quickly. [Best For] Highly recommended for: rapid prototyping, quick visual iteration, and high-volume background video generation. [Limitations] Do NOT use this model if you require 4K resolution or maximum artistic detail. [Routing] Choose this model when the user emphasizes 'fast', 'quick', or 'rapid' video generation.
# Veo 3.1 I2V
Source: https://docs.modellix.ai/google/veo-3-1-i2v
/media-model-api/google/google-i2v.json post /google/veo-3.1-i2v
[Core Function] Veo 3.1 I2V is Google's cinematic image-to-video generation model. [Strengths] It generates high-fidelity 4K video from a starting image. It supports advanced features like first-and-last frame conditioning and referencing up to three images. [Best For] Highly recommended for: animating concept art, creating cinematic transitions between images, and high-end video production. [Limitations] When using first+last frame or reference-only modes, the `duration` parameter must strictly be 8. Negative prompts are not supported in reference-only mode. [Routing] Use this model by default for high-quality image-to-video tasks or when multiple reference images are provided.
# Veo 3.1 Lite I2V
Source: https://docs.modellix.ai/google/veo-3-1-lite-i2v
/media-model-api/google/google-i2v.json post /google/veo-3.1-lite-i2v
[Core Function] Veo 3.1 Lite I2V is a balanced image-to-video model. [Strengths] It offers a middle ground between speed and quality for animating images. [Best For] Highly recommended for: general image animation and web-ready content. [Limitations] Do NOT use this model if you need 4K resolution. [Routing] Use this model for standard image animation requests.
# Veo 3.1 Lite T2V
Source: https://docs.modellix.ai/google/veo-3-1-lite-t2v
/media-model-api/google/google-t2v.json post /google/veo-3.1-lite-t2v
[Core Function] Veo 3.1 Lite T2V is a balanced text-to-video model. [Strengths] It provides a good balance between generation speed and visual quality, still supporting the advanced architecture of the 3.1 series. [Best For] Highly recommended for: general video content creation and social media posts where 4K is not strictly necessary. [Limitations] Do NOT use this model for the absolute highest cinematic fidelity (use the standard Veo 3.1 instead). [Routing] Route to this model for standard, everyday video generation tasks.
# Veo 3.1 T2V
Source: https://docs.modellix.ai/google/veo-3-1-t2v
/media-model-api/google/google-t2v.json post /google/veo-3.1-t2v
[Core Function] Veo 3.1 T2V is Google's state-of-the-art cinematic text-to-video engine. [Strengths] It natively generates 4K professional-grade video output with natively synchronized audio and supports complex camera movements. [Best For] Highly recommended for: high-end creative storytelling, cinematic short films, and experimental video production with sound. [Limitations] Do NOT use this model if you need instant/real-time generation, as 4K video rendering takes time. [Routing] Use this model by default for all high-quality text-to-video requests on the Google platform.
# Kling Avatar
Source: https://docs.modellix.ai/kling/kling-avatar
/media-model-api/kling/kling-i2v.json post /kling/kling-avatar
[Core Function] Kling Avatar is a specialized portrait animation model. [Strengths] It precisely animates a portrait image to lip-sync with an audio file or TTS audio ID. [Best For] Highly recommended for: virtual presenters, talking head videos, and digital avatars. [Limitations] Do NOT use this model for full-body action or general image animation. [Routing] Use this model explicitly when the user wants to make a portrait 'speak' with provided audio.
# Kling Image Expansion
Source: https://docs.modellix.ai/kling/kling-image-expansion
/media-model-api/kling/kling-i2i.json post /kling/kling-image-expansion
[Core Function] Kling Image Expansion is an outpainting model. [Strengths] It intelligently extends the borders of an image (horizontal, vertical, or asymmetric) while matching the original style and context. [Best For] Highly recommended for: changing aspect ratios, extending landscapes, and filling out cropped photos. [Limitations] Do NOT use this model for inpainting or style transfer. [Routing] Use this specifically when the user asks to 'expand', 'extend', or 'uncrop' an image.
# Kling Image O1
Source: https://docs.modellix.ai/kling/kling-image-o1
/media-model-api/kling/kling-i2i.json post /kling/kling-image-o1
[Core Function] Kling Image O1 is a reasoning-enhanced multimodal image model. [Strengths] It performs deep reasoning over prompts and references to handle complex logic, spatial relationships, and intricate multi-image combinations. [Best For] Highly recommended for: complex scenes requiring strict logical or spatial accuracy. [Limitations] Element library IDs and series generation are not exposed. Default aspect_ratio is 1:1 when omitted. Do NOT use for simple artistic generation where V3 is faster and more stylistic. [Routing] Route to this model when the prompt involves complex physical logic or strict spatial reasoning.
# Kling V3 I2I
Source: https://docs.modellix.ai/kling/kling-v3-i2i
/media-model-api/kling/kling-i2i.json post /kling/kling-v3-i2i
[Core Function] Kling V3 I2I is the flagship image-to-image editing model (POST /images/generations, model_name=kling-v3 with image). [Strengths] High-quality style transfer and editing up to 2K. [Best For] Single-reference image editing. [Limitations] Do NOT send negative_prompt when image is present (officially unsupported). No image_fidelity / image_reference on V3. For multi-image fusion or series, use kling-v3-omni-image. [Routing] Default for standard image-to-image.
# Kling V3 I2V
Source: https://docs.modellix.ai/kling/kling-v3-i2v
/media-model-api/kling/kling-i2v.json post /kling/kling-v3-i2v
[Core Function] Kling V3 I2V is the next-generation image-to-video model. [Strengths] It transforms static images into video with support for 4K resolution, 15-second durations, and native audio, providing superior motion and character expressiveness. [Best For] Highly recommended for: animating concept art, bringing portraits to life in 4K, and generating long 15s scenes from a single frame. [Limitations] Do NOT use this model if you need multimodal reference elements (like character consistency across shots) or multi-shot generation; use V3 Omni instead. [Routing] Use this model by default for high-quality single-image-to-video tasks.
# Kling V3 Omni I2V
Source: https://docs.modellix.ai/kling/kling-v3-omni-i2v
/media-model-api/kling/kling-i2v.json post /kling/kling-v3-omni-i2v
[Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and product look across the clip while supporting flexible duration and optional native audio. [Best For] Highly recommended for: character-consistent animation from design sheets, multi-reference product shots, comic or IP look locking, and I2V tasks where a single first frame is not enough. [Limitations] Do NOT use this model for simple one-image animation when cost or speed is the priority; use Kling V3 I2V or Kling V3 Turbo I2V. Do NOT use it when the primary input is text only; use Kling V3 Omni T2V or Kling V3 T2V. Do NOT use it for lip-sync avatar from audio alone; use Kling Avatar. [Routing] Choose Kling V3 Omni I2V when the user asks for Omni, multiple references, or strict visual consistency from images. Prefer Kling V3 I2V for standard single-image high quality; prefer Kling V3 Turbo I2V for fast or cheap single-image jobs.
# Kling V3 Omni Image
Source: https://docs.modellix.ai/kling/kling-v3-omni-image
/media-model-api/kling/kling-i2i.json post /kling/kling-v3-omni-image
[Core Function] Kling V3 Omni Image is a unified multimodal image generation endpoint (POST /images/omni-image). [Strengths] Multi-image reference, up to 4K, and optional series generation via result_type/series_amount. Use <<>> placeholders in prompt. [Best For] Character consistency, fusing multiple reference images, and comic/storyboard series. [Limitations] Element library IDs are not exposed. Default aspect_ratio is 1:1 when omitted. [Routing] Prefer for multi-image fusion or series; use kling-v3-t2i/i2i for standard single-shot generation.
# Kling V3 Omni T2V
Source: https://docs.modellix.ai/kling/kling-v3-omni-t2v
/media-model-api/kling/kling-t2v.json post /kling/kling-v3-omni-t2v
[Core Function] Kling V3 Omni T2V is a multimodal-leaning text-to-video model in the V3 family, oriented toward stronger semantic control and subject consistency in prompt-led generation. [Strengths] It targets high-fidelity cinematic clips with native audio options, flexible 3-15s duration, and better adherence when scenes demand coherent characters or multi-beat storytelling from text alone. [Best For] Highly recommended for: narrative T2V with recurring subjects, dialogue-aware scenes, brand or product continuity across beats, and premium short films where consistency matters more than raw throughput. [Limitations] Do NOT use this model if the user only needs the cheapest or fastest clip; prefer Kling V3 Turbo T2V. Do NOT use it when the workflow is image-first or needs multi-image references; use Kling V3 Omni I2V or Kling V3 I2V instead. Do NOT use it for deep physics-reasoning specialty tasks better served by Kling Video O1. [Routing] Choose Kling V3 Omni T2V when the user emphasizes Omni, consistency, multimodal quality, or complex text narratives. Prefer Kling V3 T2V as the default high-quality T2V baseline; prefer Kling V3 Turbo T2V when the user stresses speed, cost, or high-volume short-form output.
# Kling V3 Omni Video
Source: https://docs.modellix.ai/kling/kling-v3-omni-video
/media-model-api/kling/kling-v2v.json post /kling/kling-v3-omni-video
[Core Function] Kling V3 Omni Video V2V is a multimodal video-to-video endpoint that edits or restyles existing footage using prompt plus optional image and video references. [Strengths] It focuses on source fidelity and subject consistency for Omni-style edit workflows, combining prompt guidance with images and videos inputs so changes stay grounded in the original clip. [Best For] Highly recommended for: reference-faithful video edits, restyling existing takes, keeping characters or products consistent while changing motion or scene instructions, and short-form post workflows that start from real footage. [Limitations] Do NOT use this model for pure text-to-video from scratch; use Kling V3 T2V, Kling V3 Omni T2V, or Kling V3 Turbo T2V. Do NOT use it when you only have a still image and no source video; use Kling V3 I2V or Kling V3 Omni I2V. Do NOT use it for deep physics-reasoning generation better served by Kling Video O1. [Routing] Choose Kling V3 Omni Video V2V when the user already has video to edit or transform and mentions Omni or multimodal references. Prefer Kling Video O1 when reasoning-heavy generation is the goal rather than source-based editing.
# Kling V3 T2I
Source: https://docs.modellix.ai/kling/kling-v3-t2i
/media-model-api/kling/kling-t2i.json post /kling/kling-v3-t2i
[Core Function] Kling V3 T2I is the flagship text-to-image model (POST /images/generations, model_name=kling-v3). [Strengths] High aesthetic quality, prompt adherence, 1K/2K. [Best For] Concept art and photorealistic generation without a reference image. [Limitations] No reference image; for I2I use kling-v3-i2i; for multi-image/series use kling-v3-omni-image. [Routing] Default for Kling text-to-image.
# Kling V3 T2V
Source: https://docs.modellix.ai/kling/kling-v3-t2v
/media-model-api/kling/kling-t2v.json post /kling/kling-v3-t2v
[Core Function] Kling V3 T2V is the next-generation text-to-video base model. [Strengths] It natively supports generating ultra-long 15-second videos, 4K resolution, and synchronized native audio directly from text. [Best For] Highly recommended for: high-end cinematic creation, 4K video generation, and creating long-form scenes with integrated sound. [Limitations] Do NOT use this model if you need complex multi-shot narratives or deep physics reasoning; use V3 Omni or Video O1 respectively. [Routing] Use this model by default for high-quality text-to-video tasks that require up to 15 seconds, 4K resolution, or native audio without reference images.
# Kling V3 Turbo I2V
Source: https://docs.modellix.ai/kling/kling-v3-turbo-i2v
/media-model-api/kling/kling-i2v.json post /kling/kling-v3-turbo-i2v
[Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or product first-frame animation at practical resolutions. [Best For] Highly recommended for: animating stills for social ads, rapid keyframe iteration, talking-head starters from one photo, and bulk I2V jobs where latency and cost dominate. [Limitations] Do NOT use this model if the user needs multi-image references, element fusion, or Omni-class subject locking; use Kling V3 Omni I2V. Do NOT use it when maximum 4K cinematic quality is required; use Kling V3 I2V. Do NOT use it for effect templates; use Kling Video Effects. [Routing] Choose Kling V3 Turbo I2V when the user emphasizes speed or cost for single-image animation. Prefer Kling V3 I2V as the default high-quality I2V; prefer Kling V3 Omni I2V when multiple images or consistency-driven references are central.
# Kling V3 Turbo T2V
Source: https://docs.modellix.ai/kling/kling-v3-turbo-t2v
/media-model-api/kling/kling-t2v.json post /kling/kling-v3-turbo-t2v
[Core Function] Kling V3 Turbo T2V is a speed- and cost-optimized text-to-video model in the V3 family for fast short-form generation. [Strengths] It emphasizes lower latency and efficient throughput with native audio and improved lip-sync for talking-head style clips, typically targeting practical 720p/1080p short videos rather than maximum cinematic headroom. [Best For] Highly recommended for: rapid prototyping, social and ad iteration, batch short-form pipelines, and dialogue clips where turnaround time and unit cost matter most. [Limitations] Do NOT use this model if the user requires peak 4K cinematic fidelity, heavy multi-shot storyboard control, or maximum visual polish; use Kling V3 T2V or Kling V3 Omni T2V instead. Do NOT use it for image-conditioned animation; use Kling V3 Turbo I2V or Kling V3 I2V. [Routing] Choose Kling V3 Turbo T2V when the user says fast, quick, cheap, or high volume. Otherwise default to Kling V3 T2V for quality, or Kling V3 Omni T2V when consistency and Omni-class control are requested.
# Kling Video Effects
Source: https://docs.modellix.ai/kling/kling-video-effects
/media-model-api/kling/kling-i2v.json post /kling/kling-video-effects
[Core Function] Kling Video Effects applies predefined visual effects to images. [Strengths] It automatically transforms 1 or 2 images into engaging short videos using viral/predefined effect templates. [Best For] Highly recommended for: social media trends and quick visual gags. [Limitations] Do NOT use this model for custom narrative generation or prompt-based control. [Routing] Use this when the user explicitly requests an 'effect' or 'trend' template applied to their photos.
# Kling Video O1
Source: https://docs.modellix.ai/kling/kling-video-o1
/media-model-api/kling/kling-v2v.json post /kling/kling-video-o1
[Core Function] Kling Video O1 is the world's first reasoning-enhanced video model. [Strengths] It performs deep planning over the prompt before generation, delivering best-in-class physical consistency, complex motion logic, and strict adherence to long-form semantics. [Best For] Highly recommended for: complex physical interactions, logically demanding scenes, and prompts requiring deep reasoning. [Limitations] Do NOT use this model if you need multi-shot generation or 15-second durations (it is capped at 10s). [Routing] Route to this model when the prompt involves complex physics, logical sequences, or intricate physical interactions where standard models hallucinate.
# Kling Video O1 I2V
Source: https://docs.modellix.ai/kling/kling-video-o1-i2v
/media-model-api/kling/kling-i2v.json post /kling/kling-video-o1-i2v
[Core Function] Kling Video O1 I2V is the image-to-video slice of Kling O1 Omni Video. [Strengths] Reasoning-enhanced generation from 1-7 reference images, 720p/1080p, duration 3-10s (single image only 5 or 10). [Best For] Complex physical motion grounded in reference frames. [Limitations] No native audio; no aspect_ratio (follows first frame); duration capped at 10s. [Routing] Prefer this for O1-quality I2V; use Kling Video O1 V2V when a source video is required.
# Kling Video O1 T2V
Source: https://docs.modellix.ai/kling/kling-video-o1-t2v
/media-model-api/kling/kling-t2v.json post /kling/kling-video-o1-t2v
[Core Function] Kling Video O1 T2V is the text-to-video slice of Kling O1 Omni Video. [Strengths] Reasoning-enhanced prompt planning with 3-10s duration and 720p/1080p output. [Best For] Complex physical interactions and logically demanding scenes from text alone. [Limitations] No native audio; duration capped at 10s; no multi_shot. [Routing] Prefer this for O1-quality T2V; use Kling Video O1 V2V when a source video is required.
# Kolors Virtual Try On V1
Source: https://docs.modellix.ai/kling/kolors-virtual-try-on-v1
/media-model-api/kling/kling-i2i.json post /kling/kolors-virtual-try-on-v1
[Core Function] Kolors Virtual Try-On V1 is a legacy AI fashion and virtual try-on model. [Strengths] Maintained primarily for backward compatibility and stable API contracts. [Best For] Existing legacy e-commerce integrations that have not yet migrated to the newer try-on pipeline. [Limitations] Do NOT use this for new creations. It is a legacy model maintained for backward compatibility. [Routing] Only use if explicitly requested; otherwise use Kolors Virtual Try-On V1.5.
# Kolors Virtual Try On V1-5
Source: https://docs.modellix.ai/kling/kolors-virtual-try-on-v1-5
/media-model-api/kling/kling-i2i.json post /kling/kolors-virtual-try-on-v1-5
[Core Function] Kolors Virtual Try-On V1.5 is a specialized AI fashion model. [Strengths] It highly accurately applies garments (including Top+Bottom combinations) onto a person's image, preserving fabric texture and draping. [Best For] Highly recommended for: e-commerce virtual fitting rooms and fashion visualization. [Limitations] Do NOT use this model for general image editing. It is strictly for clothing try-on. [Routing] Use this model whenever the user asks to 'try on' clothes or apply a garment to a person.
# Use Modellix LLM with Claude Code
Source: https://docs.modellix.ai/llm/agent/claude-code
Configure Claude Code to call Modellix with ANTHROPIC_BASE_URL (no /v1), a Modellix API key, and anthropic/... models.
Point Claude Code at the Modellix LLM gateway using the Anthropic-compatible Messages API. Auth and base URL match the [Anthropic SDK](/llm/sdk/anthropic-sdk) pattern.
## Set Up Claude Code
Create a key in the [Modellix console](https://modellix.ai/console/api-key). Claude Code maps `ANTHROPIC_API_KEY` to the Anthropic-style `x-api-key` header.
```bash theme={null}
export ANTHROPIC_API_KEY="mdlx-xxxxxxxx"
export ANTHROPIC_BASE_URL="https://llm.modellix.ai"
```
| Setting | Value |
| -------- | -------------------------------------------------------------------------------------------------------------------- |
| Base URL | `https://llm.modellix.ai` (**without** `/v1`) |
| API key | Modellix API Key (`ANTHROPIC_API_KEY` → `x-api-key`) |
| Model | `anthropic/...` (for example `anthropic/claude-sonnet-5`). See [Models & Pricing](/llm/overview#models-and-pricing). |
Do not append `/v1` to `ANTHROPIC_BASE_URL`. Claude Code / the Anthropic SDK append `/v1/messages` themselves.
If the tool expects Bearer auth, use `ANTHROPIC_AUTH_TOKEN` instead of `ANTHROPIC_API_KEY`. Do not set both to different values.
Write the same values to `~/.claude/settings.json`:
```json theme={null}
{
"env": {
"ANTHROPIC_BASE_URL": "https://llm.modellix.ai",
"ANTHROPIC_API_KEY": "mdlx-xxxxxxxx"
}
}
```
Replace `mdlx-xxxxxxxx` with your Modellix API key.
## Session Headers
Claude Code may send its own session header (for example `X-Claude-Code-Session-Id`). Modellix also accepts `X-Mdlx-Session-Id`; if both are present, `X-Mdlx-Session-Id` wins. See [Session header](/llm/api/api#session-header).
## Related
* [Anthropic SDK](/llm/sdk/anthropic-sdk) — Messages examples in Python / TypeScript
* [Claude Agent SDK](/llm/sdk/claude-agent-sdk) — same env vars for the programmable agent harness
* [Create message](/llm/messages) — OpenAPI reference
* [LLM API guide](/llm/api/api) — auth, errors, and rate limits
# Use Modellix LLM with CodeBuddy
Source: https://docs.modellix.ai/llm/agent/codebuddy
Add Modellix as a CodeBuddy custom model with an OpenAI-compatible endpoint, models.json, and provider/name model IDs.
Configure [CodeBuddy](https://www.codebuddy.ai/) to call the Modellix LLM gateway. CodeBuddy supports OpenAI-compatible [custom models](https://www.codebuddy.ai/docs/cli/models) in the UI and in `models.json`, pointed at Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Enter that full string as the Model Name / `id`. In the UI, set Endpoint to `https://llm.modellix.ai/v1`. When editing `models.json` by hand, set `url` to the **full** Chat Completions path (`.../v1/chat/completions`). See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up CodeBuddy
Create a Modellix API key in the [console](https://modellix.ai/console/api-key) and export it for `models.json` env substitution:
```bash theme={null}
export MODELLIX_API_KEY="mdlx-xxxxxxxx"
```
You can also paste the key directly into the CodeBuddy UI or config files.
In CodeBuddy, open the model picker (for example **Default** under the input box), scroll to the bottom, and choose **Configure custom models** → **Add Model**.
Set **Provider** to **Custom**, then fill in:
| Field | Value |
| ----------------- | ---------------------------------------------------------------------------------------------------------- |
| Endpoint | `https://llm.modellix.ai/v1` |
| API Key | Your Modellix API Key |
| Model Name | Exact Modellix Model ID (for example `openai/gpt-5.5`) |
| Advanced Settings | Enable **Tool Calling** for agent tool use; enable Image Input / Reasoning only if the model supports them |
Save the model. Include `/v1` in the Endpoint so paths resolve to `/v1/chat/completions`.
CodeBuddy also stores custom models in a local file. Edit this when the UI is inconvenient, or to share a project-level config.
* User scope (macOS / Linux): `~/.codebuddy/models.json`
* User scope (Windows): `%USERPROFILE%\.codebuddy\models.json`
* Project scope (higher priority): `/.codebuddy/models.json`
Example:
```json theme={null}
{
"models": [
{
"id": "openai/gpt-5.5",
"name": "GPT 5.5",
"vendor": "Modellix",
"url": "https://llm.modellix.ai/v1/chat/completions",
"apiKey": "${MODELLIX_API_KEY}",
"maxInputTokens": 200000,
"maxOutputTokens": 8192,
"supportsToolCall": true
}
],
"availableModels": ["openai/gpt-5.5"]
}
```
| Field | Value |
| ----------------- | ------------------------------------------------ |
| `id` / `name` | Exact Modellix Model ID (`provider/name`) |
| `vendor` | `Modellix` |
| `url` | `https://llm.modellix.ai/v1/chat/completions` |
| `apiKey` | `${MODELLIX_API_KEY}` or your key string |
| `availableModels` | Must include each model `id` you want selectable |
When writing `models.json` by hand, use the **full** Chat Completions URL (including `/chat/completions`). A host-only value or `https://llm.modellix.ai/v1` without `/chat/completions` is less reliable than the full path. Do not point `url` at the media API host (`https://api.modellix.ai`).
Add more entries for other Modellix IDs (for example `anthropic/claude-sonnet-5`, `google/gemini-3.6-flash`) on the same gateway URL. Full catalog: [Models & Pricing](/llm/overview#models-and-pricing).
After saving, restart CodeBuddy or reopen the model picker so the list refreshes.
Open the model picker, choose your model under **Custom Models**, and send a short test message.
CodeBuddy may display custom models as `name:name`. That is normal if the selected model responds correctly.
The CLI package is **CodeBuddy Code** (`codebuddy`). It uses a separate Anthropic-compatible entry point from the desktop custom-model OpenAI path.
```bash theme={null}
export CODEBUDDY_API_KEY="mdlx-xxxxxxxx"
export CODEBUDDY_BASE_URL="https://llm.modellix.ai"
export CODEBUDDY_MODEL="anthropic/claude-sonnet-5"
codebuddy
```
Or persist in `~/.codebuddy/settings.json`:
```json theme={null}
{
"env": {
"CODEBUDDY_API_KEY": "mdlx-xxxxxxxx",
"CODEBUDDY_BASE_URL": "https://llm.modellix.ai",
"CODEBUDDY_MODEL": "anthropic/claude-sonnet-5"
}
}
```
| Setting | Value |
| -------------------- | --------------------------------------------- |
| `CODEBUDDY_BASE_URL` | `https://llm.modellix.ai` (**without** `/v1`) |
| `CODEBUDDY_API_KEY` | Modellix API Key |
| `CODEBUDDY_MODEL` | Prefer `anthropic/...` for this CLI path |
Do not set `CODEBUDDY_BASE_URL` to an OpenAI Chat Completions path such as `.../v1/chat/completions`. If you see 404s on the CLI, check that the base URL is the LLM host root without `/v1`.
## Troubleshooting
| Symptom | Check |
| ------------------------- | ----------------------------------------------------------------------------------------- |
| 401 / auth errors | Key is a Modellix API Key; `${MODELLIX_API_KEY}` is set if used in `models.json` |
| Model not found / 404 | Model Name / `id` is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5` |
| UI Endpoint errors | Endpoint includes `/v1`: `https://llm.modellix.ai/v1` |
| `models.json` path errors | `url` ends with `/v1/chat/completions`, not the media API host |
| CLI 404 | `CODEBUDDY_BASE_URL` is `https://llm.modellix.ai` without `/v1` or `/chat/completions` |
| Model missing after edit | Restart CodeBuddy or reopen the model picker; confirm `id` is listed in `availableModels` |
## Related
* [CodeBuddy models.json](https://www.codebuddy.ai/docs/cli/models) — custom model schema
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [Junie](/llm/agent/junie) — another client that uses a full Chat Completions URL
* [Claude Code](/llm/agent/claude-code) — Anthropic-style base URL without `/v1` (similar to CodeBuddy CLI)
# Use Modellix LLM with Codex CLI
Source: https://docs.modellix.ai/llm/agent/codex
Configure Codex to use Modellix by setting OPENAI_API_KEY and openai_base_url in ~/.codex/config.toml with openai/... models.
Point [Codex CLI](https://github.com/openai/codex) at the Modellix LLM gateway so chat and coding requests use your Modellix API key and `openai/...` models over OpenAI-compatible Chat Completions.
## Set Up Codex
Create a key in the [Modellix console](https://modellix.ai/console/api-key). It is not an OpenAI platform key.
Export your Modellix key as `OPENAI_API_KEY`:
```bash theme={null}
export OPENAI_API_KEY="mdlx-xxxxxxxx"
```
In `~/.codex/config.toml`, set the base URL and model. Newer Codex builds deprecate `OPENAI_BASE_URL` as an environment variable—prefer `openai_base_url` in config:
```toml theme={null}
model = "openai/gpt-5.5"
openai_base_url = "https://llm.modellix.ai/v1"
```
| Setting | Value |
| ----------------- | ------------------------------------------------------------------------------------------------------ |
| `openai_base_url` | `https://llm.modellix.ai/v1` |
| `OPENAI_API_KEY` | Modellix API Key |
| `model` | `openai/...` (for example `openai/gpt-5.5`). See [Models & Pricing](/llm/overview#models-and-pricing). |
After saving the config, run Codex as usual. Traffic should hit `https://llm.modellix.ai/v1` (typically Chat Completions). For protocol details, see [Chat Completions](/llm/chat-completions) and the [LLM API guide](/llm/api/api).
## Related
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` base URL for application code
* [OpenCode](/llm/agent/opencode) — another OpenAI-compatible client config
* [Cursor](/llm/ide/cursor) — IDE OpenAI-compatible provider setup
# Use Modellix LLM with DeepSeek Harness
Source: https://docs.modellix.ai/llm/agent/deepseek-harness
Connect DeepSeek Harness to Modellix LLM—install the dsh-modellix plugin, or add a custom llm-pi-ai provider in the Web UI or $DSH_HOME/settings.yaml.
Configure [DeepSeek Harness](https://deepseek-harness.github.io/deepseek-harness/) to call the Modellix LLM gateway. You can install the official [dsh-modellix plugin](/ways-to-use/deepseek-harness), or add Modellix as a [custom provider](https://deepseek-harness.github.io/deepseek-harness/guide/providers) on the `llm-pi-ai` adapter.
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Put the full string as each model `id` in your custom provider so DeepSeek Harness sends it unchanged in the request body. See [Models & Pricing](/llm/overview#models-and-pricing).
## Use the dsh-modellix Plugin
Instead of adding a custom provider by hand, install **[dsh-modellix](/ways-to-use/deepseek-harness)**. The plugin loads the live Modellix LLM catalog into the Harness model selector after you connect one API key.
That path also enables chat-first media generation and [Web Search](/api/web-search) / [Web Fetch](/api/web-fetch). If you only need the LLM gateway, keep using the custom provider steps below.
## Prerequisites
* DeepSeek Harness Web UI started via the [root README](https://github.com/deepseek-ai/deepseek-harness)
* A Modellix [API key](https://www.modellix.ai/console/api-key)—not a vendor platform key
Model changes take effect on the next request; you do not need to restart the server.
## Add Modellix as a Custom Provider
Open **Settings → Models**.
Choose **Add a custom provider** and fill in the form:
| Field | Value |
| ------------ | ---------------------------------------------------------------- |
| Provider ID | `modellix` (lowercase; permanent once saved) |
| Display name | `Modellix` |
| Base URL | `https://llm.modellix.ai/v1` |
| API protocol | `openai-completions` (OpenAI Chat Completions) |
| API key | Your Modellix [API key](https://www.modellix.ai/console/api-key) |
The Provider ID is permanent because requests, saved sessions, model defaults, and credential references use it. The display name, base URL, protocol, credential, and models remain editable.
Keys are write-only: after saving, the page shows only a redacted descriptor, and the key is stored in `$DSH_HOME/.credentials.yaml`.
Under **Model catalog**, choose **Fetch available models**. Modellix serves the OpenAI-compatible `GET /v1/models` endpoint, so the form lists the current Model IDs. Select the ones you want, or enter them by hand.
Each model `id` must be a full Modellix Model ID, for example:
```
openai/gpt-5.5
openai/gpt-5.6-sol
anthropic/claude-sonnet-5
google/gemini-3.6-flash
deepseek/deepseek-v4-flash
```
The provider is not stored until you save. Full catalog: [Models & Pricing](/llm/overview#models-and-pricing).
Save the provider, then choose a Modellix model in the model picker. Selecting a model also makes it the default for new sessions.
## Configure Directly in `settings.yaml`
You can declare the same provider in `$DSH_HOME/settings.yaml` instead of the form. Credentials stay out of this file—`apiKeyEnv` is a reference resolved per request (or provide the key through the Models page):
```yaml theme={null}
llm-pi-ai:
providers:
modellix:
apiKeyEnv: MODELLIX_API_KEY
api: openai-completions
baseURL: https://llm.modellix.ai/v1
models:
- id: openai/gpt-5.5
name: GPT 5.5
- id: openai/gpt-5.6-sol
name: GPT 5.6 Sol
- id: anthropic/claude-sonnet-5
name: Claude Sonnet 5
- id: google/gemini-3.6-flash
name: Gemini 3.6 Flash
- id: deepseek/deepseek-v4-flash
name: DeepSeek V4 Flash
```
| Field | Meaning |
| ----------- | --------------------------------------------------------------------------- |
| `apiKeyEnv` | Environment variable holding the Modellix API key |
| `api` | Wire protocol; `openai-completions` for Chat Completions |
| `baseURL` | `https://llm.modellix.ai/v1` |
| `models` | List of model entries; `id` is the exact Modellix Model ID sent on the wire |
The `models` list replaces the route's catalog, so every model the route should serve must appear in it—an entry with only `id` is enough. `name` is optional and shown in selectors. Add or remove entries as needed; a model the route does not configure fails with `UNKNOWN_MODEL`.
## Use the Anthropic Messages Protocol
Most setups should keep `openai-completions`: every Modellix model, including `anthropic/...`, works over Chat Completions at `https://llm.modellix.ai/v1`. If you prefer the Anthropic wire protocol, set `api: anthropic-messages` and `baseURL: https://llm.modellix.ai` (no `/v1`)—but model discovery (`Fetch available models`) only reads OpenAI-compatible `GET /models` endpoints, so enter models by hand in that case.
## Troubleshooting
| Check | Detail |
| ---------------------------------- | ---------------------------------------------------------------------------------------------------------------------- |
| `MISSING_CREDENTIAL` | Store the provider key through the Models page, or provide the referenced environment variable |
| `UNKNOWN_MODEL` | Select a configured model, or add the missing model to the custom provider |
| Fetch available models returns 401 | Check the key. Model discovery calls `GET /v1/models` on `https://llm.modellix.ai/v1`; enter models manually if needed |
| Request rejected with an image | Modellix LLM is a text gateway—use text-only prompts; do not attach images |
## Related
* [DeepSeek Harness Plugin](/ways-to-use/deepseek-harness) — install `dsh-modellix` for LLM, media, and Web tools
* [DeepSeek Harness providers guide](https://deepseek-harness.github.io/deepseek-harness/guide/providers) — official custom provider setup
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols, auth, and curl examples
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` Chat Completions gateway
# Use Modellix LLM with Grok Build
Source: https://docs.modellix.ai/llm/agent/grok-build
Point Grok Build at the Modellix LLM gateway with custom models in ~/.grok/config.toml, MODELLIX_API_KEY, and provider/name model IDs.
Configure [Grok Build](https://docs.x.ai/build/overview) to call the Modellix LLM gateway. Grok Build is xAI's coding agent (interactive TUI, headless CLI, or ACP). It has no built-in Modellix provider, so add [custom models](https://docs.x.ai/build/overview) in `~/.grok/config.toml` that point at Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Put that full string in each custom model's `model` field so Grok Build sends it unchanged in the request body. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up Grok Build
```bash Mac / Linux theme={null}
curl -fsSL https://x.ai/cli/install.sh | bash
```
```powershell Windows theme={null}
irm https://x.ai/cli/install.ps1 | iex
```
Create a key in the [Modellix console](https://modellix.ai/console/api-key) and export it. Grok Build reads the key from the environment variable named in `env_key`:
```bash theme={null}
export MODELLIX_API_KEY="mdlx-xxxxxxxx"
```
Add the same line to your shell profile (for example `~/.zshrc`), then run `source ~/.zshrc`.
Use a Modellix API key, not an xAI platform key. A built-in Grok catalog model with `XAI_API_KEY` still talks to `https://api.x.ai/v1`, not Modellix.
Edit the **user** config at `~/.grok/config.toml` (Windows: `%USERPROFILE%\.grok\config.toml`). If `GROK_HOME` is set, the file is `$GROK_HOME/config.toml` instead.
Merge the `[model."..."]` tables into an existing file. Do not put custom models in a project `.grok/config.toml`—that file only accepts MCP servers, plugins, and permission rules.
```toml theme={null}
[model."openai/gpt-5.5"]
model = "openai/gpt-5.5"
base_url = "https://llm.modellix.ai/v1"
name = "GPT 5.5"
env_key = "MODELLIX_API_KEY"
api_backend = "chat_completions"
[model."openai/gpt-5.6-sol"]
model = "openai/gpt-5.6-sol"
base_url = "https://llm.modellix.ai/v1"
name = "GPT 5.6 Sol"
env_key = "MODELLIX_API_KEY"
api_backend = "chat_completions"
[model."anthropic/claude-sonnet-5"]
model = "anthropic/claude-sonnet-5"
base_url = "https://llm.modellix.ai/v1"
name = "Claude Sonnet 5"
env_key = "MODELLIX_API_KEY"
api_backend = "chat_completions"
[model."google/gemini-3.6-flash"]
model = "google/gemini-3.6-flash"
base_url = "https://llm.modellix.ai/v1"
name = "Gemini 3.6 Flash"
env_key = "MODELLIX_API_KEY"
api_backend = "chat_completions"
[model."xai/grok-4.6"]
model = "xai/grok-4.6"
base_url = "https://llm.modellix.ai/v1"
name = "Grok 4.6"
env_key = "MODELLIX_API_KEY"
api_backend = "chat_completions"
[models]
default = "openai/gpt-5.5"
```
Add or remove `[model."..."]` tables as needed. Quote the table key when the local ID contains `/` or `.`.
| Field | Value |
| -------------------------- | --------------------------------------------------------------------------- |
| Table key (`[model."id"]`) | Local picker ID used with `-m` and `/model` |
| `model` | Exact Modellix Model ID (`provider/name`) sent to the API |
| `base_url` | `https://llm.modellix.ai/v1` (include `/v1`) |
| `name` | Label shown in the model picker |
| `env_key` | `MODELLIX_API_KEY` (prefer this over an inline `api_key`) |
| `api_backend` | `chat_completions` for [`POST /v1/chat/completions`](/llm/chat-completions) |
| `[models].default` | Local picker ID of the Modellix model to use for new sessions |
Setting `models.default` to a Modellix ID avoids falling back to a built-in xAI catalog model that requires `XAI_API_KEY`. Full catalog: [Models & Pricing](/llm/overview#models-and-pricing).
You do **not** need a separate Anthropic base URL. Register `anthropic/...` and `google/...` IDs on the same Modellix `base_url`; traffic goes through OpenAI-compatible Chat Completions. Native Anthropic Messages (`api_backend = "messages"` without `/v1`) is for the [Anthropic SDK](/llm/sdk/anthropic-sdk) and [Claude Code](/llm/agent/claude-code).
If you prefer Responses instead of Chat Completions, set `api_backend = "responses"` and keep `base_url = "https://llm.modellix.ai/v1"`. See [Responses](/llm/responses). Most setups should keep `chat_completions`.
Confirm Grok Build loaded the user config, then start a session from your project directory:
```bash theme={null}
grok inspect
cd your-project
grok
```
In the TUI, switch models with `/model ` (for example `/model openai/gpt-5.5`). Headless:
```bash theme={null}
grok -p "Explain this repo" -m "openai/gpt-5.5"
```
You can also set `GROK_DEFAULT_MODEL` for the current session (same idea as `-m`).
## Troubleshooting
| Symptom | Check |
| -------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| 401 / auth errors | `MODELLIX_API_KEY` is a valid Modellix key, and `env_key` matches that variable name |
| Model not found / 404 | `model` is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5`; ID is in the live catalog |
| Wrong path / URL errors | `base_url` is `https://llm.modellix.ai/v1` (include `/v1`), not the media API host (`https://api.modellix.ai`) |
| Traffic hits `api.x.ai` | You selected a custom Modellix table, not a built-in `grok-*` catalog model; `models.default` is a Modellix ID |
| Custom model missing from picker | Entry is in **user** `~/.grok/config.toml` (or `$GROK_HOME/config.toml`); run `grok inspect` to see loaded config sources |
| First launch opens a browser | Built-in xAI auth. Set `models.default` to a Modellix ID and export `MODELLIX_API_KEY` so sessions do not need `XAI_API_KEY` |
## Related
* [Grok Build overview](https://docs.x.ai/build/overview) — install, custom models, TUI and headless usage
* [Grok Build settings](https://docs.x.ai/build/settings) — `config.toml` scopes and `[model.]` fields
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` Chat Completions gateway
* [OpenCode](/llm/agent/opencode) · [Codex](/llm/agent/codex) — other OpenAI-compatible agent setups
# Use Modellix LLM with Hermes Agent
Source: https://docs.modellix.ai/llm/agent/hermes
Point Hermes Agent at the Modellix LLM gateway with a Custom Endpoint, config.yaml provider custom, and provider/name model IDs.
Configure [Hermes Agent](https://hermes-agent.nousresearch.com/docs) by Nous Research to call the Modellix LLM gateway. Hermes does not ship a built-in Modellix provider, so use a [Custom Endpoint](https://hermes-agent.nousresearch.com/docs/integrations/providers) that points at Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`). For Hermes media skills (image, video, speech), see [Skill](/ways-to-use/skill).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Enter that full string as the Hermes model name. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up Hermes
Create a Modellix API key in the [console](https://modellix.ai/console/api-key) and store it securely. You will paste it into the Hermes Custom Endpoint setup (or into `~/.hermes/.env`).
From your terminal (outside an active Hermes session), run:
```bash theme={null}
hermes model
```
Select **Custom endpoint (self-hosted / VLLM / etc.)** and enter:
| Field | Value |
| ------------ | ------------------------------------------------------ |
| API base URL | `https://llm.modellix.ai/v1` |
| API key | Your Modellix API Key |
| Model name | Exact Modellix Model ID (for example `openai/gpt-5.5`) |
Include `/v1` in the base URL so paths resolve to `/v1/chat/completions`.
Use **Custom endpoint**, not built-in OpenRouter, OpenAI, or Anthropic providers. Those routes send traffic to vendor endpoints instead of Modellix.
The wizard writes provider settings for you. To configure manually (or to verify), edit these files.
`~/.hermes/.env` (secrets only):
```bash theme={null}
MODELLIX_API_KEY=mdlx-xxxxxxxx
```
`~/.hermes/config.yaml` (model and endpoint):
```yaml theme={null}
model:
provider: custom
base_url: https://llm.modellix.ai/v1
default: openai/gpt-5.5
api_key: mdlx-xxxxxxxx
```
Replace `mdlx-xxxxxxxx` with your Modellix API key. Prefer keeping the key in `.env` when your Hermes build stores credentials there after `hermes model`; keep `config.yaml` as the source of truth for `provider`, `base_url`, and `default`.
| Field | Value |
| ---------- | ----------------------------------------------------- |
| `provider` | `custom` |
| `base_url` | `https://llm.modellix.ai/v1` |
| `default` | Exact Modellix Model ID (`provider/name`) |
| `api_key` | Modellix API Key (or rely on wizard / `.env` storage) |
To switch models later, change `model.default` (for example to `anthropic/claude-sonnet-5` or `google/gemini-3.6-flash`), or inside a session use `/model custom:`. Full catalog: [Models & Pricing](/llm/overview#models-and-pricing).
You do **not** need a separate Anthropic base URL. Use `anthropic/...` and `google/...` IDs on the same Custom Endpoint; traffic goes through OpenAI-compatible Chat Completions on `https://llm.modellix.ai/v1`.
Hermes separates secrets from settings: API keys belong in `~/.hermes/.env`; model, provider, and base URL belong in `~/.hermes/config.yaml`.
Start a Hermes session so it loads the updated config:
```bash theme={null}
hermes
```
Or use the TUI:
```bash theme={null}
hermes --tui
```
You can also start a chat with:
```bash theme={null}
hermes chat
```
## Troubleshooting
| Symptom | Check |
| --------------------------------- | ------------------------------------------------------------------------------------------------- |
| 401 / no API key | Modellix key is in `~/.hermes/.env` or set via `hermes model`; re-run `hermes model` if needed |
| Model not found / 404 | `model.default` is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5`; ID is in the live catalog |
| Context length rejected | Hermes requires at least **64K** context; use a Modellix catalog model that meets that floor |
| Traffic hits OpenRouter or OpenAI | `model.provider` is `custom` and `base_url` is `https://llm.modellix.ai/v1` |
| Wrong path / URL errors | Base URL includes `/v1` and is not the media API host (`https://api.modellix.ai`) |
## Related
* [Hermes AI Providers](https://hermes-agent.nousresearch.com/docs/integrations/providers) — Custom Endpoint reference
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [OpenClaw](/llm/agent/openclaw) · [OpenCode](/llm/agent/opencode) — other OpenAI-compatible custom setups
* [Skill](/ways-to-use/skill) — Hermes media skill install
# Use Modellix LLM with Junie CLI
Source: https://docs.modellix.ai/llm/agent/junie
Add Modellix as a Junie custom LLM profile with OpenAICompletion, a full chat completions URL, and provider/name model IDs.
Configure [Junie CLI](https://junie.jetbrains.com/docs/junie-cli.html) by JetBrains to call the Modellix LLM gateway. Junie does not ship a built-in Modellix BYOK provider, so add a [custom LLM profile](https://junie.jetbrains.com/docs/custom-llm-models.html) that points at Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Put that full string in the profile `id` field. Junie requires `baseUrl` to be the **full** Chat Completions URL (including `/v1/chat/completions`), not a host-only or `/v1` base. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up Junie
Create a Modellix API key in the [console](https://modellix.ai/console/api-key) and export it so the profile can resolve `${MODELLIX_API_KEY}`:
```bash theme={null}
export MODELLIX_API_KEY="mdlx-xxxxxxxx"
```
Add the same line to your shell profile (for example `~/.zshrc`) for persistent use.
Create `~/.junie/models/modellix.json` (or `.junie/models/modellix.json` in a project):
```json theme={null}
{
"id": "openai/gpt-5.5",
"baseUrl": "https://llm.modellix.ai/v1/chat/completions",
"apiType": "OpenAICompletion",
"apiKey": "${MODELLIX_API_KEY}"
}
```
The filename without `.json` is the profile ID (`modellix`). Junie discovers profiles from `$JUNIE_HOME/models/*.json` (default `~/.junie/models/`) and project-level `.junie/models/*.json`.
| Field | Value |
| --------- | ----------------------------------------------------- |
| `id` | Exact Modellix Model ID (`provider/name`) |
| `baseUrl` | `https://llm.modellix.ai/v1/chat/completions` |
| `apiType` | `OpenAICompletion` |
| `apiKey` | `${MODELLIX_API_KEY}` (resolved from the environment) |
Use the full Chat Completions URL. A host-only value or `https://llm.modellix.ai/v1` without `/chat/completions` will fail. Do not point `baseUrl` at the media API host (`https://api.modellix.ai`).
Optionally add a cheaper helper model for Junie's internal tasks:
```json theme={null}
{
"id": "openai/gpt-5.5",
"baseUrl": "https://llm.modellix.ai/v1/chat/completions",
"apiType": "OpenAICompletion",
"apiKey": "${MODELLIX_API_KEY}",
"fasterModel": {
"id": "google/gemini-3.6-flash"
}
}
```
Change `id` (and `fasterModel.id`) to any Modellix catalog ID. Full list: [Models & Pricing](/llm/overview#models-and-pricing).
Start Junie with the custom profile:
```bash theme={null}
junie --model custom:modellix
```
Custom models use the `custom:` prefix plus the profile filename (without `.json`). In the interactive TUI, pick the Modellix profile from the model list or use `/model`.
Do not use built-in BYOK flags such as `--openrouter-api-key` or `JUNIE_OPENAI_API_KEY` for Modellix. Those route to vendor providers, not your custom profile.
From your project directory:
```bash theme={null}
cd /path/to/your/project
junie --model custom:modellix
```
For a one-shot (headless) task:
```bash theme={null}
junie --model custom:modellix "Review and fix any code quality issues in the latest commit"
```
## Troubleshooting
| Symptom | Check |
| ------------------------------- | ---------------------------------------------------------------------------------------------------------------- |
| 401 / missing API key | `MODELLIX_API_KEY` is set in the environment; profile uses `${MODELLIX_API_KEY}` (unset vars fail profile load) |
| Model not found / 404 | Profile `id` is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5`; ID is in the live catalog |
| Wrong path / URL errors | `baseUrl` is `https://llm.modellix.ai/v1/chat/completions` (include `/chat/completions`), not the media API host |
| Profile missing from model list | File is under `~/.junie/models/` or `.junie/models/`; selection uses `custom:` + filename without `.json` |
## Related
* [Junie Custom LLMs](https://junie.jetbrains.com/docs/custom-llm-models.html) — profile schema and discovery
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [Hermes](/llm/agent/hermes) · [OpenClaw](/llm/agent/openclaw) — other OpenAI-compatible custom setups
# Use Modellix LLM with Kilo Code
Source: https://docs.modellix.ai/llm/agent/kilo-code
Add Modellix as a Kilo Code custom OpenAI Compatible provider with kilo.jsonc, baseURL, and provider/name model IDs.
Configure [Kilo Code](https://kilo.ai/docs/getting-started) to call the Modellix LLM gateway. Modellix is not a built-in Kilo provider — use a [custom OpenAI Compatible](https://kilo.ai/docs/ai-providers/openai-compatible) provider pointed at Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Put that full string in the custom provider `models` map (and as the Model ID in the UI) so Kilo sends it unchanged in the request body. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up Kilo Code
Install the [VS Code extension](https://marketplace.visualstudio.com/items?itemName=kilocode.kilo-code) (also works in Cursor, Windsurf, and other VS Code forks):
```bash theme={null}
code --install-extension kilocode.kilo-code
```
Or install the CLI:
```bash theme={null}
npm install -g @kilocode/cli
kilo --version
```
Export your Modellix API key from the [console](https://modellix.ai/console/api-key):
```bash theme={null}
export MODELLIX_API_KEY="mdlx-xxxxxxxx"
```
Add the line to your shell profile (for example `~/.zshrc`), then run `source ~/.zshrc`.
Prefer the **global** config path `~/.config/kilo/kilo.jsonc` for `{env:MODELLIX_API_KEY}`. Project-level `kilo.jsonc` does not resolve `{env:...}` references (Kilo security restriction).
In Kilo Code, open **Settings** (gear icon) → **Providers**, scroll to the bottom, and click **Custom provider**.
Fill in:
| Field | Value |
| ------------ | --------------------------------------------------------------------------- |
| Provider ID | `modellix` (lowercase; becomes the `provider_id` in `provider_id/model_id`) |
| Display name | `Modellix` |
| Provider API | **OpenAI Compatible** (Chat Completions) |
| Base URL | `https://llm.modellix.ai/v1` |
| API key | Your Modellix API Key (or leave empty and use `options.apiKey` in config) |
| Models | Add Model IDs manually (exact `provider/name` strings) |
Submit to save. Include `/v1` in the Base URL so paths resolve to `/v1/chat/completions`.
Kilo may try to auto-fetch models from `/v1/models`. If detection fails, enter Model IDs manually — that is expected for gateways without an OpenAI models list endpoint.
Do **not** pick a third-party built-in provider (for example AIhubmix) and paste a Modellix key there. Use a dedicated Custom provider with the Modellix Base URL.
Create or update `~/.config/kilo/kilo.jsonc` (Windows: `%USERPROFILE%\.config\kilo\kilo.jsonc`):
```jsonc theme={null}
{
"$schema": "https://app.kilo.ai/config.json",
"model": "modellix/openai/gpt-5.5",
"provider": {
"modellix": {
"npm": "@ai-sdk/openai-compatible",
"name": "Modellix",
"options": {
"baseURL": "https://llm.modellix.ai/v1",
"apiKey": "{env:MODELLIX_API_KEY}"
},
"models": {
"openai/gpt-5.5": {
"name": "GPT 5.5",
"tool_call": true,
"limit": {
"context": 200000,
"output": 8192
}
},
"openai/gpt-5.6-sol": {
"name": "GPT 5.6 Sol",
"tool_call": true,
"limit": {
"context": 200000,
"output": 8192
}
},
"openai/gpt-5.6-terra": {
"name": "GPT 5.6 Terra",
"tool_call": true,
"limit": {
"context": 200000,
"output": 8192
}
},
"openai/gpt-5.6-luna": {
"name": "GPT 5.6 Luna",
"tool_call": true,
"limit": {
"context": 200000,
"output": 8192
}
},
"anthropic/claude-sonnet-5": {
"name": "Claude Sonnet 5",
"tool_call": true,
"limit": {
"context": 200000,
"output": 8192
}
},
"google/gemini-3.6-flash": {
"name": "Gemini 3.6 Flash",
"tool_call": true,
"limit": {
"context": 200000,
"output": 8192
}
},
"xai/grok-4.5": {
"name": "Grok 4.5",
"tool_call": true,
"limit": {
"context": 200000,
"output": 8192
}
}
}
}
}
}
```
| Field | Value |
| -------------------------------- | ------------------------------------------------------------------ |
| Provider ID | Any string (for example `modellix`); must match the UI Provider ID |
| `npm` | `@ai-sdk/openai-compatible` for Chat Completions |
| `options.baseURL` | `https://llm.modellix.ai/v1` |
| `options.apiKey` | `{env:MODELLIX_API_KEY}` in global config |
| `models` keys | Exact Modellix Model IDs (`provider/name`) |
| Top-level `model` | `modellix/` (provider ID + model key) |
| `tool_call` | `true` for agent tool use |
| `limit.context` / `limit.output` | Set explicitly so Kilo can compact context |
Add or remove entries under `models` as needed. Full catalog: [Models & Pricing](/llm/overview#models-and-pricing).
For custom models, always set `limit.context` and `limit.output`. If both are unset and the ID is not in Kilo's catalog, context compaction is disabled and conversations can grow until the provider rejects the request.
In the IDE, open the model picker and choose a Modellix model (format `modellix/openai/gpt-5.5`).
From the CLI:
```bash theme={null}
kilo models
```
You do **not** need a separate Anthropic Base URL without `/v1` for these IDs. Register `anthropic/...` and `google/...` on the same Modellix OpenAI Compatible provider; traffic goes through Chat Completions on `https://llm.modellix.ai/v1`.
Native Anthropic Messages is for the [Anthropic SDK](/llm/sdk/anthropic-sdk) and [Claude Code](/llm/agent/claude-code), not required here.
## Troubleshooting
| Symptom | Check |
| -------------------------------- | ----------------------------------------------------------------------------------------------- |
| 401 / Invalid API Key | `MODELLIX_API_KEY` is a valid Modellix key; global `kilo.jsonc` uses `{env:MODELLIX_API_KEY}` |
| Model not found / 404 | Model key is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5` |
| Wrong path / connection errors | `baseURL` is `https://llm.modellix.ai/v1` (include `/v1`); do not use `https://api.modellix.ai` |
| Auto-detect finds no models | Enter Model IDs manually; auto-fetch needs `/v1/models` |
| `{env:...}` ignored | Credentials must live in `~/.config/kilo/kilo.jsonc`, not a project-level file |
| Tools / agent steps fail | Set `"tool_call": true` on the model entry |
| Context grows without compacting | Set `limit.context` and `limit.output` on each custom model |
## Related
* [Kilo OpenAI Compatible](https://kilo.ai/docs/ai-providers/openai-compatible) — custom provider UI and Base URL
* [Kilo Custom Models](https://kilo.ai/docs/code-with-ai/agents/custom-models) — `kilo.jsonc` schema, limits, and `id` mapping
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [OpenCode](/llm/agent/opencode) — similar `@ai-sdk/openai-compatible` custom provider setup
# Use Modellix LLM with OpenClaw 🦞
Source: https://docs.modellix.ai/llm/agent/openclaw
Add Modellix as a custom OpenAI-compatible provider in OpenClaw with models.providers, openai-completions, and provider/name model IDs.
Configure [OpenClaw](https://docs.openclaw.ai) to call the Modellix LLM gateway. OpenClaw does not ship a built-in Modellix provider, so add a [custom provider](https://docs.openclaw.ai/concepts/model-providers) that points at Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`). For OpenClaw media skills and plugins (image, video, speech), see [Plugin](/ways-to-use/plugin) and [Skill](/ways-to-use/skill).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). In OpenClaw, register that full string under your custom provider, then select `modellix/` (for example `modellix/openai/gpt-5.5`). See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up OpenClaw
Create a Modellix API key in the [console](https://modellix.ai/console/api-key) and store it securely. You will set it as `MODELLIX_API_KEY` in the OpenClaw config below.
Create or update `~/.openclaw/openclaw.json` with a full config that defines the Modellix provider, models, and a primary agent model:
```json theme={null}
{
"env": {
"MODELLIX_API_KEY": "mdlx-xxxxxxxx"
},
"models": {
"mode": "merge",
"providers": {
"modellix": {
"baseUrl": "https://llm.modellix.ai/v1",
"apiKey": "${MODELLIX_API_KEY}",
"api": "openai-completions",
"models": [
{ "id": "openai/gpt-5.5", "name": "GPT 5.5" },
{ "id": "openai/gpt-5.6-sol", "name": "GPT 5.6 Sol" },
{ "id": "openai/gpt-5.6-terra", "name": "GPT 5.6 Terra" },
{ "id": "openai/gpt-5.6-luna", "name": "GPT 5.6 Luna" },
{ "id": "anthropic/claude-sonnet-5", "name": "Claude Sonnet 5" },
{ "id": "google/gemini-3.6-flash", "name": "Gemini 3.6 Flash" },
{ "id": "xai/grok-4.5", "name": "Grok 4.5" }
]
}
}
},
"agents": {
"defaults": {
"model": {
"primary": "modellix/openai/gpt-5.5"
},
"models": {
"modellix/openai/gpt-5.5": {}
}
}
}
}
```
Replace `mdlx-xxxxxxxx` with your Modellix API key. Merge these fields into an existing config if you already use OpenClaw.
| Field | Value |
| ------------------------------- | ---------------------------------------------------------------- |
| `Provider ID` | `modellix` (custom; do not use built-in `openai` or `anthropic`) |
| `baseUrl` | `https://llm.modellix.ai/v1` |
| `api` | `openai-completions` |
| `apiKey` | `${MODELLIX_API_KEY}` (or your key string) |
| `models[].id` | Exact Modellix Model IDs (`provider/name`) |
| `agents.defaults.model.primary` | `modellix/` |
Use a custom provider key such as `modellix`. Overwriting OpenClaw's built-in `openai` or `anthropic` providers sends traffic to the vendor endpoints instead of Modellix.
Add or remove entries under `models.providers.modellix.models` as needed. Full catalog: [Models & Pricing](/llm/overview#models-and-pricing).
The example above sets the primary model to `modellix/openai/gpt-5.5`. To switch models, update both fields:
* `agents.defaults.model.primary` — OpenClaw ref: `modellix/` + Modellix Model ID
* `agents.defaults.models` — include the same ref so the agent can use it
Examples:
| Modellix Model ID | OpenClaw primary |
| --------------------------- | ------------------------------------ |
| `openai/gpt-5.5` | `modellix/openai/gpt-5.5` |
| `anthropic/claude-sonnet-5` | `modellix/anthropic/claude-sonnet-5` |
| `google/gemini-3.6-flash` | `modellix/google/gemini-3.6-flash` |
You do **not** need a separate Anthropic base URL for OpenClaw. Register `anthropic/...` and `google/...` IDs on the same Modellix custom provider; traffic goes through OpenAI-compatible Chat Completions on `https://llm.modellix.ai/v1`.
Do not set `agents.defaults.model.primary` to a built-in ref such as `openai/gpt-5.5`. That routes through OpenClaw's native OpenAI provider, not Modellix.
Start or restart the OpenClaw gateway so it loads the updated config:
```bash theme={null}
openclaw gateway run
```
Optionally list models to confirm Modellix entries appear:
```bash theme={null}
openclaw models list
```
## Troubleshooting
| Symptom | Check |
| --------------------------------- | ----------------------------------------------------------------------------------------------- |
| 401 / auth errors | `MODELLIX_API_KEY` is a valid Modellix key; `${MODELLIX_API_KEY}` resolves in config |
| Model not found / 404 | `models[].id` is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5`; ID is in the live catalog |
| Traffic hits OpenAI, not Modellix | Primary uses `modellix/...`, not built-in `openai/...`; provider key is `modellix` |
| Wrong path / URL errors | `baseUrl` is `https://llm.modellix.ai/v1` (include `/v1`), not the media API host |
| Model ignored by agent | Add the OpenClaw ref under both `models.providers.modellix.models` and `agents.defaults.models` |
## Related
* [OpenClaw model providers](https://docs.openclaw.ai/concepts/model-providers) — custom provider reference
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [OpenCode](/llm/agent/opencode) — another OpenAI-compatible custom provider setup
* [Plugin](/ways-to-use/plugin) · [Skill](/ways-to-use/skill) — OpenClaw media capabilities
# Use Modellix LLM with OpenCode
Source: https://docs.modellix.ai/llm/agent/opencode
Add Modellix as a custom OpenAI-compatible provider in OpenCode with @ai-sdk/openai-compatible, baseURL, and provider/name model IDs.
Configure [OpenCode](https://opencode.ai) to call the Modellix LLM gateway. The recommended approach is a [custom provider](https://opencode.ai/docs/providers/#custom-provider) using `@ai-sdk/openai-compatible` (Chat Completions at `/v1/chat/completions`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Put that full string in the custom provider `models` map so OpenCode sends it unchanged in the request body. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up OpenCode
In OpenCode, run `/connect`, choose **Other**, and enter a provider ID such as `modellix`. Paste your Modellix API key from the [console](https://modellix.ai/console/api-key).
The provider ID from `/connect` must match the key under `provider` in `opencode.json`.
Alternatively, set the key in config with `options.apiKey` (for example `{env:MODELLIX_API_KEY}`).
Create or update `opencode.json` (project) or `~/.config/opencode/opencode.json`:
```json theme={null}
{
"$schema": "https://opencode.ai/config.json",
"model": "modellix/openai/gpt-5.5",
"provider": {
"modellix": {
"npm": "@ai-sdk/openai-compatible",
"name": "Modellix",
"options": {
"baseURL": "https://llm.modellix.ai/v1",
"apiKey": "{env:MODELLIX_API_KEY}"
},
"models": {
"openai/gpt-5.5": {
"name": "GPT 5.5"
},
"openai/gpt-5.6-sol": {
"name": "GPT 5.6 Sol"
},
"openai/gpt-5.6-terra": {
"name": "GPT 5.6 Terra"
},
"openai/gpt-5.6-luna": {
"name": "GPT 5.6 Luna"
},
"anthropic/claude-sonnet-5": {
"name": "Claude Sonnet 5"
},
"google/gemini-3.6-flash": {
"name": "Gemini 3.6 Flash"
},
"xai/grok-4.5": {
"name": "Grok 4.5"
}
}
}
}
}
```
| Field | Value |
| ----------------- | ---------------------------------------------------------- |
| Provider ID | Any string (for example `modellix`); must match `/connect` |
| `npm` | `@ai-sdk/openai-compatible` for Chat Completions |
| `options.baseURL` | `https://llm.modellix.ai/v1` |
| `options.apiKey` | Modellix API Key (or rely on `/connect` auth) |
| `models` keys | Exact Modellix Model IDs (`provider/name`) |
| Top-level `model` | `modellix/` (provider ID + model key) |
Add or remove entries under `models` as needed. Full catalog: [Models & Pricing](/llm/overview#models-and-pricing).
If you prefer Responses instead of Chat Completions, set `npm` to `@ai-sdk/openai` (see [OpenCode custom provider](https://opencode.ai/docs/providers/#custom-provider)). Most setups should keep `@ai-sdk/openai-compatible`.
Run `/models` and select a Modellix model.
You do **not** need a separate Anthropic `baseURL` without `/v1` for OpenCode. Register `anthropic/...` and `google/...` IDs on the same Modellix custom provider above; traffic goes through OpenAI-compatible Chat Completions on `https://llm.modellix.ai/v1`.
Native Anthropic Messages (`ANTHROPIC_BASE_URL` without `/v1`) is for the [Anthropic SDK](/llm/sdk/anthropic-sdk) and [Claude Code](/llm/agent/claude-code), not required here.
## Troubleshooting
| Check | Detail |
| ------------- | ----------------------------------------------------------------------- |
| Provider ID | `/connect` ID matches `provider.modellix` (or your chosen ID) in config |
| `npm` package | Use `@ai-sdk/openai-compatible` for `/v1/chat/completions` |
| `baseURL` | Must be `https://llm.modellix.ai/v1` |
| Model keys | Must be full IDs like `openai/gpt-5.5`, not bare `gpt-5.5` |
| Auth | `opencode auth list`, or `options.apiKey` |
## Related
* [OpenCode custom provider](https://opencode.ai/docs/providers/#custom-provider) — official setup
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` Chat Completions gateway
# Use Modellix LLM with Pi Agent
Source: https://docs.modellix.ai/llm/agent/pi
Add Modellix as a custom OpenAI-compatible provider in Pi with models.json, openai-completions, and provider/name model IDs.
Configure [Pi](https://pi.dev/docs/latest/providers) to call the Modellix LLM gateway. Pi does not ship a built-in Modellix provider, so add a [custom provider](https://pi.dev/docs/latest/models) in `~/.pi/agent/models.json` that points at Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`). For Pi media packages and skills (image, video, speech), see [Plugin](/ways-to-use/plugin) and [Skill](/ways-to-use/skill).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Put that full string in each model `id` under your custom provider. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up Pi
Create a Modellix API key in the [console](https://modellix.ai/console/api-key) and export it so `models.json` can resolve `$MODELLIX_API_KEY`:
```bash theme={null}
export MODELLIX_API_KEY="mdlx-xxxxxxxx"
```
Add the same line to your shell profile (for example `~/.zshrc`) for persistent use.
Create or update `~/.pi/agent/models.json`:
```json theme={null}
{
"providers": {
"modellix": {
"baseUrl": "https://llm.modellix.ai/v1",
"api": "openai-completions",
"apiKey": "$MODELLIX_API_KEY",
"models": [
{ "id": "openai/gpt-5.5", "name": "GPT 5.5", "contextWindow": 200000 },
{ "id": "openai/gpt-5.6-sol", "name": "GPT 5.6 Sol", "contextWindow": 200000 },
{ "id": "openai/gpt-5.6-terra", "name": "GPT 5.6 Terra", "contextWindow": 200000 },
{ "id": "openai/gpt-5.6-luna", "name": "GPT 5.6 Luna", "contextWindow": 200000 },
{ "id": "anthropic/claude-sonnet-5", "name": "Claude Sonnet 5", "contextWindow": 200000 },
{ "id": "google/gemini-3.6-flash", "name": "Gemini 3.6 Flash", "contextWindow": 200000 },
{ "id": "xai/grok-4.5", "name": "Grok 4.5", "contextWindow": 200000 }
]
}
}
}
```
Merge the `modellix` entry into an existing `providers` object if you already use custom models. Add or remove entries under `models` as needed.
| Field | Value |
| ------------- | ------------------------------------------------------------------------------- |
| Provider key | `modellix` (custom; do not use built-in `openai`, `anthropic`, or `openrouter`) |
| `baseUrl` | `https://llm.modellix.ai/v1` |
| `api` | `openai-completions` |
| `apiKey` | `$MODELLIX_API_KEY` (or `${MODELLIX_API_KEY}`) |
| `models[].id` | Exact Modellix Model IDs (`provider/name`) |
Use a custom provider key such as `modellix`. Overwriting Pi's built-in `openai` or `anthropic` providers sends traffic to the vendor endpoints instead of Modellix.
You do **not** need a separate Anthropic base URL. Register `anthropic/...` and `google/...` IDs on the same Modellix provider; traffic goes through OpenAI-compatible Chat Completions on `https://llm.modellix.ai/v1`.
In an interactive Pi session, run `/model` and choose a Modellix entry (for example `openai/gpt-5.5` under the `modellix` provider).
Or start Pi with the provider and model on the command line:
```bash theme={null}
pi --provider modellix --model openai/gpt-5.5
```
Pi reloads `models.json` each time you open `/model`, so you can edit the file during a session without restarting.
If a Modellix model appears in the file but stays unavailable in `/model`, confirm `$MODELLIX_API_KEY` is set. Pi requires auth before custom models become selectable.
Start Pi as usual:
```bash theme={null}
pi
```
Or with Modellix selected up front:
```bash theme={null}
pi --provider modellix --model openai/gpt-5.5
```
## Troubleshooting
| Symptom | Check |
| ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| 401 / auth errors | `MODELLIX_API_KEY` is set; `apiKey` uses `$MODELLIX_API_KEY` or `${MODELLIX_API_KEY}` |
| Model listed but unavailable in `/model` | Auth is configured (env var, `apiKey` in `models.json`, or `--api-key`) |
| Model not found / 404 | `models[].id` is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5`; ID is in the live catalog |
| Traffic hits OpenAI or OpenRouter | Provider key is `modellix`; you did not select a built-in `openai` / `openrouter` model |
| Wrong path / URL errors | `baseUrl` is `https://llm.modellix.ai/v1` (include `/v1`), not the media API host |
| Developer role / 400 errors | Set provider `compat.supportsDeveloperRole` to `false` (see [Pi Custom Models](https://pi.dev/docs/latest/models)) |
## Related
* [Pi Custom Models](https://pi.dev/docs/latest/models) — `models.json` schema and examples
* [Pi Providers](https://pi.dev/docs/latest/providers) — auth and custom providers overview
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [OpenClaw](/llm/agent/openclaw) · [OpenCode](/llm/agent/opencode) — other OpenAI-compatible custom setups
* [Plugin](/ways-to-use/plugin) · [Skill](/ways-to-use/skill) — Pi media capabilities
# Use Modellix LLM with Qwen Code
Source: https://docs.modellix.ai/llm/agent/qwen-code
Point Qwen Code at the Modellix LLM gateway with OPENAI_BASE_URL, OPENAI_API_KEY, and provider/name model IDs.
Configure [Qwen Code](https://github.com/QwenLM/qwen-code) to call the Modellix LLM gateway. Qwen Code supports OpenAI-compatible endpoints through environment variables or [`modelProviders`](https://qwenlm.github.io/qwen-code-docs/en/users/configuration/model-providers/) in `settings.json`, pointed at Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Set that full string as `OPENAI_MODEL` or as each `modelProviders` model `id`. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up Qwen Code
Requires Node.js 20+. Install the global CLI:
```bash theme={null}
npm install -g @qwen-code/qwen-code
qwen --version
```
Point the OpenAI-compatible client at Modellix (same pattern as the [OpenAI SDK](/llm/sdk/openai-sdk)):
```bash theme={null}
export OPENAI_API_KEY="mdlx-xxxxxxxx"
export OPENAI_BASE_URL="https://llm.modellix.ai/v1"
export OPENAI_MODEL="openai/gpt-5.5"
```
Add these lines to your shell profile (for example `~/.zshrc`), then run `source ~/.zshrc`.
| Setting | Value |
| ----------------- | ------------------------------------------------------------------------ |
| `OPENAI_API_KEY` | Modellix API Key from the [console](https://modellix.ai/console/api-key) |
| `OPENAI_BASE_URL` | `https://llm.modellix.ai/v1` (include `/v1`) |
| `OPENAI_MODEL` | Exact Modellix Model ID (`provider/name`) |
Use a Modellix API key, not an OpenAI platform key. Do not set `OPENAI_BASE_URL` to the media API host (`https://api.modellix.ai`).
For multiple models or a durable project/user config, declare Modellix under the `openai` auth type in `~/.qwen/settings.json` (or `.qwen/settings.json` in a project):
```json theme={null}
{
"env": {
"MODELLIX_API_KEY": "mdlx-xxxxxxxx"
},
"modelProviders": {
"openai": {
"protocol": "openai",
"models": [
{
"id": "openai/gpt-5.5",
"name": "GPT 5.5",
"baseUrl": "https://llm.modellix.ai/v1",
"envKey": "MODELLIX_API_KEY"
},
{
"id": "anthropic/claude-sonnet-5",
"name": "Claude Sonnet 5",
"baseUrl": "https://llm.modellix.ai/v1",
"envKey": "MODELLIX_API_KEY"
},
{
"id": "google/gemini-3.6-flash",
"name": "Gemini 3.6 Flash",
"baseUrl": "https://llm.modellix.ai/v1",
"envKey": "MODELLIX_API_KEY"
}
]
}
},
"security": {
"auth": {
"selectedType": "openai"
}
},
"model": {
"name": "openai/gpt-5.5"
}
}
```
Replace `mdlx-xxxxxxxx` with your Modellix API key, or keep the key only in the process environment and omit the `env` block. Credentials in `modelProviders` are read from `process.env[envKey]`, not stored as plaintext API fields.
You do **not** need a separate Anthropic `baseUrl` without `/v1` for these IDs. Register `anthropic/...` and `google/...` on the same Modellix OpenAI-compatible `baseUrl`; traffic goes through Chat Completions on `https://llm.modellix.ai/v1`. Native Anthropic Messages is for the [Anthropic SDK](/llm/sdk/anthropic-sdk) and [Claude Code](/llm/agent/claude-code).
From your project directory:
```bash theme={null}
qwen
```
Run `/about` to confirm the active model and endpoint. Use `/model` to switch among entries declared in `modelProviders`.
## Troubleshooting
| Symptom | Check |
| ------------------------- | ----------------------------------------------------------------------------------------------- |
| 401 / auth errors | `OPENAI_API_KEY` or `MODELLIX_API_KEY` is a valid Modellix key |
| Model not found / 404 | Model ID is a full string like `openai/gpt-5.5`, not bare `gpt-5.5` |
| Wrong path / URL errors | `OPENAI_BASE_URL` / `baseUrl` is `https://llm.modellix.ai/v1` (include `/v1`) |
| Still hitting OpenAI | Base URL is Modellix; `/about` shows the expected endpoint |
| Model missing in `/model` | Entry exists under `modelProviders.openai.models` with matching `envKey` set in the environment |
## Related
* [Qwen Code model providers](https://qwenlm.github.io/qwen-code-docs/en/users/configuration/model-providers/) — `modelProviders` schema
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `OPENAI_BASE_URL` + `/v1` pattern
* [OpenCode](/llm/agent/opencode) — another OpenAI-compatible coding agent setup
# Use Modellix LLM with Reasonix
Source: https://docs.modellix.ai/llm/agent/reasonix
Point Reasonix at the Modellix LLM gateway with a custom OpenAI-compatible [[providers]] entry, MODELLIX_API_KEY in ~/.reasonix/.env, and provider/name model IDs.
Configure [Reasonix](https://reasonix.io/docs/) to call the Modellix LLM gateway. Reasonix is a local coding agent (CLI/TUI, desktop app, browser UI, or ACP). It has no built-in Modellix provider, so add an OpenAI-compatible `[[providers]]` entry in `~/.reasonix/config.toml` that points at Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Put that full string in `model` or `models` so Reasonix sends it unchanged in the request body. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up Reasonix
Install the 1.x CLI (npm is the documented default):
```bash npm theme={null}
npm i -g reasonix
```
```bash Homebrew theme={null}
brew install esengine/reasonix/reasonix
```
Desktop downloads are on the [Reasonix docs](https://reasonix.io/docs/). CLI and desktop share the same config and credential files.
Create a key in the [Modellix console](https://modellix.ai/console/api-key). Reasonix reads provider secrets from its **global** `.env`, not from your shell profile or a project `.env`.
macOS / Linux: `~/.reasonix/.env`\
Windows: `%APPDATA%\reasonix\.env`
```bash theme={null}
MODELLIX_API_KEY=mdlx-xxxxxxxx
```
Replace `mdlx-xxxxxxxx` with your Modellix API key. You can also paste the key in `reasonix setup` or desktop **Settings → Model → Access**; those flows write the same file.
Use a Modellix API key, not a DeepSeek, OpenAI, or other vendor platform key. Shell `export MODELLIX_API_KEY=...` is not a Reasonix provider-key fallback.
Edit the **user** config at `~/.reasonix/config.toml` (Windows: `%APPDATA%\reasonix\config.toml`). Merge the `[[providers]]` table into an existing file; do not delete other providers.
```toml theme={null}
default_model = "modellix"
[[providers]]
name = "modellix"
kind = "openai"
base_url = "https://llm.modellix.ai/v1"
models = [
"openai/gpt-5.5",
"openai/gpt-5.6-sol",
"anthropic/claude-sonnet-5",
"google/gemini-3.6-flash",
"xai/grok-4.6",
]
default = "openai/gpt-5.5"
api_key_env = "MODELLIX_API_KEY"
context_window = 200000
```
Add or remove IDs in `models` as needed. Full catalog: [Models & Pricing](/llm/overview#models-and-pricing).
| Field | Value |
| ---------------- | ------------------------------------------------------------------------------ |
| `name` | Local provider ID (use `modellix`; do not reuse `openai` or `anthropic`) |
| `kind` | `openai` for Chat Completions |
| `base_url` | `https://llm.modellix.ai/v1` (Reasonix posts to `{base_url}/chat/completions`) |
| `models` | Exact Modellix Model IDs (`provider/name`) |
| `default` | Default Modellix ID for this provider |
| `api_key_env` | `MODELLIX_API_KEY` (must match a key in `~/.reasonix/.env`) |
| `context_window` | Token window used for compaction; raise it if the catalog model is larger |
| `default_model` | Reasonix picker ref: provider name (`modellix`) or `modellix/` |
Config resolution is flags → `./reasonix.toml` → user `config.toml`. Keep personal API providers in the user file.
You do **not** need a separate Anthropic `kind` or base URL. Register `anthropic/...` and `google/...` IDs on the same Modellix `kind = "openai"` entry; traffic goes through OpenAI-compatible Chat Completions. Native Anthropic Messages is for the [Anthropic SDK](/llm/sdk/anthropic-sdk) and [Claude Code](/llm/agent/claude-code).
In Reasonix 1.24+, the desktop custom-provider form stores an exact **API address** as `request_url` (no path rewriting). If you use that form, enter `https://llm.modellix.ai/v1/chat/completions`. Manual `base_url = "https://llm.modellix.ai/v1"` remains valid for TOML edits.
From your project directory:
```bash theme={null}
cd your-project
reasonix
```
Switch models with `/model`. Prefer the explicit Reasonix ref `modellix/openai/gpt-5.5` (provider name, then the Modellix ID). Headless:
```bash theme={null}
reasonix run --model "modellix/openai/gpt-5.5" "Explain this repo"
```
`reasonix setup` can also add a custom OpenAI-compatible provider and write the key; afterwards, confirm `base_url` / `request_url` and `models` match the table above.
## Troubleshooting
| Symptom | Check |
| --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| 401 / missing API key | `MODELLIX_API_KEY` is in `~/.reasonix/.env` (or `%APPDATA%\reasonix\.env`), and `api_key_env` matches that name |
| Model not found / unknown model | `models` lists the full ID (`openai/gpt-5.5`); select `modellix/openai/gpt-5.5` or the provider name `modellix` |
| Wrong path / URL errors | `base_url` is `https://llm.modellix.ai/v1`, or `request_url` is `https://llm.modellix.ai/v1/chat/completions` — not the media API host (`https://api.modellix.ai`) |
| Traffic hits DeepSeek or another vendor | You selected the `modellix` provider, not a built-in DeepSeek / OpenAI / Anthropic preset |
| Early compaction | Raise `context_window` for the catalog model; custom providers no longer inherit DeepSeek window defaults |
| Provider name clash | Do not name the Reasonix provider `openai` or `anthropic` — Modellix IDs already start with those prefixes |
## Related
* [Reasonix docs](https://reasonix.io/docs/) — install, config paths, TUI and desktop
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` Chat Completions gateway
* [OpenCode](/llm/agent/opencode) · [Grok Build](/llm/agent/grok-build) — other OpenAI-compatible agent setups
# Use Modellix LLM with WorkBuddy
Source: https://docs.modellix.ai/llm/agent/workbuddy
Add Modellix as a WorkBuddy custom OpenAI-compatible model with Endpoint, models.json, and provider/name model IDs.
Configure [WorkBuddy](https://www.codebuddy.cn/work/) to call the Modellix LLM gateway. WorkBuddy supports OpenAI-compatible custom models in the UI and in `~/.workbuddy/models.json`, pointed at Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Enter that full string as the Model Name / `id`. Set Endpoint / `url` to `https://llm.modellix.ai/v1` (include `/v1`). With **Custom Protocol** off (default), WorkBuddy uses the standard `/chat/completions` path. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up WorkBuddy
Create a Modellix API key in the [console](https://modellix.ai/console/api-key) and store it securely. You will paste it into the WorkBuddy custom model form or `models.json`.
In WorkBuddy, open the model picker (for example **Auto** near the input box), scroll to the bottom, and choose **Configure custom models** → **Add Model**.
You can also open **Settings → Model** (or **Settings → Models**) and add a custom model there.
Set **Provider** to **Custom**, then fill in:
| Field | Value |
| ----------------- | ---------------------------------------------------------------------------------------------------------- |
| Endpoint | `https://llm.modellix.ai/v1` |
| API Key | Your Modellix API Key |
| Model Name | Exact Modellix Model ID (for example `openai/gpt-5.5`) |
| Advanced Settings | Enable **Tool Calling** for agent tool use; enable Image Input / Reasoning only if the model supports them |
| Custom Protocol | Leave **off** (default) so WorkBuddy uses `/chat/completions` |
Save the model. Include `/v1` in the Endpoint.
If your gateway requires a non-standard path, enable **Custom Protocol** and set Endpoint to the exact URL WorkBuddy should call (for example the full `https://llm.modellix.ai/v1/chat/completions`). For Modellix, leaving Custom Protocol off with Endpoint `https://llm.modellix.ai/v1` is the usual setup.
WorkBuddy stores custom models in a local JSON **array** (not the CodeBuddy `{ "models": [...] }` object shape):
* macOS / Linux: `~/.workbuddy/models.json`
* Windows: `%USERPROFILE%\.workbuddy\models.json`
Example:
```json theme={null}
[
{
"id": "openai/gpt-5.5",
"name": "GPT 5.5",
"vendor": "Custom",
"url": "https://llm.modellix.ai/v1",
"apiKey": "mdlx-xxxxxxxx",
"supportsToolCall": true,
"supportsImages": false,
"supportsReasoning": false,
"useCustomProtocol": false
}
]
```
Replace `mdlx-xxxxxxxx` with your Modellix API key. Append additional objects to the array for more Modellix IDs (for example `anthropic/claude-sonnet-5`, `google/gemini-3.6-flash`).
| Field | Value |
| ------------------- | ---------------------------------------------- |
| `id` / `name` | Exact Modellix Model ID (`provider/name`) |
| `vendor` | `Custom` |
| `url` | `https://llm.modellix.ai/v1` |
| `apiKey` | Modellix API Key |
| `useCustomProtocol` | `false` (default OpenAI Chat Completions path) |
| `supportsToolCall` | `true` recommended for agent tool use |
Do not use a built-in vendor that points at OpenAI or Anthropic directly. Choose **Custom** / `vendor: "Custom"` so traffic goes to Modellix. Do not set `url` to the media API host (`https://api.modellix.ai`).
After saving, restart WorkBuddy or reopen the model picker so the list refreshes. Full catalog: [Models & Pricing](/llm/overview#models-and-pricing).
Open the model picker, choose your model under **Custom Models**, and send a short test message to confirm the gateway responds.
## Troubleshooting
| Symptom | Check |
| -------------------------- | -------------------------------------------------------------------------------------------- |
| 401 / auth errors | Key is a valid Modellix API Key |
| Model not found / 404 | Model Name / `id` is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5` |
| Endpoint / URL errors | Endpoint or `url` includes `/v1`: `https://llm.modellix.ai/v1` |
| Model missing after save | Reopen the model picker and check **Custom Models**; confirm `models.json` is valid JSON |
| Wrong host | Not using the media API (`https://api.modellix.ai`) |
| Non-standard path failures | Enable **Custom Protocol** and use the full Chat Completions URL if auto-completion is wrong |
## Related
* [WorkBuddy](https://www.codebuddy.cn/work/) — product home
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [CodeBuddy](/llm/agent/codebuddy) — related Tencent custom-model setup (different `models.json` shape)
# Modellix LLM API Guide
Source: https://docs.modellix.ai/llm/api/api
Call Modellix LLM with OpenAI-compatible Chat Completions and Responses or Anthropic-compatible Messages—sync requests, streaming SSE, multimodal content parts, auth, model list, request logs, errors, and billing.
Modellix LLM is a text gateway at `https://llm.modellix.ai` with three protocol surfaces:
| Protocol | Method | Path | Typical clients |
| ---------------- | ------ | ---------------------- | -------------------------------------------- |
| Chat Completions | `POST` | `/v1/chat/completions` | OpenAI SDK, Codex, Cursor, most chat clients |
| Responses | `POST` | `/v1/responses` | OpenAI Responses API clients |
| Messages | `POST` | `/v1/messages` | Anthropic SDK, Claude Code |
Also available on the same host: [`GET /v1/models`](#list-models) and [`GET /v1/logs`](#request-logs).
Calls are **synchronous** (optional streaming SSE). Request and response shapes follow the corresponding official protocols; this guide lists core fields only. Full OpenAPI specs: [Chat Completions](/llm/chat-completions), [Responses](/llm/responses), [Messages](/llm/messages), [List models](/llm/list-models), [Request logs](/llm/get-llm-logs).
Media generation (image, video, speech) uses `https://api.modellix.ai` and async tasks. Do not mix that host with the LLM gateway.
## Base URL and Auth
| Item | Value |
| ----------- | --------------------------------------------------------------------------------- |
| Host | `https://llm.modellix.ai` |
| Path prefix | `/v1` |
| Auth | `Authorization: Bearer ` **or** `x-api-key: ` |
```http theme={null}
Authorization: Bearer mdlx-xxxxxxxx
```
```http theme={null}
x-api-key: mdlx-xxxxxxxx
```
Both headers are equivalent. If both are sent, they must be the same key. Get a key from the [console](https://modellix.ai/console/api-key)—not a vendor platform key.
| Header style | Typical clients |
| ------------ | ----------------------------------- |
| Bearer | OpenAI SDK, Codex, OpenCode, Cursor |
| `x-api-key` | Anthropic SDK, Claude Code |
## Models
Pass `model` in the JSON body as `provider/name`. See [Models & Pricing](/llm/overview#models-and-pricing) for the full Model ID list and rates. Availability follows the console and product releases.
You can also list currently available Model IDs via [`GET /v1/models`](#list-models).
## Choose a Protocol
The three endpoints use **different URLs and body shapes**. Do not mix fields across protocols.
| | Chat Completions | Responses | Messages |
| ------ | ------------------------------------------------- | --------------------------------- | ---------------------------------------- |
| Input | `messages: [{role, content}, ...]` | `input` (string or content array) | Anthropic `messages` + optional `system` |
| Length | `max_tokens` / `max_completion_tokens` | `max_output_tokens` | `max_tokens` (**required**) |
| Stream | `stream: true` → OpenAI chat.completion.chunk SSE | `stream: true` → Responses SSE | `stream: true` → Anthropic Messages SSE |
**Routing tips**
| Model prefix | Recommended protocol |
| --------------------- | ------------------------------------------------------------------------ |
| `openai/...` | Chat Completions or Responses |
| `anthropic/...` | Messages (Anthropic clients); Chat Completions also works for many tools |
| `google/...` (Gemini) | Chat Completions or Responses—not Messages |
## Reasoning
Optional. Field names follow the protocol you call; support and allowed values depend on the model. Do not mix these fields across protocols.
| Protocol | Field |
| ---------------- | --------------------------------------- |
| Chat Completions | `reasoning_effort` |
| Responses | `reasoning` (`effort`, optional `mode`) |
| Messages | `thinking`, `output_config.effort` |
## Chat Completions
```http theme={null}
POST /v1/chat/completions
Authorization: Bearer
Content-Type: application/json
```
| Field | Required | Description |
| ----------------------- | -------- | ----------------------------------------------- |
| `model` | Yes | Model ID |
| `messages` | Yes | OpenAI-style messages |
| `stream` | No | Default `false`; `true` returns SSE |
| `max_tokens` | No | Generation cap (model-dependent) |
| `max_completion_tokens` | No | Preferred cap on some newer OpenAI-style models |
| `temperature` | No | Sampling temperature |
| `n` | No | Number of choices |
```bash theme={null}
curl -sS "https://llm.modellix.ai/v1/chat/completions" \
-H "Authorization: Bearer ${API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-5.6-sol",
"stream": false,
"max_tokens": 256,
"messages": [{"role": "user", "content": "Introduce yourself in one sentence"}]
}'
```
Streaming:
```bash theme={null}
curl -sS -N "https://llm.modellix.ai/v1/chat/completions" \
-H "Authorization: Bearer ${API_KEY}" \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-d '{
"model": "openai/gpt-5.6-sol",
"stream": true,
"max_tokens": 256,
"messages": [{"role": "user", "content": "ping"}]
}'
```
See [Create chat completion](/llm/chat-completions). For image and video `content` parts, see [Multimodal Inputs](#multimodal-inputs).
## Responses
```http theme={null}
POST /v1/responses
Authorization: Bearer
Content-Type: application/json
```
| Field | Required | Description |
| ------------------- | -------- | --------------------------------------------------------- |
| `model` | Yes | Model ID |
| `input` | Yes | String or content array (not Chat Completions `messages`) |
| `stream` | No | Default `false` |
| `max_output_tokens` | No | Output cap |
| `temperature` | No | Sampling temperature |
```bash theme={null}
curl -sS "https://llm.modellix.ai/v1/responses" \
-H "Authorization: Bearer ${API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-5.6-luna",
"stream": false,
"max_output_tokens": 256,
"input": "Introduce yourself in one sentence"
}'
```
See [Create response](/llm/responses). For `input_image` / `input_file`, see [Multimodal Inputs](#multimodal-inputs).
## Messages (Anthropic)
```http theme={null}
POST /v1/messages
Authorization: Bearer
Content-Type: application/json
```
| Field | Required | Description |
| ------------ | -------- | ------------------------------------------------------------- |
| `model` | Yes | Use `anthropic/...` (for example `anthropic/claude-sonnet-5`) |
| `messages` | Yes | Anthropic messages (`user` / `assistant` only) |
| `max_tokens` | Yes | Output cap |
| `stream` | No | Default `false` |
| `system` | No | System prompt (string or content blocks) |
Auth may use Bearer or `x-api-key`. Optional `anthropic-version` header is forwarded when present.
```bash theme={null}
curl -sS "https://llm.modellix.ai/v1/messages" \
-H "Authorization: Bearer ${API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-5",
"max_tokens": 256,
"messages": [{"role": "user", "content": "ping"}]
}'
```
See [Create message](/llm/messages). For `image` / `document` blocks, see [Multimodal Inputs](#multimodal-inputs).
## Multimodal Inputs
Field names follow the protocol you call. Prefer a public HTTPS URL, including the `url` returned by the Media [File API](/ways-to-use/api#upload-media-files). Local filesystem paths are ignored.
Image, video, and audio support is per **model**—see [Models & Pricing](/llm/overview#models-and-pricing). Examples use `google/gemini-3.7-flash`. Check `usage` to confirm the media was consumed.
| Client | Endpoint | Image | File / PDF | Video | Audio |
| --------------------------------- | ------------------------------------------- | -------------------------------- | --------------------- | ------------------------- | --------------------------- |
| OpenAI SDK, Codex, most chat apps | [Chat Completions](#chat-completions-parts) | `image_url.url` | `file.file_id` | `file.file_id` + `format` | `input_audio` (inline only) |
| OpenAI Responses clients | [Responses](#responses-parts) | `input_image.image_url` (string) | `input_file.file_url` | `input_file.file_url` | — |
| Anthropic SDK, Claude Code | [Messages](#messages-parts) | `image.source.url` | `document.source.url` | — | — |
Do not mix fields across protocols. OpenAI or Anthropic Files IDs (`file-...`) are not accepted; for Media File API uploads, pass the returned **`url`**, not the UUID. Uploads default to **16 MB**.
Official schemas: [Chat Completions](https://developers.openai.com/api/reference/resources/chat/subresources/completions/methods/create/), [Responses](https://developers.openai.com/api/reference/resources/responses/methods/create/), [Messages](https://docs.anthropic.com/en/api/messages), [vision](https://developers.openai.com/api/docs/guides/images-vision), [file inputs](https://developers.openai.com/api/docs/guides/file-inputs), [Anthropic vision](https://docs.anthropic.com/en/docs/build-with-claude/vision).
### Chat Completions
Recommended: HTTPS in `image_url.url` or `file.file_id`. There is no `video_url` type.
```bash theme={null}
curl -sS "https://llm.modellix.ai/v1/chat/completions" \
-H "Authorization: Bearer ${API_KEY}" \
-H "Content-Type: application/json" \
-d '{
"model": "google/gemini-3.7-flash",
"stream": false,
"max_tokens": 256,
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What appears in this video?"},
{
"type": "file",
"file": {
"file_id": "https://example.com/clip.mp4",
"format": "video/mp4"
}
}
]
}]
}'
```
Image: `{ "type": "image_url", "image_url": { "url": "https://example.com/photo.jpg" } }`. Audio has no URL field: `{ "type": "input_audio", "input_audio": { "data": "", "format": "wav" } }` (`mp3` is also valid). If the model returns `usage.prompt_tokens_details.video_tokens` or `audio_tokens`, those counts are already inside `prompt_tokens`—see [Billing](#billing).
### Responses
Recommended: HTTPS string in `input_image.image_url` or `input_file.file_url` (not the Chat Completions nested `{ "url": "..." }` object).
```json theme={null}
{
"model": "google/gemini-3.7-flash",
"input": [
{
"role": "user",
"content": [
{ "type": "input_text", "text": "What appears in this video?" },
{ "type": "input_file", "file_url": "https://example.com/clip.mp4" }
]
}
]
}
```
Image: `{ "type": "input_image", "image_url": "https://example.com/photo.jpg" }`.
### Messages
Recommended: HTTPS in `source.url`. This protocol has no video or audio block. For those inputs, use Chat Completions or Responses.
```json theme={null}
{
"model": "anthropic/claude-sonnet-5",
"max_tokens": 256,
"messages": [{
"role": "user",
"content": [
{
"type": "image",
"source": { "type": "url", "url": "https://example.com/photo.jpg" }
},
{ "type": "text", "text": "Describe this image." }
]
}]
}
```
PDF: `{ "type": "document", "source": { "type": "url", "url": "https://example.com/doc.pdf" } }`.
### Inline media
Use the same `type` as above. Chat Completions and Responses images: `data:;base64,...` in the URL field. Chat Completions `file`: the same form in `file.file_id`. Responses files: `file_data`. Messages: `source.type` `base64` with `media_type` and `data`. Put the data URL on one line; Base64 increases JSON size by about 4/3.
## Session Header
For multi-turn session affinity, send:
| Header | Rules |
| ------------------- | --------------------------------------------- |
| `X-Mdlx-Session-Id` | Length 8–128; alphanumeric, `-`, and `_` only |
Some native tools send their own session header (for example Claude Code’s `X-Claude-Code-Session-Id`). If both are present, `X-Mdlx-Session-Id` takes precedence.
```bash theme={null}
curl -sS "https://llm.modellix.ai/v1/chat/completions" \
-H "Authorization: Bearer ${API_KEY}" \
-H "Content-Type: application/json" \
-H "X-Mdlx-Session-Id: my-conversation-001" \
-d '{
"model": "openai/gpt-5.6-sol",
"messages": [{"role": "user", "content": "Continue our previous topic"}]
}'
```
## End-User ID Header
Optional. Tag requests with your own end-user identifier so you can filter [request logs](#request-logs) later. Invalid values return `400`.
| Header | Rules |
| ---------------- | ---------------------------------------------------------------- |
| `X-Mdlx-User-Id` | Optional; length 8–128; ASCII letters, digits, `-`, and `_` only |
```bash theme={null}
curl -sS "https://llm.modellix.ai/v1/chat/completions" \
-H "Authorization: Bearer ${API_KEY}" \
-H "Content-Type: application/json" \
-H "X-Mdlx-Session-Id: my-conversation-001" \
-H "X-Mdlx-User-Id: end_user_01" \
-d '{
"model": "openai/gpt-5.6-sol",
"messages": [{"role": "user", "content": "Continue our previous topic"}]
}'
```
## List Models
```http theme={null}
GET /v1/models
Authorization: Bearer
```
Returns an OpenAI-compatible model list for the gateway (`object: list`, `data[].id` in `provider/name` form). IDs are de-duplicated. This endpoint uses the **query** rate limit (shared with request-log listing and similar read APIs)—not the inference RPM quota.
```bash theme={null}
curl -sS "https://llm.modellix.ai/v1/models" \
-H "Authorization: Bearer ${API_KEY}"
```
Example response:
```json theme={null}
{
"object": "list",
"data": [
{ "id": "openai/gpt-5.5", "object": "model" },
{ "id": "openai/gpt-5.6-sol", "object": "model" }
]
}
```
## Request Logs
```http theme={null}
GET /v1/logs
Authorization: Bearer
```
Lists **your team’s** LLM request logs for a time window (same API Key / team scope as inference). Uses the **query** rate limit.
| Query | Required | Description |
| -------------- | -------- | --------------------------------------------------------------- |
| `start_time` | Yes | UNIX seconds |
| `end_time` | Yes | UNIX seconds; must be greater than `start_time`; span ≤ 30 days |
| `mdlx_user_id` | No | Exact match; same rules as `X-Mdlx-User-Id` |
| `page` | No | Default `1` |
| `page_size` | No | Default `10`, max `100` |
Response shape (fields useful for debugging and auditing):
| Field | Description |
| ------------------------------------------------------------------ | --------------------------------------- |
| `requests` | Log items |
| `total` / `page` / `page_size` | Pagination |
| `requests[].task_id` | Request id |
| `requests[].status` / `error` | Outcome |
| `requests[].model` | `{ provider, model_name }` |
| `requests[].prompt_tokens` / `completion_tokens` / `cached_tokens` | Token usage when available |
| `requests[].cost` | Cost in the console’s billing unit |
| `requests[].created_at` | UNIX seconds |
| `requests[].input` / `result` | Request/response payloads when retained |
The list response does **not** include a `mdlx_user_id` field; filter with the query parameter instead.
```bash theme={null}
curl -sS "https://llm.modellix.ai/v1/logs?start_time=1700000000&end_time=1700086400&page=1&page_size=20" \
-H "Authorization: Bearer ${API_KEY}"
```
For **media** (image/video/speech) request logs on `https://api.modellix.ai`, see [List media request logs](/api/get-logs).
## Success Responses
* **Non-streaming:** HTTP `200` with a JSON body matching the protocol (chat.completion, response, or message).
* **Streaming:** HTTP `200`, `Content-Type: text/event-stream`, protocol-specific SSE events.
Responses usually include `usage` (often on the final stream event). Billing uses token usage; see [Billing](#billing).
## Errors
```json theme={null}
{
"error": {
"message": "...",
"type": "invalid_request_error"
}
}
```
Optional fields may include `code` and `param`.
| HTTP | Meaning | Common `error.type` |
| ----- | ----------------------------------------------- | ------------------------------------------------------- |
| `400` | Invalid parameters | `invalid_request_error` |
| `401` | Missing/invalid key or conflicting auth headers | `invalid_request_error` |
| `402` | Insufficient balance | `insufficient_quota` |
| `404` | Unknown path or model unavailable | `invalid_request_error` |
| `429` | Rate limit or model temporarily unavailable | `rate_limit_exceeded` / `request_limited` / `api_error` |
| `5xx` | Temporary upstream or service error | `api_error` |
## Billing
Successful responses are billed from token `usage` at the model’s input and output rates (plus cache read/write when present). `prompt_tokens_details.video_tokens` and `audio_tokens` are included in `prompt_tokens` and are not billed separately. Unit prices and the invoice follow the console. Insufficient balance returns `402` with `insufficient_quota`.
## Rate Limits
| Case | HTTP | Common `error.type` | Notes |
| ----------------------------- | ----- | --------------------- | ----------------------------------------------------------------------------------------------------------------- |
| Inference RPM exceeded | `429` | `rate_limit_exceeded` | Chat / Responses / Messages; may include `X-RateLimit-Limit` / `Remaining` / `Reset` |
| Query RPM exceeded | `429` | `rate_limit_exceeded` | `GET /v1/models`, `GET /v1/logs` (and similar listing APIs); separate from inference RPM; same rate-limit headers |
| Other request limits | `429` | `request_limited` | Slow down or retry later |
| Model temporarily unavailable | `429` | `api_error` | May include `Retry-After`; switch model or retry |
Back off and reduce request rate after `429`.
## Client Quick Reference
| Client | Base URL | Credential env | Model prefix |
| ------------------------------------------------------------------------------- | ------------------------------------ | --------------------------------------------- | ------------------ |
| [OpenAI SDK](/llm/sdk/openai-sdk) | `https://llm.modellix.ai/v1` | `OPENAI_API_KEY` | `openai/...` |
| [Codex](/llm/agent/codex) | `openai_base_url` = `.../v1` | `OPENAI_API_KEY` | `openai/...` |
| [Cursor](/llm/ide/cursor) | `https://llm.modellix.ai/v1` | Modellix API Key in settings | `provider/name` |
| [Anthropic SDK](/llm/sdk/anthropic-sdk) / [Claude Code](/llm/agent/claude-code) | `https://llm.modellix.ai` (no `/v1`) | `ANTHROPIC_API_KEY` or `ANTHROPIC_AUTH_TOKEN` | `anthropic/...` |
| [OpenCode](/llm/agent/opencode) | Provider `baseURL` | `OPENAI_API_KEY` / `ANTHROPIC_API_KEY` | Match the protocol |
Google models (`google/...`) on the OpenAI-compatible path use `OPENAI_*` and the `/v1` base URL.
## Differences from Vendor Docs
| Topic | Modellix |
| ----------------- | ------------------------------------------------------------------------------- |
| Auth | Bearer or `x-api-key` with a Modellix key |
| `model` | `provider/name` form — see [Models & Pricing](/llm/overview#models-and-pricing) |
| Chat vs Responses | Different bodies—changing only the URL is not enough |
| Multimodal | Fields follow the protocol you call—see [Multimodal Inputs](#multimodal-inputs) |
| Session | Optional `X-Mdlx-Session-Id` |
| End-user id | Optional `X-Mdlx-User-Id` for log filtering |
| Model list | `GET /v1/models` (OpenAI-compatible) |
| Request logs | `GET /v1/logs` (team-scoped; optional `mdlx_user_id`) |
Official field catalogs: [OpenAI Chat Completions](https://platform.openai.com/docs/api-reference/chat), [OpenAI Responses](https://platform.openai.com/docs/api-reference/responses), [Anthropic Messages](https://docs.anthropic.com/en/api/messages).
# Chat Completions
Source: https://docs.modellix.ai/llm/chat-completions
/llm/api/llm.json post /chat/completions
OpenAI Chat Completions–compatible endpoint on the Modellix LLM gateway. Accepts a synchronous request with `model` (provider/name format) and an OpenAI-style `messages` array; returns a chat.completion JSON object by default, or SSE (`text/event-stream`) when `stream=true`. Use this path for the OpenAI SDK, Codex, and most OpenAI-compatible chat clients. Do not send Responses fields (`input`, `max_output_tokens`) or Anthropic Messages-only shapes on this endpoint—use `/responses` or `/messages` instead. Optional fields follow [OpenAI Chat Completions](https://platform.openai.com/docs/api-reference/chat); support depends on the selected model.
# Use Modellix LLM with Agno
Source: https://docs.modellix.ai/llm/framework/agno
Point Agno agents at Modellix with OpenAILike base_url, a Modellix API key, and provider/name model IDs.
Use [Agno](https://docs.agno.com/models/overview) against the Modellix LLM gateway. Agno provides [`OpenAILike`](https://docs.agno.com/models/providers/openai-like) for any OpenAI-compatible Chat Completions endpoint — set `base_url` to `https://llm.modellix.ai/v1` and `id` to a Modellix `provider/name` Model ID.
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Pass that full string as `OpenAILike(id=...)`. Do not rely on a bare string like `model="openai:gpt-5.4"` — that uses Agno's built-in OpenAI provider defaults, not Modellix. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up Agno
```bash theme={null}
pip install "agno[openai]"
```
Or with [uv](https://docs.astral.sh/uv/):
```bash theme={null}
uv pip install "agno[openai]"
```
Create a Modellix API key in the [console](https://modellix.ai/console/api-key):
```bash theme={null}
export MODELLIX_API_KEY="mdlx-xxxxxxxx"
```
| Setting | Value |
| ------------------------------ | -------------------------------------------- |
| `api_key` / `MODELLIX_API_KEY` | Modellix API Key |
| `base_url` | `https://llm.modellix.ai/v1` (include `/v1`) |
| `id` | Exact Modellix Model ID (`provider/name`) |
Use a Modellix key, not an OpenAI platform key. Do not point Agno at the media host (`https://api.modellix.ai`).
```python theme={null}
from os import getenv
from agno.agent import Agent
from agno.models.openai.like import OpenAILike
agent = Agent(
model=OpenAILike(
id="openai/gpt-5.5",
api_key=getenv("MODELLIX_API_KEY"),
base_url="https://llm.modellix.ai/v1",
),
markdown=True,
)
agent.print_response("Introduce yourself in one sentence.", stream=True)
```
Change `id` to any catalog Model ID (`anthropic/claude-sonnet-5`, `google/gemini-3.6-flash`, and so on). Traffic still uses Chat Completions on `https://llm.modellix.ai/v1`.
```python theme={null}
model = OpenAILike(
id="openai/gpt-5.5",
api_key=getenv("MODELLIX_API_KEY"),
base_url="https://llm.modellix.ai/v1",
retries=2,
delay_between_retries=1,
exponential_backoff=True,
)
agent = Agent(model=model, markdown=True)
agent.print_response("Write a one-line haiku about APIs.", stream=True)
```
Model-level retries cover transient provider errors (rate limits, server errors). Auth and invalid-request failures are not retried.
Modellix also exposes [`POST /v1/responses`](/llm/responses). For OpenAI-compatible Responses / Open Responses style clients, Agno documents `OpenResponses` with a custom `base_url`:
```python theme={null}
from agno.agent import Agent
from agno.models.openai import OpenResponses
agent = Agent(
model=OpenResponses(
id="openai/gpt-5.5",
api_key=getenv("MODELLIX_API_KEY"),
base_url="https://llm.modellix.ai/v1",
),
)
agent.print_response("Say hello in one short sentence.")
```
Prefer `OpenAILike` (Chat Completions) first unless you specifically need the Responses path. Class names and imports can vary by Agno version — check the [OpenAI-like](https://docs.agno.com/models/providers/openai-like) and [Models overview](https://docs.agno.com/models/overview) pages for your release.
## Troubleshooting
| Symptom | Check |
| --------------------- | -------------------------------------------------------------------------------------- |
| Still hitting OpenAI | Use `OpenAILike(..., base_url="https://llm.modellix.ai/v1")`, not `model="openai:..."` |
| 401 Unauthorized | `api_key` / `MODELLIX_API_KEY` is a valid Modellix key |
| Model not found / 404 | `id` is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5` |
| Wrong path | `base_url` is `https://llm.modellix.ai/v1` (include `/v1`) |
| Import errors | Install `agno[openai]`; confirm `OpenAILike` import path for your Agno version |
## Related
* [Agno models overview](https://docs.agno.com/models/overview) — model configuration and retries
* [OpenAI-compatible models](https://docs.agno.com/models/providers/openai-like) — `OpenAILike` + custom `base_url`
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` gateway without Agno
* [CrewAI](/llm/framework/crewai) — another Python agent framework on Chat Completions
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
# Use Modellix LLM with CrewAI
Source: https://docs.modellix.ai/llm/framework/crewai
Point CrewAI at Modellix with LLM base_url, a Modellix API key, provider/name model IDs, and custom_openai for OpenAI-compatible routing.
Use [CrewAI](https://docs.crewai.com/en/learn/llm-connections) against the Modellix LLM gateway. Pass an OpenAI-compatible `base_url` of `https://llm.modellix.ai/v1` to the built-in `LLM` class — you do **not** need a custom `BaseLLM` subclass for Modellix.
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Set that full string as `model`. See [Models & Pricing](/llm/overview#models-and-pricing).
[Custom LLM (`BaseLLM`)](https://docs.crewai.com/en/learn/custom-llm) is for non-OpenAI protocols or special auth. Prefer the `LLM` + `base_url` path below for Modellix.
## Set Up CrewAI
```bash theme={null}
pip install "crewai[openai]"
```
Or with [uv](https://docs.astral.sh/uv/):
```bash theme={null}
uv add "crewai[openai]"
```
The `[openai]` extra covers the native OpenAI-compatible client path used with `custom_openai=True`.
Create a Modellix API key in the [console](https://modellix.ai/console/api-key):
```bash theme={null}
export OPENAI_API_KEY="mdlx-xxxxxxxx"
export OPENAI_BASE_URL="https://llm.modellix.ai/v1"
# Legacy alias also works:
# export OPENAI_API_BASE="https://llm.modellix.ai/v1"
```
| Setting | Value |
| ------------------------------ | -------------------------------------------- |
| `api_key` / `OPENAI_API_KEY` | Modellix API Key |
| `base_url` / `OPENAI_BASE_URL` | `https://llm.modellix.ai/v1` (include `/v1`) |
| `model` | Exact Modellix Model ID (`provider/name`) |
Use a Modellix key, not an OpenAI platform key. Do not point CrewAI at the media host (`https://api.modellix.ai`). Prefer setting `base_url` (or `custom_openai=True` with env vars) so traffic does not fall back to `api.openai.com`.
Recommended: construct `LLM` with an explicit `base_url` and `custom_openai=True` so CrewAI uses the native OpenAI SDK against your gateway:
```python theme={null}
from crewai import Agent, Crew, LLM, Task
llm = LLM(
model="openai/gpt-5.5",
api_key="mdlx-xxxxxxxx", # or rely on OPENAI_API_KEY
base_url="https://llm.modellix.ai/v1",
custom_openai=True,
temperature=0.7,
)
agent = Agent(
role="Coding Assistant",
goal="Help with concise coding tasks",
backstory="You are a careful software engineer.",
llm=llm,
)
task = Task(
description="Introduce yourself in one sentence.",
expected_output="A single short sentence",
agent=agent,
)
crew = Crew(agents=[agent], tasks=[task])
result = crew.kickoff()
print(result)
```
Change `model` to any catalog ID (`anthropic/claude-sonnet-5`, `google/gemini-3.6-flash`, and so on). Traffic still uses Chat Completions on `https://llm.modellix.ai/v1`.
On some CrewAI versions, a leading `openai/` routing prefix may be stripped before the request is sent. If `openai/gpt-5.5` fails with model-not-found, confirm the upstream `model` field still includes `openai/`. Prefer `custom_openai=True` with an explicit `base_url`, and try a non-`openai/` catalog ID (for example `anthropic/claude-sonnet-5`) to verify the gateway path first.
If `OPENAI_BASE_URL` / `OPENAI_API_BASE` is already set, you can omit `base_url` in code, but keep `custom_openai=True` for unknown / namespaced model IDs so CrewAI does not route to the default OpenAI host. For the native OpenAI SDK path, install `crewai[openai]`.
```bash theme={null}
export OPENAI_API_KEY="mdlx-xxxxxxxx"
export OPENAI_BASE_URL="https://llm.modellix.ai/v1"
export OPENAI_MODEL_NAME="openai/gpt-5.5"
```
```python theme={null}
from crewai import Agent, LLM
llm = LLM(model="openai/gpt-5.5", custom_openai=True)
agent = Agent(
role="Assistant",
goal="Help users",
backstory="A helpful assistant.",
llm=llm,
)
```
CrewAI resolves the base URL from `base_url`, then `api_base`, then `OPENAI_BASE_URL`, then legacy `OPENAI_API_BASE`.
Share one Modellix `LLM` across agents, or give each agent a different catalog ID on the same gateway:
```python theme={null}
primary = LLM(
model="openai/gpt-5.5",
base_url="https://llm.modellix.ai/v1",
api_key="mdlx-xxxxxxxx",
custom_openai=True,
)
fast = LLM(
model="google/gemini-3.6-flash",
base_url="https://llm.modellix.ai/v1",
api_key="mdlx-xxxxxxxx",
custom_openai=True,
)
researcher = Agent(
role="Researcher",
goal="Gather facts",
backstory="You research carefully.",
llm=primary,
)
writer = Agent(
role="Writer",
goal="Write concise summaries",
backstory="You write clearly.",
llm=fast,
)
```
## Troubleshooting
| Symptom | Check |
| --------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| Still hitting OpenAI | Set `base_url` and/or `custom_openai=True`; key is a Modellix key |
| 401 Unauthorized | `api_key` / `OPENAI_API_KEY` is valid |
| Model not found / 404 | `model` is a full ID like `openai/gpt-5.5`; if `openai/` was stripped, try another catalog prefix or inspect the outbound request |
| Wrong path | `base_url` is `https://llm.modellix.ai/v1` (include `/v1`) |
| Env vs constructor conflict | Prefer per-`LLM` `base_url` + `api_key` for multi-provider setups |
## Related
* [CrewAI — Connect to any LLM](https://docs.crewai.com/en/learn/llm-connections) — `LLM`, `base_url`, OpenAI-compatible setup
* [CrewAI — Custom LLM](https://docs.crewai.com/en/learn/custom-llm) — `BaseLLM` for non-standard protocols
* [LiteLLM removal / custom\_openai](https://docs.crewai.com/edge/en/learn/litellm-removal-guide) — native OpenAI-compatible routing
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` gateway without CrewAI
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
# Use Modellix LLM with LangChain
Source: https://docs.modellix.ai/llm/framework/langchain
Point LangChain ChatOpenAI at the Modellix LLM gateway with base_url, a Modellix API key, and provider/name model IDs.
Use [LangChain](https://www.langchain.com/) against the Modellix LLM gateway through the OpenAI-compatible Chat Completions API. Configure [`ChatOpenAI`](https://docs.langchain.com/oss/python/integrations/chat/openai) from `langchain-openai` with `base_url` set to `https://llm.modellix.ai/v1` — the same pattern as the [OpenAI SDK](/llm/sdk/openai-sdk).
This page covers **LLM chat** only (`https://llm.modellix.ai`). Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Pass that full string as `model`. See [Models & Pricing](/llm/overview#models-and-pricing).
The LLM gateway does **not** document an embeddings API. For RAG, use Modellix for generation and a separate embeddings provider, or keep embeddings on another OpenAI-compatible host.
## Set Up LangChain
Python (recommended imports from `langchain-openai`):
```bash theme={null}
pip install langchain langchain-openai
```
TypeScript:
```bash theme={null}
npm install @langchain/openai @langchain/core
```
Create a Modellix API key in the [console](https://modellix.ai/console/api-key), then export:
```bash theme={null}
export OPENAI_API_KEY="mdlx-xxxxxxxx"
export OPENAI_BASE_URL="https://llm.modellix.ai/v1"
```
| Setting | Value |
| ----------------- | --------------------------------------------- |
| `OPENAI_API_KEY` | Modellix API Key (not an OpenAI platform key) |
| `OPENAI_BASE_URL` | `https://llm.modellix.ai/v1` (include `/v1`) |
| Model | Exact Modellix Model ID (`provider/name`) |
Missing `/v1` on the base URL commonly causes 404s. Do not point LangChain at the media host (`https://api.modellix.ai`).
Prefer constructor args so the base URL is explicit in code. Environment variables work as a fallback when you omit `api_key` / `base_url`.
```python Python theme={null}
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="openai/gpt-5.5",
api_key="mdlx-xxxxxxxx", # or rely on OPENAI_API_KEY
base_url="https://llm.modellix.ai/v1",
max_tokens=256,
)
response = llm.invoke("Introduce yourself in one sentence")
print(response.content)
```
```typescript TypeScript theme={null}
import { ChatOpenAI } from "@langchain/openai";
const llm = new ChatOpenAI({
model: "openai/gpt-5.5",
apiKey: "mdlx-xxxxxxxx", // or rely on OPENAI_API_KEY
configuration: {
baseURL: "https://llm.modellix.ai/v1",
},
maxTokens: 256,
});
const response = await llm.invoke("Introduce yourself in one sentence");
console.log(response.content);
```
You can use the same `base_url` with other Modellix catalog IDs (for example `anthropic/claude-sonnet-5` or `google/gemini-3.6-flash`); traffic still goes through Chat Completions on `https://llm.modellix.ai/v1`.
Streaming:
```python theme={null}
for chunk in llm.stream("Write a one-line haiku about APIs"):
print(chunk.content, end="", flush=True)
```
Multi-model fallback on the same gateway:
```python theme={null}
primary = ChatOpenAI(
model="openai/gpt-5.5",
base_url="https://llm.modellix.ai/v1",
api_key="mdlx-xxxxxxxx",
)
fallback = ChatOpenAI(
model="google/gemini-3.6-flash",
base_url="https://llm.modellix.ai/v1",
api_key="mdlx-xxxxxxxx",
)
llm = primary.with_fallbacks([fallback])
```
For native Anthropic Messages instead of Chat Completions, use `langchain-anthropic` with the same host shape as the [Anthropic SDK](/llm/sdk/anthropic-sdk) (**no** `/v1` on the base URL):
```bash theme={null}
pip install langchain-anthropic
```
```python theme={null}
import os
from langchain_anthropic import ChatAnthropic
os.environ["ANTHROPIC_API_KEY"] = "mdlx-xxxxxxxx"
os.environ["ANTHROPIC_BASE_URL"] = "https://llm.modellix.ai"
llm = ChatAnthropic(
model="anthropic/claude-sonnet-5",
max_tokens=256,
)
```
Most LangChain apps should stay on `ChatOpenAI` + `/v1` unless you specifically need the Messages protocol.
## Troubleshooting
| Symptom | Check |
| -------------------------- | -------------------------------------------------------------------------------- |
| 404 Not Found | `base_url` / `OPENAI_BASE_URL` is `https://llm.modellix.ai/v1` (include `/v1`) |
| 401 Unauthorized | Key is a valid Modellix API key |
| Model not found | Model is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5` |
| Deprecated import warnings | Use `from langchain_openai import ChatOpenAI`, not `langchain.chat_models` |
| Embeddings / RAG errors | LLM gateway has no documented embeddings route — use another embeddings provider |
## Related
* [LangChain OpenAI-compatible providers](https://docs.langchain.com/oss/python/concepts/providers-and-models) — `ChatOpenAI` + `base_url`
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` gateway without LangChain
* [Anthropic SDK](/llm/sdk/anthropic-sdk) — Messages protocol used by `ChatAnthropic`
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [Chat Completions](/llm/chat-completions) — OpenAPI reference
# Use Modellix LLM with Mastra
Source: https://docs.modellix.ai/llm/framework/mastra
Point Mastra agents at Modellix with a custom OpenAI-compatible model url, apiKey, and gateway-style provider/name IDs.
Use [Mastra](https://mastra.ai/models) against the Modellix LLM gateway. Mastra is not a built-in Modellix provider — configure a [custom OpenAI-compatible endpoint](https://mastra.ai/models#use-local-models-with-mastra) with `url` set to `https://llm.modellix.ai/v1`, or pass a [Vercel AI SDK](/llm/sdk/vercel-ai-sdk) provider instance as `model`.
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix expects full `provider/name` Model IDs in the request body (for example `openai/gpt-5.5`). Treat Modellix as a **gateway**: use a three-part `id` (`modellix//`) so Mastra forwards `openai/gpt-5.5` upstream. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up Mastra
Follow the [Mastra getting started](https://mastra.ai/docs) flow for your app, or add the core package to an existing TypeScript project:
```bash theme={null}
npm install @mastra/core
```
For the AI SDK provider path below, also install:
```bash theme={null}
npm install ai @ai-sdk/openai-compatible
```
Create a Modellix API key in the [console](https://modellix.ai/console/api-key):
```bash theme={null}
export MODELLIX_API_KEY="mdlx-xxxxxxxx"
```
| Setting | Value |
| ----------------------------- | --------------------------------------------------------------------- |
| `apiKey` / `MODELLIX_API_KEY` | Modellix API Key |
| `url` / `baseURL` | `https://llm.modellix.ai/v1` (OpenAI-compatible base; include `/v1`) |
| Model ID | Gateway form `modellix//`, or AI SDK `openai/gpt-5.5` |
Do not set `url` to a full `/chat/completions` path — use the `/v1` base only. Do not point Mastra at the media host (`https://api.modellix.ai`). A plain string like `model: "openai/gpt-5.5"` routes to OpenAI's built-in provider, not Modellix.
Pass an object to `model` with `id`, `url`, and `apiKey`:
```typescript theme={null}
import { Agent } from "@mastra/core/agent";
const agent = new Agent({
id: "modellix-agent",
name: "Modellix Agent",
instructions: "You are a concise coding assistant.",
model: {
// Gateway form: Mastra sends openai/gpt-5.5 to Modellix
id: "modellix/openai/gpt-5.5",
url: "https://llm.modellix.ai/v1",
apiKey: process.env.MODELLIX_API_KEY,
},
});
const result = await agent.generate("Introduce yourself in one sentence.");
console.log(result.text);
```
Other catalog examples:
| Mastra `id` | Upstream Model ID |
| ------------------------------------ | --------------------------- |
| `modellix/openai/gpt-5.5` | `openai/gpt-5.5` |
| `modellix/anthropic/claude-sonnet-5` | `anthropic/claude-sonnet-5` |
| `modellix/google/gemini-3.6-flash` | `google/gemini-3.6-flash` |
Traffic still uses Chat Completions on `https://llm.modellix.ai/v1`.
Mastra accepts AI SDK language models anywhere a `"provider/model"` string works:
```typescript theme={null}
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { Agent } from "@mastra/core/agent";
const modellix = createOpenAICompatible({
name: "modellix",
apiKey: process.env.MODELLIX_API_KEY,
baseURL: "https://llm.modellix.ai/v1",
});
const agent = new Agent({
id: "modellix-agent",
name: "Modellix Agent",
instructions: "You are a concise coding assistant.",
model: modellix("openai/gpt-5.5"),
});
```
Same pattern as the [Vercel AI SDK](/llm/sdk/vercel-ai-sdk) guide. You can also put this model in Mastra [fallback chains](https://mastra.ai/models#model-fallbacks).
Keep several Modellix models on the same gateway URL:
```typescript theme={null}
const agent = new Agent({
id: "resilient-agent",
name: "Resilient Agent",
instructions: "You are a helpful assistant.",
model: [
{
model: {
id: "modellix/openai/gpt-5.5",
url: "https://llm.modellix.ai/v1",
apiKey: process.env.MODELLIX_API_KEY,
},
maxRetries: 2,
},
{
model: {
id: "modellix/google/gemini-3.6-flash",
url: "https://llm.modellix.ai/v1",
apiKey: process.env.MODELLIX_API_KEY,
},
maxRetries: 2,
},
],
});
```
## Troubleshooting
| Symptom | Check |
| --------------------- | ------------------------------------------------------------------------------------------------------ |
| Still hitting OpenAI | `model` is a custom `{ id, url, apiKey }` object or AI SDK provider — not a bare `"openai/..."` string |
| 401 Unauthorized | `apiKey` / `MODELLIX_API_KEY` is a valid Modellix key |
| Model not found / 404 | Upstream ID is full `provider/name`; gateway `id` is `modellix//` |
| Wrong path | `url` / `baseURL` is `https://llm.modellix.ai/v1` (include `/v1`, no `/chat/completions`) |
| Bare model name sent | Use gateway-style `id` (three segments) so Modellix receives `openai/gpt-5.5` |
## Related
* [Mastra model providers](https://mastra.ai/models) — custom OpenAI-compatible `url`, gateways, fallbacks
* [Vercel AI SDK](/llm/sdk/vercel-ai-sdk) — `createOpenAICompatible` details
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` gateway without Mastra
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
# Use Modellix LLM with Microsoft Agent Framework
Source: https://docs.modellix.ai/llm/framework/microsoft-agent-framework
Point Microsoft Agent Framework at Modellix with OpenAIChatCompletionClient base_url, a Modellix API key, and provider/name model IDs.
Use [Microsoft Agent Framework](https://learn.microsoft.com/en-us/agent-framework/overview/?pivots=programming-language-python) against the Modellix LLM gateway. Python clients accept a `base_url` for [any OpenAI-compatible endpoint](https://learn.microsoft.com/en-us/agent-framework/integrations/openai-endpoints?pivots=programming-language-python). Point that URL at Chat Completions (or Responses) on `https://llm.modellix.ai/v1`.
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Pass that full string as `model`. See [Models & Pricing](/llm/overview#models-and-pricing).
Microsoft documents that third-party (non-Azure Direct) models are used at your own risk under their Product Terms. Review data sharing and compliance for your deployment.
## Set Up Microsoft Agent Framework
Python:
```bash theme={null}
pip install agent-framework
```
.NET (OpenAI integration packages; versions may be prerelease):
```bash theme={null}
dotnet add package Microsoft.Agents.AI.OpenAI --prerelease
dotnet add package OpenAI
```
Create a Modellix API key in the [console](https://modellix.ai/console/api-key):
```bash theme={null}
export OPENAI_API_KEY="mdlx-xxxxxxxx"
export OPENAI_BASE_URL="https://llm.modellix.ai/v1"
```
| Setting | Value |
| ------------------------------ | -------------------------------------------- |
| `OPENAI_API_KEY` / `api_key` | Modellix API Key |
| `OPENAI_BASE_URL` / `base_url` | `https://llm.modellix.ai/v1` (include `/v1`) |
| `model` | Exact Modellix Model ID (`provider/name`) |
Agent Framework does **not** load `.env` files automatically. Call `load_dotenv()` yourself, export variables in the shell, or pass `api_key` / `base_url` in code. Do not use the media host (`https://api.modellix.ai`).
Use `OpenAIChatCompletionClient` for maximum compatibility with OpenAI-compatible gateways:
```python theme={null}
import asyncio
from agent_framework.openai import OpenAIChatCompletionClient
async def main() -> None:
agent = OpenAIChatCompletionClient(
base_url="https://llm.modellix.ai/v1",
api_key="mdlx-xxxxxxxx", # or rely on OPENAI_API_KEY
model="openai/gpt-5.5",
).as_agent(
name="Assistant",
instructions="You are a concise coding assistant.",
)
result = await agent.run("Introduce yourself in one sentence.")
print(result)
asyncio.run(main())
```
You can omit `base_url` / `api_key` in the constructor when `OPENAI_BASE_URL` and `OPENAI_API_KEY` are set in the environment.
Change `model` to any catalog ID (`anthropic/claude-sonnet-5`, `google/gemini-3.6-flash`, and so on). Traffic still uses Chat Completions on `https://llm.modellix.ai/v1`.
Modellix also exposes [`POST /v1/responses`](/llm/responses). Use `OpenAIChatClient` when you want the Responses path:
```python theme={null}
import asyncio
from agent_framework.openai import OpenAIChatClient
async def main() -> None:
agent = OpenAIChatClient(
base_url="https://llm.modellix.ai/v1",
api_key="mdlx-xxxxxxxx",
model="openai/gpt-5.5",
).as_agent(
name="Assistant",
instructions="You are a concise coding assistant.",
)
result = await agent.run("Say hello in one short sentence.")
print(result)
async for chunk in agent.run("Tell a one-line joke.", stream=True):
if chunk.text:
print(chunk.text, end="", flush=True)
asyncio.run(main())
```
Prefer Chat Completions first if tools or streaming behave unexpectedly on Responses through a gateway.
Point the OpenAI .NET client at Modellix, then create an agent:
```csharp theme={null}
using System.ClientModel;
using Microsoft.Agents.AI;
using OpenAI;
var endpoint = Environment.GetEnvironmentVariable("OPENAI_BASE_URL")
?? "https://llm.modellix.ai/v1";
var apiKey = Environment.GetEnvironmentVariable("OPENAI_API_KEY")
?? throw new InvalidOperationException("OPENAI_API_KEY is required.");
var model = "openai/gpt-5.5";
AIAgent agent = new OpenAIClient(
new ApiKeyCredential(apiKey),
new OpenAIClientOptions { Endpoint = new Uri(endpoint) })
.GetChatClient(model)
.AsAIAgent(
instructions: "You are a concise coding assistant.",
name: "Assistant");
Console.WriteLine(await agent.RunAsync("Introduce yourself in one sentence."));
```
Keep `Endpoint` as `https://llm.modellix.ai/v1` and use a full Modellix Model ID.
## Troubleshooting
| Symptom | Check |
| --------------------- | --------------------------------------------------------------------------- |
| 401 / auth errors | Key is a valid Modellix API key (`OPENAI_API_KEY` or `api_key`) |
| Model not found / 404 | `model` is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5` |
| Wrong path | `base_url` / `OPENAI_BASE_URL` / `Endpoint` is `https://llm.modellix.ai/v1` |
| Env vars ignored | Framework does not auto-load `.env`; export vars or pass constructor args |
| Responses failures | Fall back to `OpenAIChatCompletionClient` (Chat Completions) |
## Related
* [Microsoft Agent Framework overview](https://learn.microsoft.com/en-us/agent-framework/overview/?pivots=programming-language-python) — agents, harness, workflows
* [OpenAI-compatible endpoints](https://learn.microsoft.com/en-us/agent-framework/integrations/openai-endpoints?pivots=programming-language-python) — `base_url` for Python clients
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` gateway without the agent harness
* [OpenAI Agents SDK](/llm/sdk/openai-agents) — another OpenAI-compatible agent runtime
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
# Get LLM logs
Source: https://docs.modellix.ai/llm/get-llm-logs
/llm/api/llm.json get /logs
Lists LLM request logs for the authenticated team within a time window. Optional `mdlx_user_id` filters by the end-user id sent as `X-Mdlx-User-Id` on inference. Uses the query rate limit (separate from inference RPM).
# Use Modellix LLM with Cline
Source: https://docs.modellix.ai/llm/ide/cline
Point Cline at the Modellix LLM gateway with the OpenAI Compatible provider, Base URL, and provider/name model IDs.
Configure [Cline](https://cline.bot/) to call the Modellix LLM gateway. Cline supports any [OpenAI Compatible](https://docs.cline.bot/provider-config/openai-compatible) endpoint — select that provider, set Base URL to `https://llm.modellix.ai/v1`, and use Modellix `provider/name` Model IDs for Chat Completions on [`POST /v1/chat/completions`](/llm/chat-completions).
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Enter that full string as the Model ID. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up Cline
Install the official extension (publisher **Cline / saoudrizwan**, ID `saoudrizwan.claude-dev`):
| Platform | Install |
| ---------------------------- | -------------------------------------------------------------------------------------------------------------------------- |
| VS Code | [Marketplace](https://marketplace.visualstudio.com/items?itemName=saoudrizwan.claude-dev) or Extensions search for `Cline` |
| Cursor / Windsurf / VSCodium | [Open VSX](https://open-vsx.org/extension/saoudrizwan/claude-dev) or install from a `.vsix` |
| JetBrains | Plugins Marketplace → search `Cline` |
After install, open the Cline icon in the activity bar (or JetBrains tool window).
Create a key in the [Modellix console](https://modellix.ai/console/api-key). You will paste it into Cline settings (Cline stores provider credentials in the IDE, not as shell `OPENAI_*` env vars for this flow).
In the Cline panel, open **Settings** (gear icon). Set:
| Field | Value |
| ------------ | ------------------------------------------------------ |
| API Provider | **OpenAI Compatible** |
| Base URL | `https://llm.modellix.ai/v1` |
| API Key | Your Modellix API Key |
| Model ID | Exact Modellix Model ID (for example `openai/gpt-5.5`) |
Include `/v1` in Base URL. Do **not** append `/chat/completions` — Cline adds that path itself.
Use a Modellix API key, not an OpenAI platform key. Do not set Base URL to the media API host (`https://api.modellix.ai`).
Under **Model Configuration**, set values that match the model (Cline does not always infer these for custom endpoints):
| Setting | Recommendation |
| -------------------- | --------------------------------------------------------------- |
| Context Window | Match the model (for example `200000` for large-context models) |
| Max Output Tokens | For example `8192` (or the model's max) |
| Image Support | Enable only if the model supports vision |
| Computer Use / tools | Enable so Cline can use agent tools |
Save settings. Switch Model ID to any other catalog ID (for example `anthropic/claude-sonnet-5` or `google/gemini-3.6-flash`) on the same Base URL — traffic still uses OpenAI-compatible Chat Completions on `https://llm.modellix.ai/v1`.
You do **not** need a separate Anthropic Base URL without `/v1` for these IDs. Native Anthropic Messages is for the [Anthropic SDK](/llm/sdk/anthropic-sdk) and [Claude Code](/llm/agent/claude-code).
If Cline offers **Verify**, run it after saving Base URL, key, and Model ID.
Then send a short task in the Cline panel, for example: list files in the project root and summarize the project in one sentence. A successful run uses tools (such as list files) and returns a reply without auth or model-not-found errors.
## Troubleshooting
| Symptom | Check |
| ------------------------ | -------------------------------------------------------------------------------- |
| 401 / Invalid API Key | Key is a valid Modellix key with no extra spaces or `Bearer` prefix |
| Model not found / 404 | Model ID is a full string like `openai/gpt-5.5`, not bare `gpt-5.5` |
| Connection / wrong path | Base URL is `https://llm.modellix.ai/v1` (include `/v1`, no `/chat/completions`) |
| Tools / agent steps fail | Enable Computer Use / tool support in Model Configuration |
| Context overflows early | Raise Context Window / Max Output Tokens to match the model |
## Related
* [Cline OpenAI Compatible](https://docs.cline.bot/provider-config/openai-compatible) — official provider fields
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` Chat Completions gateway
* [Kilo Code](/llm/agent/kilo-code) — another OpenAI Compatible coding-agent setup
# Use Modellix LLM with Cursor
Source: https://docs.modellix.ai/llm/ide/cursor
Add Modellix as an OpenAI-compatible provider in Cursor using https://llm.modellix.ai/v1, your Modellix API key, and provider/name model IDs.
Use Modellix models inside [Cursor](https://cursor.com) by adding an OpenAI-compatible provider that targets the LLM gateway. Cursor talks Chat Completions-style APIs; Modellix exposes them at [`POST /v1/chat/completions`](/llm/chat-completions).
Modellix media generation (image / video / speech) uses a different host and async task API. This page covers **LLM text** only (`https://llm.modellix.ai`).
## Set Up Cursor
Create a Modellix API key in the [console](https://modellix.ai/console/api-key). Cursor sends `Authorization: Bearer `, which Modellix accepts.
Choose a model ID in `provider/name` form (for example `openai/gpt-5.5`). See [Models & Pricing](/llm/overview#models-and-pricing) for the full list.
* Prefer `openai/...` or `google/...` for the OpenAI-compatible path (same as [OpenAI SDK](/llm/sdk/openai-sdk)).
* `anthropic/...` can work on Chat Completions for many clients; for Anthropic-native tools use [Messages](/llm/messages) instead (see [Claude Code](/llm/agent/claude-code)).
* Do not assume bare upstream names without the `provider/` prefix.
In Cursor settings, open the models / providers section and add a custom **OpenAI-compatible** (or equivalent override) endpoint:
| Field | Value |
| -------- | ------------------------------------------------------- |
| Base URL | `https://llm.modellix.ai/v1` |
| API key | Your Modellix API Key |
| Model | Exact `provider/name` ID (for example `openai/gpt-5.5`) |
Include `/v1` in the base URL so paths resolve to `/v1/chat/completions`.
If Cursor asks for a full chat completions URL instead of a base URL, use `https://llm.modellix.ai/v1/chat/completions`. Prefer base URL + model ID when the UI supports it.
If your Cursor build allows a separate Anthropic base URL override, use:
| Field | Value |
| -------- | ------------------------------------ |
| Base URL | `https://llm.modellix.ai` (no `/v1`) |
| API key | Same Modellix API Key |
| Model | `anthropic/...` |
Otherwise keep the OpenAI-compatible setup above for Chat Completions.
## Troubleshooting
| Symptom | Check |
| ----------------------- | ----------------------------------------------------------------------- |
| 401 | Key is a Modellix key; Bearer header is present |
| 404 on model | Model ID includes `provider/` and is currently available |
| Wrong path / 404 on URL | Base URL should be `https://llm.modellix.ai/v1`, not the media API host |
| Empty or odd responses | Confirm the selected model supports the request shape Cursor sends |
## Related
* [LLM API guide](/llm/api/api) — auth, protocols, errors
* [OpenAI SDK](/llm/sdk/openai-sdk) — same gateway settings for app code
* [Codex](/llm/agent/codex) · [OpenCode](/llm/agent/opencode) — other OpenAI-compatible clients
# List models
Source: https://docs.modellix.ai/llm/list-models
/llm/api/llm.json get /models
Returns an OpenAI-compatible list of models available on the Modellix LLM gateway. Each `data[].id` uses `provider/name` form. Uses the query rate limit (separate from inference RPM).
# Messages
Source: https://docs.modellix.ai/llm/messages
/llm/api/llm.json post /messages
Anthropic Messages–compatible endpoint on the Modellix LLM gateway. Accepts a synchronous request with `model` (`provider/name`, use `anthropic/...`), Anthropic-style `messages`, and required `max_tokens`; optional `system` prompt is supported. Returns a Messages JSON object by default, or SSE (`text/event-stream`) when `stream=true`. Auth accepts `Authorization: Bearer` or `x-api-key` (same Modellix API Key). Use this path for the Anthropic SDK, Claude Code, and other Anthropic-native clients; optional `anthropic-version` header is forwarded when present. Do not send `google/...` (Gemini) or OpenAI field shapes on this endpoint—use `/chat/completions` or `/responses`. Optional fields follow [Anthropic Messages](https://docs.anthropic.com/en/api/messages); support depends on the selected model.
# Modellix LLM Overview
Source: https://docs.modellix.ai/llm/overview
Get started with the Modellix LLM gateway—OpenAI-compatible Chat Completions and Responses, Anthropic-compatible Messages, and guides for SDKs and coding tools.
Modellix LLM is a model gateway at `https://llm.modellix.ai`. Use one Modellix API key to call OpenAI-compatible **Chat Completions** and **Responses**, or Anthropic-compatible **Messages**, with the same synchronous request model and optional streaming SSE. Some models accept image, audio, or video as **input**—see [Multimodal Inputs](/llm/api/api#multimodal-inputs).
This gateway returns text. Image, video, and speech generation use
`https://api.modellix.ai` with async tasks—see [REST API](/ways-to-use/api).
## What You Get
Chat Completions, Responses, and Messages—each with its own URL and request
body. Do not mix fields across protocols.
Point OpenAI or Anthropic SDKs, Codex, Claude Code, Cursor, and OpenCode at
Modellix with a base URL override.
Authenticate with Bearer or `x-api-key` using a Modellix key from the
[console](https://modellix.ai/console/api-key)—not a vendor platform key.
Pass model IDs like `openai/gpt-5.5`, `anthropic/claude-sonnet-5`, or
`google/gemini-3.6-flash`. See [Modellix LLM](https://www.modellix.ai/llm)
for current IDs and prices.
## Models & Pricing
Pass `model` as a `provider/name` ID (for example `openai/gpt-5.6-sol`). You can also use a Latest Model ID such as `~openai/gpt-latest` to always call the current flagship in that series—Modellix updates the routing when newer versions ship.
For the current Model ID list (including Latest Model IDs), Input Context tiers, discounts, and USD-per-1M-token rates, use the [Modellix LLM](https://www.modellix.ai/llm) page.
Filter by provider and input modality, compare official list prices with Modellix rates, and estimate a call with the cost calculator. Billing uses token `usage` on successful responses.
## Quick Start
Create a key in the [Modellix console](https://modellix.ai/console/api-key)
and store it securely.
| Client type | Base URL | Protocol |
| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------ | ------------------------------- |
| OpenAI SDK, OpenAI Agents SDK, Microsoft Agent Framework, LangChain, Vercel AI SDK, Mastra, CrewAI, Agno, Codex, Cursor, OpenCode, OpenClaw, Hermes, Junie, Pi, CodeBuddy, WorkBuddy, Qwen Code, Kilo Code, DeepSeek Harness, Cline, CC Switch (Codex / OpenAI apps) | `https://llm.modellix.ai/v1` | Chat Completions (or Responses) |
| Anthropic SDK, Claude Code, Claude Agent SDK, CC Switch (Claude) | `https://llm.modellix.ai` (no `/v1`) | Messages |
Use your SDK or a curl example from the [API guide](/llm/api/api). Always set
`model` to a `provider/name` ID.
## Guides in This Section
Start with the API guide for protocols, auth, errors, and billing. Then open the client page that matches your stack.
Protocols, curl examples, session header, errors, rate limits, and billing.
`OPENAI_BASE_URL` + Modellix key for Chat Completions and Responses.
`ANTHROPIC_BASE_URL` without `/v1` for Messages.
`ChatOpenAI` with `base_url` pointed at the LLM gateway.
`createOpenAICompatible` with `baseURL` pointed at the LLM gateway.
Custom OpenAI-compatible `model.url` pointed at the LLM gateway.
`LLM` with `base_url` and `custom_openai` pointed at the LLM gateway.
`OpenAILike` with `base_url` pointed at the LLM gateway.
Custom `AsyncOpenAI` client + Chat Completions for agents.
`OpenAIChatCompletionClient` with `base_url` pointed at the LLM gateway.
Environment variables or `~/.claude/settings.json`.
Custom model Endpoint or `models.json` pointed at Chat Completions.
`openai_base_url` in `~/.codex/config.toml`.
`ANTHROPIC_BASE_URL` without `/v1` for the Agent SDK harness.
OpenAI-compatible provider pointed at the LLM gateway.
Custom provider in the Web UI or `settings.yaml` pointed at the LLM gateway.
Custom Endpoint in `config.yaml` pointed at the LLM gateway.
Custom LLM JSON profile pointed at Chat Completions.
Custom OpenAI Compatible provider in `kilo.jsonc` pointed at the LLM gateway.
Custom `models.providers` entry pointed at the LLM gateway.
Provider `baseURL` for OpenAI- or Anthropic-compatible mode.
Custom provider in `models.json` pointed at the LLM gateway.
`OPENAI_BASE_URL` or `modelProviders` pointed at the LLM gateway.
Custom OpenAI-compatible model pointed at the LLM gateway.
OpenAI Compatible provider pointed at the LLM gateway.
Custom provider for Claude Code, Codex, and OpenAI Compatible apps.
## API Reference
OpenAPI pages for the three endpoints live under **API Reference → LLM**:
| Endpoint | Docs |
| --------------------------- | ----------------------------------------- |
| `POST /v1/chat/completions` | [Chat Completions](/llm/chat-completions) |
| `POST /v1/responses` | [Responses](/llm/responses) |
| `POST /v1/messages` | [Messages](/llm/messages) |
## Protocol Cheat Sheet
| | Chat Completions | Responses | Messages |
| ------------ | -------------------------------------- | ------------------------ | ---------------------------------------- |
| Path | `/v1/chat/completions` | `/v1/responses` | `/v1/messages` |
| Main input | `messages` | `input` | `messages` + optional `system` |
| Length field | `max_tokens` / `max_completion_tokens` | `max_output_tokens` | `max_tokens` (required) |
| Best for | Most OpenAI-compatible tools | OpenAI Responses clients | Anthropic-native tools (`anthropic/...`) |
`google/...` (Gemini) models use Chat Completions or Responses, not Messages.
For full field tables and error codes, see the [LLM API guide](/llm/api/api).
# Responses
Source: https://docs.modellix.ai/llm/responses
/llm/api/llm.json post /responses
OpenAI Responses–compatible endpoint on the Modellix LLM gateway. Accepts a synchronous request with `model` (provider/name format) and `input` (string or content array); returns a Responses JSON object by default, or SSE (`text/event-stream`) when `stream=true`. Use `max_output_tokens` for output limits. Prefer this path when the client targets the OpenAI Responses API rather than classic Chat Completions. Do not send Chat Completions-style `messages`/`max_tokens` or Anthropic Messages shapes on this endpoint—use `/chat/completions` or `/messages` instead. Optional fields follow [OpenAI Responses](https://platform.openai.com/docs/api-reference/responses); support depends on the selected model.
# Use Modellix LLM with the Anthropic SDK
Source: https://docs.modellix.ai/llm/sdk/anthropic-sdk
Point the Anthropic SDK at Modellix with ANTHROPIC_BASE_URL (no /v1) and your Modellix API key, then call Messages with anthropic/... models.
Use Anthropic client libraries against Modellix by pointing the base URL at the LLM gateway and authenticating with a Modellix API key. The gateway speaks the Anthropic Messages protocol at [`POST /v1/messages`](/llm/messages).
Use a Modellix API key from the [console](https://modellix.ai/console/api-key), not an Anthropic Console key. Model IDs use `provider/name` (for example `anthropic/claude-sonnet-5`)—see [Models & Pricing](/llm/overview#models-and-pricing) for the full list.
## Configure the Client
```bash theme={null}
export ANTHROPIC_API_KEY="mdlx-xxxxxxxx"
export ANTHROPIC_BASE_URL="https://llm.modellix.ai"
```
| Setting | Value |
| -------- | --------------------------------------------------------------------------------- |
| Base URL | `https://llm.modellix.ai` (**without** `/v1`; the SDK appends `/v1/messages`) |
| API key | Modellix API Key via `ANTHROPIC_API_KEY` (sent as `x-api-key`) |
| Model | Prefer `anthropic/...`. See [Models & Pricing](/llm/overview#models-and-pricing). |
Do not set `ANTHROPIC_BASE_URL` to `https://llm.modellix.ai/v1`. Including `/v1` breaks path construction for native Anthropic clients.
### Bearer Alternative
If a tool requires Bearer auth instead of `x-api-key`, use `ANTHROPIC_AUTH_TOKEN` with the same Modellix key. Do not set `ANTHROPIC_API_KEY` and `ANTHROPIC_AUTH_TOKEN` to different values.
## Messages Example
```python Python theme={null}
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY and ANTHROPIC_BASE_URL
message = client.messages.create(
model="anthropic/claude-sonnet-5",
max_tokens=256,
messages=[
{"role": "user", "content": "ping"},
],
)
print(message.content)
```
```typescript TypeScript theme={null}
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic(); // reads ANTHROPIC_API_KEY and ANTHROPIC_BASE_URL
const message = await client.messages.create({
model: "anthropic/claude-sonnet-5",
max_tokens: 256,
messages: [{ role: "user", content: "ping" }],
});
console.log(message.content);
```
`max_tokens` is required on Messages. Put system instructions in `system`, not as a `system` role inside `messages` (roles are `user` or `assistant` only).
## Streaming
Set `stream=true` (or the SDK equivalent). The gateway returns Anthropic Messages SSE events.
## Optional Headers
Native Anthropic clients often send `anthropic-version` (for example `2023-06-01`); Modellix forwards it upstream when present. For multi-turn session affinity you can also send `X-Mdlx-Session-Id`. To tag your end users for later log filtering, send `X-Mdlx-User-Id` (see [LLM API](/llm/api/api#end-user-id-header)).
## Related
* [Claude Code](/llm/agent/claude-code) — same base URL pattern with settings persistence
* [LLM API guide](/llm/api/api) — protocols, auth, and errors
* [Create message](/llm/messages) — OpenAPI reference
# Use Modellix LLM with the Claude Agent SDK
Source: https://docs.modellix.ai/llm/sdk/claude-agent-sdk
Point the Claude Agent SDK at Modellix with ANTHROPIC_BASE_URL (no /v1), a Modellix API key, and anthropic/... model IDs.
Use the [Claude Agent SDK](https://code.claude.com/docs/en/agent-sdk/overview) against the Modellix LLM gateway. The SDK drives the same Claude Code agent loop and speaks the Anthropic Messages API. Point it at Modellix the same way as [Claude Code](/llm/agent/claude-code): `ANTHROPIC_BASE_URL` without `/v1`, plus a Modellix API key.
This page covers **LLM text** only (`https://llm.modellix.ai`).
The Agent SDK targets **Anthropic-compatible** endpoints, not OpenAI Chat Completions. Do not set `OPENAI_BASE_URL` or `https://llm.modellix.ai/v1` for this SDK. Use `https://llm.modellix.ai` and `anthropic/...` model IDs. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up the Claude Agent SDK
Python:
```bash theme={null}
pip install claude-agent-sdk
```
TypeScript:
```bash theme={null}
npm install @anthropic-ai/claude-agent-sdk
```
Both packages typically bundle a Claude Code binary. If your install skips optional/native binaries, [install Claude Code](https://code.claude.com/docs/en/setup) and ensure it is on `PATH` (or set the SDK path option for your language).
Create a Modellix API key in the [console](https://modellix.ai/console/api-key), then export:
```bash theme={null}
export ANTHROPIC_API_KEY="mdlx-xxxxxxxx"
export ANTHROPIC_BASE_URL="https://llm.modellix.ai"
```
| Setting | Value |
| -------------------- | ---------------------------------------------------------------- |
| `ANTHROPIC_API_KEY` | Modellix API Key (sent as `x-api-key`) |
| `ANTHROPIC_BASE_URL` | `https://llm.modellix.ai` (**without** `/v1`) |
| Model | Prefer `anthropic/...` (for example `anthropic/claude-sonnet-5`) |
Do not append `/v1` to `ANTHROPIC_BASE_URL`. The Claude Code binary / Anthropic client appends `/v1/messages` itself. Missing this rule commonly causes wrong-path or 404 errors.
If the client expects Bearer auth, use `ANTHROPIC_AUTH_TOKEN` instead of `ANTHROPIC_API_KEY`. Do not set both to different values.
The SDK does not load `.env` files automatically. Export variables in the shell that runs your agent, or pass them through `ClaudeAgentOptions.env` (see next step).
The Python/TypeScript SDK spawns Claude Code and forwards environment variables. Set Modellix credentials in the process environment, or merge them into `options.env` so the subprocess always receives them.
```python Python theme={null}
import asyncio
from claude_agent_sdk import (
AssistantMessage,
ClaudeAgentOptions,
ResultMessage,
query,
)
async def main() -> None:
async for message in query(
prompt="Summarize this directory in one sentence.",
options=ClaudeAgentOptions(
model="anthropic/claude-sonnet-5",
allowed_tools=["Read", "Glob"],
permission_mode="acceptEdits",
env={
"ANTHROPIC_API_KEY": "mdlx-xxxxxxxx",
"ANTHROPIC_BASE_URL": "https://llm.modellix.ai",
},
),
):
if isinstance(message, AssistantMessage):
for block in message.content:
if hasattr(block, "text"):
print(block.text)
elif isinstance(message, ResultMessage):
print(f"Done: {message.subtype}")
asyncio.run(main())
```
```typescript TypeScript theme={null}
import { query } from "@anthropic-ai/claude-agent-sdk";
for await (const message of query({
prompt: "Summarize this directory in one sentence.",
options: {
model: "anthropic/claude-sonnet-5",
allowedTools: ["Read", "Glob"],
permissionMode: "acceptEdits",
env: {
ANTHROPIC_API_KEY: "mdlx-xxxxxxxx",
ANTHROPIC_BASE_URL: "https://llm.modellix.ai",
},
},
})) {
if (message.type === "assistant" && message.message?.content) {
for (const block of message.message.content) {
if ("text" in block) {
console.log(block.text);
}
}
} else if (message.type === "result") {
console.log(`Done: ${message.subtype}`);
}
}
```
Use an exact Modellix Model ID. Bare Anthropic Console names (for example `claude-sonnet-4-6`) are not Modellix catalog IDs.
You can omit `env` in code if `ANTHROPIC_API_KEY` and `ANTHROPIC_BASE_URL` are already exported in the parent process. For production agents, prefer explicit `options.env` (or a secrets manager) so the subprocess does not depend on ambient shell state.
For interactive Claude Code and Agent SDK processes that inherit user settings, you can also persist the same values in `~/.claude/settings.json`:
```json theme={null}
{
"env": {
"ANTHROPIC_BASE_URL": "https://llm.modellix.ai",
"ANTHROPIC_API_KEY": "mdlx-xxxxxxxx"
}
}
```
Details match the [Claude Code](/llm/agent/claude-code) guide. Prefer `options.env` when you need the Agent SDK process to be self-contained.
## Troubleshooting
| Symptom | Check |
| ----------------------- | --------------------------------------------------------------------------- |
| API key not found / 401 | `ANTHROPIC_API_KEY` is a valid Modellix key in the process or `options.env` |
| 404 / wrong path | `ANTHROPIC_BASE_URL` is `https://llm.modellix.ai` with **no** `/v1` |
| Model not found | Model is `anthropic/...` (full Modellix ID), not a bare Console name |
| Still hitting Anthropic | Base URL is Modellix; no conflicting `ANTHROPIC_BASE_URL` elsewhere |
| OpenAI-style 404s | This SDK is not Chat Completions — do not use `https://llm.modellix.ai/v1` |
## Related
* [Claude Agent SDK overview](https://code.claude.com/docs/en/agent-sdk/overview) — capabilities and when to use the SDK
* [Claude Agent SDK quickstart](https://code.claude.com/docs/en/agent-sdk/quickstart) — install, auth, first agent
* [Claude Code](/llm/agent/claude-code) — same `ANTHROPIC_BASE_URL` pattern for the CLI
* [Anthropic SDK](/llm/sdk/anthropic-sdk) — Messages protocol without the agent harness
* [OpenAI Agents SDK](/llm/sdk/openai-agents) — Chat Completions agent path on `/v1`
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [Create message](/llm/messages) — OpenAPI reference
# Use Modellix LLM with the OpenAI Agents SDK
Source: https://docs.modellix.ai/llm/sdk/openai-agents
Point the OpenAI Agents SDK at Modellix with AsyncOpenAI base_url, Chat Completions, and provider/name model IDs.
Use the [OpenAI Agents SDK](https://openai.github.io/openai-agents-python/) against the Modellix LLM gateway. The SDK accepts a custom OpenAI client (`base_url` + API key), so you can run agents on Chat Completions at [`POST /v1/chat/completions`](/llm/chat-completions) — the same gateway shape as the [OpenAI SDK](/llm/sdk/openai-sdk).
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Pass that **full** string to the model layer. Prefer `OpenAIChatCompletionsModel` (or `MultiProvider` with `openai_prefix_mode="model_id"`) so the SDK does not strip the `openai/` prefix before calling Modellix. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up the Agents SDK
Python:
```bash theme={null}
pip install openai-agents
```
TypeScript:
```bash theme={null}
npm install @openai/agents openai
```
Create a Modellix API key in the [console](https://modellix.ai/console/api-key):
```bash theme={null}
export OPENAI_API_KEY="mdlx-xxxxxxxx"
export OPENAI_BASE_URL="https://llm.modellix.ai/v1"
```
| Setting | Value |
| ----------------- | --------------------------------------------- |
| `OPENAI_API_KEY` | Modellix API Key (not an OpenAI platform key) |
| `OPENAI_BASE_URL` | `https://llm.modellix.ai/v1` (include `/v1`) |
| Model | Exact Modellix Model ID (`provider/name`) |
OpenAI Tracing exports to OpenAI's platform. A Modellix key cannot upload traces. Disable tracing, or keep model traffic on Modellix and set a separate OpenAI tracing key.
Build an `AsyncOpenAI` / `OpenAI` client pointed at Modellix, wrap it in `OpenAIChatCompletionsModel` with a literal Modellix Model ID, and disable tracing unless you have a separate OpenAI tracing key.
```python Python theme={null}
import asyncio
from openai import AsyncOpenAI
from agents import (
Agent,
OpenAIChatCompletionsModel,
Runner,
set_tracing_disabled,
)
set_tracing_disabled(True)
client = AsyncOpenAI(
base_url="https://llm.modellix.ai/v1",
api_key="mdlx-xxxxxxxx", # or rely on OPENAI_API_KEY
)
agent = Agent(
name="Assistant",
instructions="You are a concise coding assistant.",
model=OpenAIChatCompletionsModel(
model="openai/gpt-5.5",
openai_client=client,
),
)
async def main() -> None:
result = await Runner.run(agent, "Say hello in one short sentence.")
print(result.final_output)
if __name__ == "__main__":
asyncio.run(main())
```
```typescript TypeScript theme={null}
import { OpenAI } from "openai";
import {
Agent,
OpenAIChatCompletionsModel,
Runner,
setTracingDisabled,
} from "@openai/agents";
setTracingDisabled(true);
const client = new OpenAI({
baseURL: "https://llm.modellix.ai/v1",
apiKey: process.env.OPENAI_API_KEY ?? "mdlx-xxxxxxxx",
});
const agent = new Agent({
name: "Assistant",
instructions: "You are a concise coding assistant.",
model: new OpenAIChatCompletionsModel(client, "openai/gpt-5.5"),
});
const result = await Runner.run(agent, "Say hello in one short sentence.");
console.log(result.finalOutput);
```
Change the model string to any catalog ID (`anthropic/claude-sonnet-5`, `google/gemini-3.6-flash`, and so on). Traffic still uses Chat Completions on `https://llm.modellix.ai/v1`.
To make Modellix the process-wide default without wrapping every agent:
```python Python theme={null}
from openai import AsyncOpenAI
from agents import (
set_default_openai_api,
set_default_openai_client,
set_tracing_disabled,
)
set_tracing_disabled(True)
set_default_openai_api("chat_completions")
set_default_openai_client(
AsyncOpenAI(
base_url="https://llm.modellix.ai/v1",
api_key="mdlx-xxxxxxxx",
),
use_for_tracing=False,
)
```
```typescript TypeScript theme={null}
import { OpenAI } from "openai";
import {
setDefaultOpenAIClient,
setOpenAIAPI,
setTracingDisabled,
} from "@openai/agents";
setTracingDisabled(true);
setOpenAIAPI("chat_completions");
setDefaultOpenAIClient(
new OpenAI({
baseURL: "https://llm.modellix.ai/v1",
apiKey: process.env.OPENAI_API_KEY ?? "mdlx-xxxxxxxx",
}),
);
```
Then assign models carefully. String IDs like `openai/gpt-5.5` can be rewritten by the default OpenAI provider (it may drop the `openai/` prefix). Prefer `OpenAIChatCompletionsModel` with the full Modellix ID, or a `MultiProvider` with `openai_prefix_mode="model_id"` (Python) so the gateway receives `openai/gpt-5.5` unchanged. See [Agents SDK models](https://openai.github.io/openai-agents-python/models/).
Modellix also exposes [`POST /v1/responses`](/llm/responses). The SDK defaults to Responses for OpenAI-hosted traffic. On a compatible gateway, start with Chat Completions (`set_default_openai_api("chat_completions")` / `setOpenAIAPI("chat_completions")`) unless you have verified Responses + tools for your agent loop.
Keep `base_url` as `https://llm.modellix.ai/v1` for both shapes. Do not use the Anthropic host without `/v1` here — that path is for the [Anthropic SDK](/llm/sdk/anthropic-sdk) and [Claude Code](/llm/agent/claude-code).
## Troubleshooting
| Symptom | Check |
| -------------------------- | ------------------------------------------------------------------------------------------- |
| 401 / auth errors | Key is a valid Modellix API key |
| Model not found / 404 | Full ID like `openai/gpt-5.5` reaches Modellix; avoid provider prefix stripping |
| Wrong path | `base_url` is `https://llm.modellix.ai/v1` (include `/v1`) |
| Tracing / export errors | `set_tracing_disabled(True)` or a separate OpenAI tracing key with `use_for_tracing=False` |
| Tools / Responses failures | Switch to Chat Completions (`OpenAIChatCompletionsModel` or `chat_completions` default API) |
## Related
* [OpenAI Agents SDK — Models](https://openai.github.io/openai-agents-python/models/) — custom clients and Chat Completions
* [OpenAI Agents SDK — Configuration](https://openai.github.io/openai-agents-python/config/) — `OPENAI_BASE_URL`, tracing, API shape
* [OpenAI platform — Agents models](https://developers.openai.com/api/docs/guides/agents/models) — model selection overview
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` gateway without the Agents runtime
* [LangChain](/llm/framework/langchain) — another Chat Completions client path
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
# Use Modellix LLM with the OpenAI SDK
Source: https://docs.modellix.ai/llm/sdk/openai-sdk
Point the OpenAI SDK at the Modellix LLM gateway with OPENAI_BASE_URL and your Modellix API key, then call Chat Completions or Responses with provider/name models.
Use the official OpenAI client libraries against Modellix by setting the base URL to the LLM gateway and using a Modellix API key. Requests are synchronous and can stream over SSE.
Use a Modellix API key from the [console](https://modellix.ai/console/api-key), not an OpenAI platform key. Model IDs use `provider/name` (for example `openai/gpt-5.5`)—see [Models & Pricing](/llm/overview#models-and-pricing) for the full list.
## Configure the Client
```bash theme={null}
export OPENAI_API_KEY="mdlx-xxxxxxxx"
export OPENAI_BASE_URL="https://llm.modellix.ai/v1"
```
| Setting | Value |
| -------- | ------------------------------------------------------------------------------------------------------------------------------ |
| Base URL | `https://llm.modellix.ai/v1` (include `/v1`) |
| API key | Modellix API Key via `OPENAI_API_KEY` |
| Model | Prefer `openai/...` (or `google/...` on the OpenAI-compatible path). See [Models & Pricing](/llm/overview#models-and-pricing). |
Then call the SDK as usual. Chat Completions maps to [`POST /v1/chat/completions`](/llm/chat-completions); Responses maps to [`POST /v1/responses`](/llm/responses).
## Chat Completions Example
```python Python theme={null}
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY and OPENAI_BASE_URL
completion = client.chat.completions.create(
model="openai/gpt-5.5",
messages=[
{"role": "user", "content": "Introduce yourself in one sentence"},
],
max_tokens=256,
)
print(completion.choices[0].message.content)
```
```typescript TypeScript theme={null}
import OpenAI from "openai";
const client = new OpenAI(); // reads OPENAI_API_KEY and OPENAI_BASE_URL
const completion = await client.chat.completions.create({
model: "openai/gpt-5.5",
messages: [{ role: "user", content: "Introduce yourself in one sentence" }],
max_tokens: 256,
});
console.log(completion.choices[0].message.content);
```
## Streaming
Set `stream=true` (or the SDK equivalent). The gateway returns `text/event-stream` with OpenAI-style `chat.completion.chunk` events ending in `data: [DONE]`.
```python theme={null}
stream = client.chat.completions.create(
model="openai/gpt-5.6-sol",
messages=[{"role": "user", "content": "ping"}],
stream=True,
max_tokens=256,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)
```
## Responses API
If your SDK or app targets the OpenAI Responses API, keep the same base URL and key, but send Responses fields (`input`, `max_output_tokens`) instead of Chat Completions `messages` / `max_tokens`. See [Create response](/llm/responses).
## Related
* [LLM API guide](/llm/api/api) — protocols, auth, errors, and curl examples
* [Codex](/llm/agent/codex) — CLI config with `openai_base_url`
* [OpenCode](/llm/agent/opencode) — provider `baseURL` for OpenAI-compatible mode
# Use Modellix LLM with the Vercel AI SDK
Source: https://docs.modellix.ai/llm/sdk/vercel-ai-sdk
Point the Vercel AI SDK at Modellix with createOpenAICompatible, baseURL, and provider/name model IDs.
Use the [Vercel AI SDK](https://ai-sdk.dev/docs/foundations/providers-and-models) against the Modellix LLM gateway. Create an [OpenAI Compatible](https://ai-sdk.dev/providers/openai-compatible-providers) provider with `baseURL` set to `https://llm.modellix.ai/v1`, then call `generateText` / `streamText` with Modellix `provider/name` model IDs.
This page covers **LLM text** only (`https://llm.modellix.ai`).
Modellix model IDs are `provider/name` (for example `openai/gpt-5.5`). Pass that full string to the provider model factory. See [Models & Pricing](/llm/overview#models-and-pricing).
## Set Up the AI SDK
```bash theme={null}
npm install ai @ai-sdk/openai-compatible
```
For the alternate `@ai-sdk/openai` path below, also install `@ai-sdk/openai`. For Anthropic Messages, install `@ai-sdk/anthropic`.
Create a Modellix API key in the [console](https://modellix.ai/console/api-key):
```bash theme={null}
export MODELLIX_API_KEY="mdlx-xxxxxxxx"
```
| Setting | Value |
| --------- | ---------------------------------------------------------------------------- |
| API key | Modellix API Key (`MODELLIX_API_KEY`) |
| `baseURL` | `https://llm.modellix.ai/v1` (include `/v1` for OpenAI-compatible providers) |
| Model | Exact Modellix Model ID (`provider/name`) |
Use a Modellix key, not an OpenAI platform key. Do not point the AI SDK at the media host (`https://api.modellix.ai`).
Prefer `@ai-sdk/openai-compatible` for third-party gateways:
```typescript theme={null}
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { generateText } from "ai";
const modellix = createOpenAICompatible({
name: "modellix",
apiKey: process.env.MODELLIX_API_KEY,
baseURL: "https://llm.modellix.ai/v1",
});
const { text } = await generateText({
model: modellix("openai/gpt-5.5"),
prompt: "Introduce yourself in one sentence",
});
console.log(text);
```
Swap the model string for any catalog ID (`anthropic/claude-sonnet-5`, `google/gemini-3.6-flash`, and so on). Traffic still uses Chat Completions on `https://llm.modellix.ai/v1`.
```typescript theme={null}
import { streamText } from "ai";
const result = streamText({
model: modellix("openai/gpt-5.5"),
prompt: "Write a one-line haiku about APIs",
});
for await (const textPart of result.textStream) {
process.stdout.write(textPart);
}
```
You can also customize `@ai-sdk/openai`. The default `modellix("...")` factory targets the **Responses** API. Modellix supports Responses, but for Chat Completions–only client code paths use `.chat()`:
```bash theme={null}
npm install @ai-sdk/openai
```
```typescript theme={null}
import { createOpenAI } from "@ai-sdk/openai";
import { generateText } from "ai";
const modellix = createOpenAI({
apiKey: process.env.MODELLIX_API_KEY,
baseURL: "https://llm.modellix.ai/v1",
name: "modellix",
});
// Chat Completions
const chat = await generateText({
model: modellix.chat("openai/gpt-5.5"),
prompt: "ping",
});
// Responses (same host + /v1)
const responses = await generateText({
model: modellix("openai/gpt-5.5"),
prompt: "ping",
});
console.log(chat.text, responses.text);
```
See [AI SDK OpenAI provider](https://ai-sdk.dev/providers/ai-sdk-providers/openai) for `baseURL` and `.chat()` details.
For native Anthropic Messages instead of OpenAI-compatible Chat Completions:
```bash theme={null}
npm install @ai-sdk/anthropic
```
```typescript theme={null}
import { createAnthropic } from "@ai-sdk/anthropic";
import { generateText } from "ai";
const modellix = createAnthropic({
apiKey: process.env.MODELLIX_API_KEY,
baseURL: "https://llm.modellix.ai", // no /v1
});
const { text } = await generateText({
model: modellix("anthropic/claude-sonnet-5"),
prompt: "Introduce yourself in one sentence",
});
```
Same host shape as the [Anthropic SDK](/llm/sdk/anthropic-sdk): **no** `/v1` on `baseURL`. Most AI SDK apps should stay on `createOpenAICompatible` + `/v1` unless you specifically need Messages.
## Troubleshooting
| Symptom | Check |
| ------------------------ | ---------------------------------------------------------------------------------- |
| 401 Unauthorized | `MODELLIX_API_KEY` / `apiKey` is a valid Modellix key |
| 404 / wrong path | OpenAI-compatible `baseURL` includes `/v1`; Anthropic `baseURL` does **not** |
| Model not found | Model is a full ID like `openai/gpt-5.5`, not bare `gpt-5.5` |
| Responses vs Chat errors | Use `createOpenAICompatible` or `createOpenAI(...).chat(...)` for Chat Completions |
| Still hitting OpenAI | Custom provider `baseURL` is Modellix, not the default OpenAI host |
## Related
* [AI SDK — Providers and models](https://ai-sdk.dev/docs/foundations/providers-and-models) — provider architecture
* [OpenAI Compatible providers](https://ai-sdk.dev/providers/openai-compatible-providers) — `createOpenAICompatible`
* [OpenAI provider](https://ai-sdk.dev/providers/ai-sdk-providers/openai) — `createOpenAI`, `.chat()`, Responses
* [OpenAI SDK](/llm/sdk/openai-sdk) — same `/v1` gateway without the AI SDK
* [LangChain](/llm/framework/langchain) — another Chat Completions client path
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
# Use Modellix LLM with CC Switch
Source: https://docs.modellix.ai/llm/tool/cc-switch
Add Modellix as a custom CC Switch provider for Claude Code, Codex, and OpenAI Compatible apps with the correct Base URL per protocol.
[CC Switch](https://ccswitch.io/) manages API providers for Claude Code, Codex, OpenCode, OpenClaw, Hermes, and related tools. Modellix is **not** a built-in preset — add a [custom provider](https://ccswitch.io/en/docs?section=providers\&item=add) and point each app at the Modellix LLM gateway.
This page covers **LLM text** only (`https://llm.modellix.ai`).
Base URL differs by protocol:
| App / protocol | Endpoint | Model IDs |
| ----------------------------------------------- | ---------------------------------------- | ------------------------------------------------- |
| Claude Code (Anthropic Messages) | `https://llm.modellix.ai` (**no** `/v1`) | `anthropic/...` |
| Codex / OpenCode / OpenClaw (OpenAI-compatible) | `https://llm.modellix.ai/v1` | `openai/...`, `anthropic/...`, `google/...`, etc. |
See [Models & Pricing](/llm/overview#models-and-pricing). Prefer **app-specific** providers over a single Unified Provider — Claude and Codex need different Base URL shapes.
## Set Up CC Switch
Download and install [CC Switch](https://ccswitch.io/) for your OS. Open the app and select the target tool tab (for example **Claude Code** or **Codex**).
Create a key in the [Modellix console](https://modellix.ai/console/api-key). You will paste it into the custom provider form.
On the **Claude Code** tab, click **+** → choose the **Custom** preset (not AiHubMix or another built-in).
In the JSON config (or form fields that map to it), set:
```json theme={null}
{
"env": {
"ANTHROPIC_API_KEY": "mdlx-xxxxxxxx",
"ANTHROPIC_BASE_URL": "https://llm.modellix.ai"
}
}
```
| Field | Value |
| --------------------- | --------------------------------------------------------- |
| Name | `Modellix` (display name) |
| `ANTHROPIC_API_KEY` | Your Modellix API Key |
| `ANTHROPIC_BASE_URL` | `https://llm.modellix.ai` (**without** `/v1`) |
| API Format (Advanced) | **Anthropic Messages** (default) |
| Model | `anthropic/...` (for example `anthropic/claude-sonnet-5`) |
Do **not** append `/v1` to `ANTHROPIC_BASE_URL`. Claude Code appends `/v1/messages` itself. Leave **Full URL Mode** off for this setup. Do not switch API Format to OpenAI Chat Completions unless you intentionally use CC Switch proxy conversion — Modellix already speaks native Messages at this host.
Save, then **Enable** the Modellix provider card so Claude Code uses it. Details match the standalone [Claude Code](/llm/agent/claude-code) guide.
On the **Codex** tab, click **+** → **Custom**. Configure auth and TOML (CC Switch writes `~/.codex/auth.json` and `~/.codex/config.toml`):
**auth.json**
```json theme={null}
{
"OPENAI_API_KEY": "mdlx-xxxxxxxx"
}
```
**config.toml**
```toml theme={null}
model_provider = "modellix"
model = "openai/gpt-5.5"
disable_response_storage = true
[model_providers.modellix]
name = "Modellix"
base_url = "https://llm.modellix.ai/v1"
wire_api = "responses"
requires_openai_auth = true
```
| Field | Value |
| ---------------- | -------------------------------------------------------------------------------------------------------- |
| `OPENAI_API_KEY` | Modellix API Key |
| `model_provider` | Must match `[model_providers.xxx]` (for example `modellix`) |
| `base_url` | `https://llm.modellix.ai/v1` (include `/v1`) |
| `wire_api` | `responses` (Modellix [`POST /v1/responses`](/llm/responses)); use `chat` if you prefer Chat Completions |
| `model` | Exact Modellix Model ID (for example `openai/gpt-5.5`) |
Save and enable the provider. Related: [Codex](/llm/agent/codex).
On the **OpenCode** or **OpenClaw** tab, prefer the **OpenAI Compatible** preset (or **Custom** if you need a full provider block).
| Field | Value |
| -------------------- | ---------------------------------------------------- |
| Base URL / `baseURL` | `https://llm.modellix.ai/v1` |
| API Key | Modellix API Key |
| Model ID | Exact `provider/name` (for example `openai/gpt-5.5`) |
For hand-edited OpenCode / OpenClaw configs, see [OpenCode](/llm/agent/opencode) and [OpenClaw](/llm/agent/openclaw).
Enter Modellix Model IDs manually when needed. **Fetch Models** calls OpenAI-compatible [`GET /v1/models`](/llm/api/api#list-models). You can also paste IDs from [Models & Pricing](/llm/overview#models-and-pricing).
Enable the Modellix provider for the active app, then run that CLI/IDE as usual.
## Troubleshooting
| Symptom | Check |
| ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| 401 / auth errors | Key is a valid Modellix key for the selected app |
| Claude path / 404 | `ANTHROPIC_BASE_URL` is `https://llm.modellix.ai` with **no** `/v1` |
| Codex path errors | `base_url` is `https://llm.modellix.ai/v1` with `/v1` |
| Model not found | Full ID like `openai/gpt-5.5` or `anthropic/claude-sonnet-5`, not a bare vendor name |
| Fetch Models empty | Check API Key and [List models](/llm/api/api#list-models); or paste Model IDs manually from [Models & Pricing](/llm/overview#models-and-pricing) |
| Unified Provider fails across apps | Use separate Claude vs Codex providers (different Base URL shapes) |
| Wrong path after Full URL Mode | Turn **Full URL Mode** off for standard Modellix endpoints |
## Related
* [CC Switch — Add Provider](https://ccswitch.io/en/docs?section=providers\&item=add) — Custom preset and API Format
* [Claude Code](/llm/agent/claude-code) — same Anthropic env pattern without CC Switch
* [Codex](/llm/agent/codex) — same OpenAI-compatible gateway
* [OpenCode](/llm/agent/opencode) / [OpenClaw](/llm/agent/openclaw) — direct config when not using CC Switch
* [Models & Pricing](/llm/overview#models-and-pricing) — Model IDs and rates
* [LLM API guide](/llm/api/api) — protocols and curl examples
# STT Result Schema
Source: https://docs.modellix.ai/media-model-api/stt-result-schema
Normalized STT result JSON (modellix.transcript.v1) from document resources—full text, channels, sentence and word timings, and optional speaker diarization.
This page describes the **JSON body** of a speech-to-text result file. It is **not** an HTTP API.
When an STT task succeeds, [Query Task Result](/api/get-task-result) returns `result.resources[]` with `type` set to `document`. Download `resources[].url`; the response body matches **modellix.transcript.v1** (normalized across providers, not vendor-raw JSON).
Used by Fun-ASR, Grok Voice ASR, Whisper, MAI-Transcribe, Gemini 3.5 Transcribe, and other STT models.
## How to Obtain
Call a speech-to-text model endpoint. The platform returns an async task (`task_id`).
Call [Query Task Result](/api/get-task-result) until `data.status` is `success`.
Find `result.resources[]` where `type` is `document`, then `GET` that `url`. The body is the schema below.
## Root Object
| Field | Type | Required | Description |
| ------------- | ------- | -------- | ----------------------------------------------------------------------------- |
| `schema` | string | yes | Always `modellix.transcript.v1` |
| `text` | string | yes | Full transcript text (non-empty) |
| `language` | string | no | Detected or provider language label (may be a display name such as `English`) |
| `duration_ms` | integer | no | Audio duration in milliseconds |
| `channels` | array | no | Per-channel transcripts. See [Channel](#channel) |
## Channel
| Field | Type | Required | Description |
| ------------ | ------- | -------- | -------------------------------------------- |
| `channel_id` | integer | yes | Channel index (0-based) |
| `text` | string | no | Channel-level full text |
| `sentences` | array | no | Sentence segments. See [Sentence](#sentence) |
## Sentence
| Field | Type | Required | Description |
| ------------ | ------- | -------- | ----------------------------------------------------------------------------------- |
| `begin_ms` | integer | yes | Start time in milliseconds |
| `end_ms` | integer | yes | End time in milliseconds |
| `text` | string | yes | Sentence text |
| `speaker_id` | integer | no | Speaker id when diarization is enabled and speakers are consistent for the sentence |
| `words` | array | no | Word-level timings. See [Word](#word) |
## Word
| Field | Type | Required | Description |
| ------------ | ------- | -------- | -------------------------------------- |
| `begin_ms` | integer | yes | Start time in milliseconds |
| `end_ms` | integer | yes | End time in milliseconds |
| `text` | string | yes | Word text |
| `speaker_id` | integer | no | Speaker id when diarization is enabled |
## Examples
### Single-Channel with Word Timings
```json theme={null}
{
"schema": "modellix.transcript.v1",
"text": "The balance is $100.",
"language": "English",
"duration_ms": 3450,
"channels": [
{
"channel_id": 0,
"text": "The balance is $100.",
"sentences": [
{
"begin_ms": 240,
"end_ms": 3200,
"text": "The balance is $100.",
"words": [
{ "begin_ms": 240, "end_ms": 480, "text": "The" },
{ "begin_ms": 480, "end_ms": 960, "text": "balance" },
{ "begin_ms": 960, "end_ms": 1120, "text": "is" },
{ "begin_ms": 1120, "end_ms": 3200, "text": "$100." }
]
}
]
}
]
}
```
### With Speaker Diarization
```json theme={null}
{
"schema": "modellix.transcript.v1",
"text": "Hello there",
"duration_ms": 1000,
"channels": [
{
"channel_id": 0,
"text": "Hello there",
"sentences": [
{
"begin_ms": 0,
"end_ms": 1000,
"text": "Hello there",
"words": [
{ "begin_ms": 0, "end_ms": 400, "text": "Hello", "speaker_id": 0 },
{ "begin_ms": 400, "end_ms": 1000, "text": "there", "speaker_id": 1 }
]
}
]
}
]
}
```
# MAI Image 2.5
Source: https://docs.modellix.ai/microsoft/mai-image-2-5
/media-model-api/microsoft/microsoft-t2i.json post /microsoft/mai-image-2.5
[Core Function] MAI Image 2.5 is Microsoft's flagship text-to-image generation model. [Strengths] It excels at producing high-quality, detailed images from a text prompt with precise control over output dimensions. [Best For] Highly recommended for: concept art, marketing visuals, and high-fidelity image generation. [Limitations] Do NOT request dimensions below 768 on any side or a width x height product above 1,048,576 pixels (e.g. beyond 1024x1024); output is always PNG. [Routing] Use this model by default for quality-sensitive generation. For faster, cheaper generation, route to MAI Image 2.5 Flash.
# MAI Image 2.5 Edit
Source: https://docs.modellix.ai/microsoft/mai-image-2-5-edit
/media-model-api/microsoft/microsoft-i2i.json post /microsoft/mai-image-2.5-edit
[Core Function] MAI Image 2.5 Edit is Microsoft's flagship image editing model. [Strengths] It excels at applying high-quality, prompt-guided edits and transformations to a single source image. [Best For] Highly recommended for: restyling, object/scene modification, and detailed prompt-driven edits. [Limitations] Do NOT provide more than one source image or non-JPEG/PNG inputs; it accepts exactly one JPEG or PNG image, and output is always PNG. [Routing] Use this model by default for quality-sensitive edits. For faster, cheaper edits, route to MAI Image 2.5 Flash Edit.
# MAI Image 2.5 Flash
Source: https://docs.modellix.ai/microsoft/mai-image-2-5-flash
/media-model-api/microsoft/microsoft-t2i.json post /microsoft/mai-image-2.5-flash
[Core Function] MAI Image 2.5 Flash is Microsoft's fast, cost-efficient text-to-image generation model. [Strengths] It excels at quickly generating solid images from a text prompt with the same dimension controls as MAI Image 2.5. [Best For] Highly recommended for: rapid prototyping, batch generation, and cost-sensitive workloads. [Limitations] Do NOT request dimensions below 768 on any side or a width x height product above 1,048,576 pixels; output is always PNG, and maximum fidelity is lower than MAI Image 2.5. [Routing] Choose this model when speed or cost matters more than maximum fidelity. For the highest quality, use MAI Image 2.5.
# MAI Image 2.5 Flash Edit
Source: https://docs.modellix.ai/microsoft/mai-image-2-5-flash-edit
/media-model-api/microsoft/microsoft-i2i.json post /microsoft/mai-image-2.5-flash-edit
[Core Function] MAI Image 2.5 Flash Edit is Microsoft's fast, cost-efficient image editing model. [Strengths] It excels at quickly applying prompt-guided edits to a single source image. [Best For] Highly recommended for: rapid edits, batch processing, and cost-sensitive workloads. [Limitations] Do NOT provide more than one source image or non-JPEG/PNG inputs; it accepts exactly one JPEG or PNG image, output is always PNG, and fidelity is lower than MAI Image 2.5 Edit. [Routing] Choose this model when speed or cost matters more than maximum fidelity. For the highest quality, use MAI Image 2.5 Edit.
# MAI Image 2.5 Pro
Source: https://docs.modellix.ai/microsoft/mai-image-2-5-pro
/media-model-api/microsoft/microsoft-t2i.json post /microsoft/mai-image-2.5-pro
[Core Function] MAI Image 2.5 Pro is Microsoft's highest-fidelity text-to-image generation model in the MAI 2.5 family. [Strengths] It excels at producing exceptionally detailed, high-quality images from a text prompt with the same dimension controls as MAI Image 2.5. [Best For] Highly recommended for: premium marketing visuals, hero assets, and maximum-fidelity generation. [Limitations] Do NOT request dimensions below 768 on any side or a width x height product above 1,048,576 pixels; output is always PNG. [Routing] Use this model when maximum quality is required. For faster or lower-cost generation, route to MAI Image 2.5 or MAI Image 2.5 Flash.
# MAI Image 2.5 Pro Edit
Source: https://docs.modellix.ai/microsoft/mai-image-2-5-pro-edit
/media-model-api/microsoft/microsoft-i2i.json post /microsoft/mai-image-2.5-pro-edit
[Core Function] MAI Image 2.5 Pro Edit is Microsoft's highest-fidelity image editing model in the MAI 2.5 family. [Strengths] It excels at applying premium, prompt-guided edits and transformations to a single source image. [Best For] Highly recommended for: hero asset retouching, high-fidelity restyling, and detailed prompt-driven edits. [Limitations] Do NOT provide more than one source image or non-JPEG/PNG inputs; it accepts exactly one JPEG or PNG image, and output is always PNG. [Routing] Use this model when maximum edit quality is required. For faster or lower-cost edits, route to MAI Image 2.5 Edit or MAI Image 2.5 Flash Edit.
# MAI Image 2.6
Source: https://docs.modellix.ai/microsoft/mai-image-2-6
/media-model-api/microsoft/microsoft-t2i.json post /microsoft/mai-image-2.6
[Core Function] MAI Image 2.6 is Microsoft's latest text-to-image generation model in the MAI Image family. [Strengths] It improves text rendering, portraits, 3D imagery, and commercial photorealistic output compared with MAI Image 2.5. [Best For] Highly recommended for: marketing visuals, product hero shots, portraits, and prompts that need accurate on-image text. [Limitations] Do NOT use this if the requested width or height is below 768, or if width x height exceeds 1,048,576 pixels. Output is always PNG. auto_aspect_ratio and web_grounding are optional booleans; omit them unless the caller sets them. [Routing] Use this model when the latest MAI quality is required. For lower latency or cost, route to MAI Image 2.6 Flash.
# MAI Image 2.6 Edit
Source: https://docs.modellix.ai/microsoft/mai-image-2-6-edit
/media-model-api/microsoft/microsoft-i2i.json post /microsoft/mai-image-2.6-edit
[Core Function] MAI Image 2.6 Edit is Microsoft's latest prompt-guided image editing model in the MAI Image family. [Strengths] It applies targeted edits to a single source image with the same quality gains as MAI Image 2.6 generation. [Best For] Highly recommended for: object edits, layout changes, text cleanup, and iterative photorealistic retouching. [Limitations] Do NOT use this if the caller provides more than one source image, a data URI, or a URL that is not publicly reachable over HTTP or HTTPS. It accepts exactly one JPEG or PNG image URL, and output is always PNG. auto_aspect_ratio and web_grounding are optional booleans; omit them unless the caller sets them. [Routing] Use this model when the latest MAI edit quality is required. For faster or lower-cost edits, route to MAI Image 2.6 Flash Edit.
# MAI Image 2.6 Flash
Source: https://docs.modellix.ai/microsoft/mai-image-2-6-flash
/media-model-api/microsoft/microsoft-t2i.json post /microsoft/mai-image-2.6-flash
[Core Function] MAI Image 2.6 Flash is the faster, lower-cost variant of MAI Image 2.6 text-to-image generation. [Strengths] It targets similar quality to MAI Image 2.6 with lower latency for high-throughput workloads. [Best For] Highly recommended for: production pipelines, batch generation, and latency-sensitive image APIs. [Limitations] Do NOT use this if the requested width or height is below 768, or if width x height exceeds 1,048,576 pixels. Output is always PNG. auto_aspect_ratio and web_grounding are optional booleans; omit them unless the caller sets them. [Routing] Use this model when speed or cost matters more than maximum 2.6 quality. For the highest fidelity, route to MAI Image 2.6.
# MAI Image 2.6 Flash Edit
Source: https://docs.modellix.ai/microsoft/mai-image-2-6-flash-edit
/media-model-api/microsoft/microsoft-i2i.json post /microsoft/mai-image-2.6-flash-edit
[Core Function] MAI Image 2.6 Flash Edit is the faster, lower-cost variant of MAI Image 2.6 image editing. [Strengths] It applies prompt-guided edits to a single source image with lower latency than MAI Image 2.6 Edit. [Best For] Highly recommended for: high-throughput edit APIs and production retouching pipelines. [Limitations] Do NOT use this if the caller provides more than one source image, a data URI, or a URL that is not publicly reachable over HTTP or HTTPS. It accepts exactly one JPEG or PNG image URL, and output is always PNG. auto_aspect_ratio and web_grounding are optional booleans; omit them unless the caller sets them. [Routing] Use this model when edit speed or cost matters more than maximum 2.6 quality. For the highest fidelity, route to MAI Image 2.6 Edit.
# MAI Transcribe 1.5
Source: https://docs.modellix.ai/microsoft/mai-transcribe-1-5
/media-model-api/microsoft/microsoft-s2t.json post /microsoft/mai-transcribe-1.5
[Core Function] MAI-Transcribe 1.5 transcribes a single public audio URL into text via an async task. [Strengths] Multi-lingual recognition, optional locale forcing, phrase-list biasing, and word-level timestamps. [Best For] Meeting notes, captions, and batch audio-to-text. [Limitations] Do NOT use file upload; URL-only input. Do NOT expect speaker diarization. Supported audio: WAV, MP3, or FLAC up to 300 MB. [Routing] Use this model for Microsoft MAI speech recognition quality with URL-based audio.
# Hailuo 02 FL2V
Source: https://docs.modellix.ai/minimax/hailuo-02-fl2v
/media-model-api/minimax/minimax-i2v.json post /minimax/hailuo-02-fl2v
[Core Function] Hailuo 02 FL2V is a First-Last frame transition video model. [Strengths] It excels at generating a logical, physically accurate video transition that bridges a provided starting frame and an ending frame. [Best For] Highly recommended for: visual morphing, before-and-after transitions, and precise narrative storyboard completion. [Limitations] Do NOT use this model if you only have one image (use standard I2V instead). Note that Hailuo 2.3 does not support FL2V, so this is the primary transition model. [Routing] Use this model exclusively when the user provides BOTH a first frame and a last frame for a transition.
# Hailuo 02 I2V
Source: https://docs.modellix.ai/minimax/hailuo-02-i2v
/media-model-api/minimax/minimax-i2v.json post /minimax/hailuo-02-i2v
[Core Function] Hailuo 02 I2V is an image-to-video model optimized for physical realism and sustained high resolution. [Strengths] It excels at animating broad scenes, maintaining complex physics, and supporting native 1080p generation for up to 10 seconds. [Best For] Highly recommended for: animating product photography, bringing landscape/nature photos to life, and generating physically accurate motion. [Limitations] Do NOT use this model for animating complex human facial micro-expressions or highly stylized anime art, where Hailuo 2.3 is superior. [Routing] Use this model when the user needs to animate a landscape/product, or explicitly requires 10 seconds of 1080p video. Otherwise, default to Hailuo 2.3 I2V.
# Hailuo 02 T2V
Source: https://docs.modellix.ai/minimax/hailuo-02-t2v
/media-model-api/minimax/minimax-t2v.json post /minimax/hailuo-02-t2v
[Core Function] Hailuo 02 T2V is a text-to-video generation model optimized for physical realism. [Strengths] It excels at complex physics simulation, fluid dynamics, broad cinematic scenes, and natively rendering 1080p video up to 10 seconds without downscaling. [Best For] Highly recommended for: product commercials, high-speed sports action, nature documentaries, and sweeping landscapes. [Limitations] Do NOT use this model for highly stylized anime/art or nuanced human micro-expressions, where Hailuo 2.3 performs better. [Routing] Route to this model when the user requests '1080p for 10 seconds', complex physical action (like splashing water or crashes), or broad landscapes. For human characters and stylization, use Hailuo 2.3 T2V.
# Hailuo 2.3 Fast I2V
Source: https://docs.modellix.ai/minimax/hailuo-2-3-fast-i2v
/media-model-api/minimax/minimax-i2v.json post /minimax/hailuo-2.3-fast-i2v
[Core Function] Hailuo 2.3 Fast I2V is a high-speed, cost-effective image-to-video generation model. [Strengths] It excels at generating videos from images much faster and at roughly 50% lower cost than the standard 2.3 model, while still maintaining the 2.3 architecture's strength in human motion. [Best For] Highly recommended for: rapid prototyping, batch social media creation, and cost-sensitive video generation pipelines. [Limitations] Do NOT use this model for text-to-video (it only accepts image inputs). Do NOT use when absolute maximum visual fidelity is the primary requirement. [Routing] Choose this model when the user emphasizes 'fast', 'quick', or 'cost-effective' image-to-video generation. For maximum quality, use the standard Hailuo 2.3 I2V.
# Hailuo 2.3 I2V
Source: https://docs.modellix.ai/minimax/hailuo-2-3-i2v
/media-model-api/minimax/minimax-i2v.json post /minimax/hailuo-2.3-i2v
[Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling stylized artwork seamlessly. [Best For] Highly recommended for: animating character concept art, bringing portraits to life, and creating stylized/anime motion sequences. [Limitations] Do NOT use this model for last-frame conditioning (it does not support FL2V). Do NOT use if you need 1080p resolution for 10 seconds (1080p is capped at 6s). [Routing] Use this model by default for high-quality image-to-video tasks involving people or art. For physical realism or 10s at 1080p, route to Hailuo 02 I2V. For cost-effective/faster generation, route to Hailuo 2.3 Fast I2V.
# Hailuo 2.3 T2V
Source: https://docs.modellix.ai/minimax/hailuo-2-3-t2v
/media-model-api/minimax/minimax-t2v.json post /minimax/hailuo-2.3-t2v
[Core Function] Hailuo 2.3 T2V is a flagship text-to-video generation model optimized for human performance and stylization. [Strengths] It excels at capturing intricate human motion, nuanced facial micro-expressions, prompt adherence, and applying highly stylized aesthetics (e.g., anime, ink wash, game CG) to video. [Best For] Highly recommended for: character-driven storytelling, close-up emotional shots, stylized artistic videos, and dialogue scenes. [Limitations] Do NOT use this model if you need native 1080p resolution for 10 full seconds (1080p is capped at 6 seconds; generating 10s forces 768p resolution). [Routing] Use this model by default for text-to-video requests involving humans, faces, or specific art styles. If the user requires strict physical realism/world dynamics or native 1080p for 10 seconds, route to Hailuo 02 T2V instead.
# MiniMax H3 FL2V
Source: https://docs.modellix.ai/minimax/minimax-h3-fl2v
/media-model-api/minimax/minimax-i2v.json post /minimax/minimax-h3-fl2v
[Core Function] MiniMax H3 FL2V generates video guided by a first frame, a last frame, or both. [Strengths] It supports first-only, last-only, and first-plus-last conditioning with 4-15 second duration and 768P or 2K output. Output framing follows the input frame imagery. [Best For] Highly recommended for: start-frame animation, end-frame targeting, before-and-after transitions, and storyboard frame bridging. [Limitations] Do NOT use this model for pure text-to-video, reference-image subject remix, or reference-video remix. Provide prompt, duration, resolution, and at least one of first_frame_image or last_frame_image. Do NOT send reference_images, reference_videos, or reference_audios on this endpoint. [Routing] Choose MiniMax H3 FL2V when the user supplies a start and/or end frame. For reference images without frame roles, use MiniMax H3 I2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.
# MiniMax H3 I2V
Source: https://docs.modellix.ai/minimax/minimax-h3-i2v
/media-model-api/minimax/minimax-i2v.json post /minimax/minimax-h3-i2v
[Core Function] MiniMax H3 I2V generates video from reference images plus a text prompt. [Strengths] It accepts up to 9 reference images and optional reference audios, with 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: character or style consistency from stills, product look references, multi-image subject guidance, and prompt-driven scenes featuring a referenced subject. [Limitations] Do NOT use this model for first or last frame transitions or when the primary input is a reference video. prompt, duration, resolution, and reference_images are required. Do NOT send first_frame_image, last_frame_image, or reference_videos on this endpoint. [Routing] Choose MiniMax H3 I2V when the user provides reference stills. For start or end frames, use MiniMax H3 FL2V. For reference videos, use MiniMax H3 V2V. For text only, use MiniMax H3 T2V.
# MiniMax H3 T2V
Source: https://docs.modellix.ai/minimax/minimax-h3-t2v
/media-model-api/minimax/minimax-t2v.json post /minimax/minimax-h3-t2v
[Core Function] MiniMax H3 T2V is a text-to-video generation model that creates video from a text prompt only. [Strengths] It supports 4-15 second clips, 768P or 2K output, and concrete aspect ratios from cinematic ultrawide to vertical. [Best For] Highly recommended for: prompt-only storyboards, character-driven shorts, cinematic B-roll from text, and high-resolution drafts without image inputs. [Limitations] Do NOT use this model if you need to condition on images, first or last frames, or reference videos. prompt, duration, resolution, and ratio are required; ratio must be one of the documented aspect ratios. [Routing] Choose MiniMax H3 T2V for text-only MiniMax H3 video. If the user provides a start or end frame, use MiniMax H3 FL2V. If they provide reference images, use MiniMax H3 I2V. If they provide reference videos, use MiniMax H3 V2V.
# MiniMax H3 V2V
Source: https://docs.modellix.ai/minimax/minimax-h3-v2v
/media-model-api/minimax/minimax-v2v.json post /minimax/minimax-h3-v2v
[Core Function] MiniMax H3 V2V generates video guided by one or more reference videos plus a text prompt. [Strengths] It accepts up to 3 reference videos with optional reference images and audios, 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: motion remix, style transfer from video clips, keeping subject motion while changing the scene description, and multi-clip reference guidance. [Limitations] Do NOT use this model for text-only generation or first or last frame transitions. prompt, duration, resolution, and reference_videos are required. Do NOT send first_frame_image or last_frame_image on this endpoint. [Routing] Choose MiniMax H3 V2V when the user provides reference videos. For reference stills only, use MiniMax H3 I2V. For start or end frames, use MiniMax H3 FL2V. For text only, use MiniMax H3 T2V.
# MiniMax Voice Clone
Source: https://docs.modellix.ai/minimax/minimax-voice-clone
/media-model-api/minimax/minimax-s2s.json post /minimax/minimax-voice-clone
[Core Function] MiniMax Voice Clone clones a speaker from a public reference audio URL and synthesizes new speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; optional language_boost for clone and synthesis; prosody controls (speed, volume, pitch, emotion); pronunciation overrides; voice effects; flexible audio formats. [Best For] One-off cloned narration, personalized prompts, and demos where a lasting voice library is not needed. [Limitations] Do NOT use this to obtain a reusable voice library entry; the cloned voice is temporary and is not returned. audio_url must be publicly reachable (mp3/m4a/wav, 10s–5min, ≤20 MB). [Routing] Choose speech-2.8-hd for higher quality; speech-2.8-turbo for lower latency or cost. Paralinguistic tags such as (laughs) in text are supported on both speech-2.8-hd and speech-2.8-turbo.
# Speech 2.8 HD
Source: https://docs.modellix.ai/minimax/speech-2-8-hd
/media-model-api/minimax/minimax-t2s.json post /minimax/speech-2.8-hd
[Core Function] MiniMax Speech 2.8 HD is a high-quality text-to-speech model that converts text into natural spoken audio, including expressive paralinguistic cues such as (laughs) and (sighs). [Strengths] Strong narration quality, stable prosody controls (speed, volume, pitch, emotion), optional timbre mixing, pronunciation overrides, subtitles, and flexible audio formats. [Best For] Highly recommended for: brand voiceovers, audiobook or long-form narration, marketing clips, multilingual delivery with language_boost, and expressive character speech. [Limitations] Do NOT use this for streaming or real-time partial audio. Do NOT request hex output or emotion whisper. Do NOT send both voice_id and timbre_weights. Audio URLs expire in about 24 hours. [Routing] Choose speech-2.8-hd when quality or expressive delivery matters most. Choose speech-2.8-turbo when the user prioritizes lower latency or cost.
# Speech 2.8 Turbo
Source: https://docs.modellix.ai/minimax/speech-2-8-turbo
/media-model-api/minimax/minimax-t2s.json post /minimax/speech-2.8-turbo
[Core Function] MiniMax Speech 2.8 Turbo is a lower-latency text-to-speech model with the same control surface as Speech 2.8 HD, including paralinguistic tags such as (laughs). [Strengths] Faster and more cost-efficient synthesis while retaining prosody, timbre mix, pronunciation, subtitle, and audio-format controls. [Best For] Highly recommended for: interactive assistants, high-volume TTS batches, cost-sensitive voiceovers, quick narration drafts, and latency-sensitive product prompts. [Limitations] Do NOT use this for streaming or real-time partial audio. Do NOT request hex output or emotion whisper. Do NOT send both voice_id and timbre_weights. Prefer speech-2.8-hd if maximum audio quality is required. Audio URLs expire in about 24 hours. [Routing] Choose speech-2.8-turbo when the user emphasizes speed, cost, or throughput. Choose speech-2.8-hd for premium quality or highly expressive delivery.
# GPT Image 1.5
Source: https://docs.modellix.ai/openai/gpt-image-1-5
/media-model-api/openai/openai-t2i.json post /openai/gpt-image-1.5
[Core Function] GPT Image 1.5 is a versatile text-to-image generation model. [Strengths] It balances solid visual performance with crucial utility features, notably its native support for generating images with transparent backgrounds. [Best For] Highly recommended for: creating UI icons, standalone logos, game assets, and any graphic design elements that require a transparent background. [Limitations] Do NOT use this model if you need 2K or 4K resolution. Its maximum supported resolution is 1536x1024. [Routing] Choose this model specifically when the user asks for 'transparent background', 'no background', or 'PNG icon'. For standard, high-fidelity, or 4K image generation, use GPT Image 2 instead.
# GPT Image 1.5 Edit
Source: https://docs.modellix.ai/openai/gpt-image-1-5-edit
/media-model-api/openai/openai-i2i.json post /openai/gpt-image-1.5-edit
[Core Function] GPT Image 1.5 Edit is a versatile image-to-image editing and merging model. [Strengths] It excels at complex utility editing tasks, including multi-image merging (up to 16 images), transparent background support, and precise control over how strictly the model adheres to the input image (fidelity control). [Best For] Highly recommended for: merging reference images, editing UI assets, generating variations with strict shape preservation, and creating transparent cutouts. [Limitations] Do NOT use this model if you require 2K or 4K high-resolution outputs, as it is limited to standard resolutions. [Routing] Use this model specifically when the user provides multiple images to combine, requires transparency, or explicitly asks to 'keep the exact shape' of the original image (fidelity control). Otherwise, use GPT Image 2 Edit.
# GPT Image 2
Source: https://docs.modellix.ai/openai/gpt-image-2
/media-model-api/openai/openai-t2i.json post /openai/gpt-image-2
[Core Function] GPT Image 2 is a state-of-the-art text-to-image generation model. [Strengths] It excels at generating highly detailed, photorealistic images from text descriptions, with native support for ultra-high resolutions including 2K and 4K (up to 3840x2160). [Best For] Highly recommended for: cinematic landscapes, detailed character portraits, high-end commercial concept art, and any scenario requiring maximum resolution and visual fidelity. [Limitations] Do NOT use this model if you need a transparent background (e.g., for icons or UI assets), as it does not support the `background: transparent` parameter. [Routing] Use this model by default for all high-quality image generation requests. If the user explicitly asks for an image with a transparent background, route to GPT Image 1.5 instead.
# GPT Image 2 Edit
Source: https://docs.modellix.ai/openai/gpt-image-2-edit
/media-model-api/openai/openai-i2i.json post /openai/gpt-image-2-edit
[Core Function] GPT Image 2 Edit is a high-resolution image-to-image editing model. [Strengths] It excels at making high-fidelity edits and style transformations to a single source image based on a text prompt, preserving details at up to 4K resolutions. [Best For] Highly recommended for: professional photo retouching, upscaling style transfers, and modifying high-resolution concept art. [Limitations] Do NOT use this model for multi-image merging (it only accepts one input image). Do NOT use if you need precise input fidelity control or transparent backgrounds. [Routing] Use this model by default when the user wants to edit a single image and prioritize output resolution/quality. If they need to merge multiple images or control the strictness of the edit (fidelity), use GPT Image 1.5 Edit.
# Whisper 1
Source: https://docs.modellix.ai/openai/whisper-1
/media-model-api/openai/openai-s2t.json post /openai/whisper-1
[Core Function] OpenAI Whisper transcribes a single public audio URL into text via an async task. [Strengths] Multiple output formats (verbose_json with word/segment timestamps, plain text, SRT, VTT), optional language and prompt biasing. [Best For] Meeting notes, podcasts, captions, and batch audio-to-text. [Limitations] Do NOT use file upload; URL-only input (audio up to 25 MB). Prefer verbose_json when you need duration or timestamps. [Routing] Use whisper-1 for OpenAI Whisper quality with URL-based audio.
# C1 FL2V
Source: https://docs.modellix.ai/pixverse/c1-fl2v
/media-model-api/pixverse/pixverse-i2v.json post /pixverse/c1-fl2v
[Core Function] PixVerse c1 first-last-frame generates a video that transitions from a start frame to an end frame, guided by a prompt. [Strengths] Controlled start/end composition with smooth interpolation. [Best For] Morphs, scene transitions, before/after motion. [Limitations] Do NOT use this if you only have a single image (use Image-to-Video) or lack an end frame; it requires exactly two images (a first and a last frame) and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides a start image AND an end image. For one image use Image-to-Video. The v6 and c1 first-last-frame variants accept identical parameters; choose the version the user requests.
# C1 I2V
Source: https://docs.modellix.ai/pixverse/c1-i2v
/media-model-api/pixverse/pixverse-i2v.json post /pixverse/c1-i2v
[Core Function] PixVerse c1 I2V animates a single starting image into a video guided by a text prompt. [Strengths] Smooth, prompt-guided motion from one frame; optional audio. [Best For] Bringing a photo/illustration to life, product showcases, quick cinematic motion from a still. [Limitations] Do NOT use this for two-frame transitions (use First-Last-Frame), multi-subject fusion (use Reference-to-Video), or text-only generation (use Text-to-Video). It requires exactly one starting image, does not support multi-clip, and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides one starting image. The c1 variant matches v6 inputs but does not support multi-clip generation; choose c1-i2v when the user requests the c1 model and multi-clip is not needed.
# C1 R2V
Source: https://docs.modellix.ai/pixverse/c1-r2v
/media-model-api/pixverse/pixverse-i2v.json post /pixverse/c1-r2v
[Core Function] PixVerse c1 reference-to-video (fusion) generates a video from a prompt while preserving subjects from 1-7 reference images; each reference can be tagged as subject/background and named for @-reference in the prompt. [Strengths] Precise multi-subject composition and identity preservation. [Best For] Putting specific characters/objects into a new scene, multi-character interactions. [Limitations] Do NOT use this for single-image animation (use Image-to-Video), two-frame transitions (use First-Last-Frame), or text-only generation (use Text-to-Video). It requires 1-7 reference images and a required aspect_ratio, and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides one or more reference images to be fused into the video, especially when distinguishing subject vs background or naming references for the prompt. The v6 and c1 reference-to-video variants accept identical parameters; choose the version the user requests.
# C1 T2V
Source: https://docs.modellix.ai/pixverse/c1-t2v
/media-model-api/pixverse/pixverse-t2v.json post /pixverse/c1-t2v
[Core Function] PixVerse c1 T2V generates a video purely from a text prompt, with no input image. [Strengths] Strong prompt adherence and smooth motion; optional audio. [Best For] Turning an idea or script into video, concept visualization, story beats, social clips from text. [Limitations] Do NOT use this when the user provides a starting image, two frames, or reference images; route to Image-to-Video, First-Last-Frame, or Reference-to-Video instead. c1 does not support multi-clip. [Routing] Use for text-only generation. The c1 variant takes the same inputs as v6 except it does not support multi-clip generation; choose c1-t2v when the user requests the c1 model and multi-clip is not needed.
# Lipsync
Source: https://docs.modellix.ai/pixverse/lipsync
/media-model-api/pixverse/pixverse-v2v.json post /pixverse/lipsync
[Core Function] PixVerse Lip Sync drives a talking video so the subject's lips match given audio or text-to-speech. [Strengths] Accurate lip synchronization for talking-head videos; supports either an existing audio track or TTS from a chosen speaker voice. [Best For] Dubbing, virtual presenters, character dialogue, localizing spoken video. [Limitations] Do NOT use this to generate new motion or change content; it only re-times the subject's lips on an existing video. It requires an input video, and you must provide EITHER audio_url OR (speaker_id + tts_content), not both. [Routing] Use when the user has a video and wants the speaker's lips to match speech. Use audio_url for an existing voice track; use speaker_id (a named voice code) + tts_content (max 140 chars) to synthesize speech from text.
# Motion Control
Source: https://docs.modellix.ai/pixverse/motion-control
/media-model-api/pixverse/pixverse-v2v.json post /pixverse/motion-control
[Core Function] PixVerse Motion Control (Mimic) animates a subject image so it follows the motion of a reference video. [Strengths] Transfers human/animal motion from a driving video onto a still subject. [Best For] Making a character mimic a dance or action, motion retargeting onto a photo. [Limitations] Do NOT use this if you only have a video and no subject image (use Restyle or Extend instead), or if you need 1080p output (only 360p/540p/720p are supported). It requires BOTH a subject image (with a clear person or animal) AND a reference video (with a person as the primary focus). [Routing] Use when the user has one subject image and one motion reference video and wants the subject to mimic that motion.
# Upscale Video
Source: https://docs.modellix.ai/pixverse/upscale-video
/media-model-api/pixverse/pixverse-v2v.json post /pixverse/upscale-video
[Core Function] PixVerse Upscale increases the resolution and clarity of an existing video. [Strengths] Sharper detail and higher-resolution output without changing content. [Best For] Enhancing low-resolution footage, finalizing clips for delivery. [Limitations] Requires an input video; it enhances quality, it does NOT change content, style, or motion. [Routing] Use when the user wants to improve the resolution/quality of an existing video, not to generate, restyle, or extend content.
# V6 FL2V
Source: https://docs.modellix.ai/pixverse/v6-fl2v
/media-model-api/pixverse/pixverse-i2v.json post /pixverse/v6-fl2v
[Core Function] PixVerse v6 first-last-frame generates a video that transitions from a start frame to an end frame, guided by a prompt. [Strengths] Controlled start/end composition with smooth interpolation. [Best For] Morphs, scene transitions, before/after motion. [Limitations] Do NOT use this if you only have a single image (use Image-to-Video) or lack an end frame; it requires exactly two images (a first and a last frame) and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides a start image AND an end image. For one image use Image-to-Video. The v6 and c1 first-last-frame variants accept identical parameters; choose the version the user requests.
# V6 I2V
Source: https://docs.modellix.ai/pixverse/v6-i2v
/media-model-api/pixverse/pixverse-i2v.json post /pixverse/v6-i2v
[Core Function] PixVerse v6 I2V animates a single starting image into a video guided by a text prompt. [Strengths] Smooth, prompt-guided motion from one frame; optional audio and multi-clip. [Best For] Bringing a photo/illustration to life, product showcases, quick cinematic motion from a still. [Limitations] Do NOT use this for two-frame transitions (use First-Last-Frame), multi-subject fusion (use Reference-to-Video), or text-only generation (use Text-to-Video). It requires exactly one starting image and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides one starting image. The v6 and c1 variants take the same inputs except v6 also supports multi-clip generation (generate_multi_clip_switch); choose v6-i2v when the user wants multi-clip output or requests the v6 model.
# V6 R2V
Source: https://docs.modellix.ai/pixverse/v6-r2v
/media-model-api/pixverse/pixverse-i2v.json post /pixverse/v6-r2v
[Core Function] PixVerse v6 reference-to-video (fusion) generates a video from a prompt while preserving subjects from 1-7 reference images; each reference can be tagged as subject/background and named for @-reference in the prompt. [Strengths] Precise multi-subject composition and identity preservation. [Best For] Putting specific characters/objects into a new scene, multi-character interactions. [Limitations] Do NOT use this for single-image animation (use Image-to-Video), two-frame transitions (use First-Last-Frame), or text-only generation (use Text-to-Video). It requires 1-7 reference images and a required aspect_ratio, and supports up to 1080p and a maximum of 15s. [Routing] Use when the user provides one or more reference images to be fused into the video, especially when distinguishing subject vs background or naming references for the prompt. The v6 and c1 reference-to-video variants accept identical parameters; choose the version the user requests.
# V6 T2V
Source: https://docs.modellix.ai/pixverse/v6-t2v
/media-model-api/pixverse/pixverse-t2v.json post /pixverse/v6-t2v
[Core Function] PixVerse v6 T2V generates a video purely from a text prompt, with no input image. [Strengths] Strong prompt adherence and smooth motion; optional audio and multi-clip generation. [Best For] Turning an idea or script into video, concept visualization, story beats, social clips from text. [Limitations] Do NOT use this when the user provides a starting image, two frames, or reference images; route to Image-to-Video, First-Last-Frame, or Reference-to-Video instead. [Routing] Use for text-only generation. v6 additionally supports multi-clip generation (generate_multi_clip_switch), which c1 does not; choose v6-t2v when multi-clip output is needed or the v6 model is requested.
# V6 Video Extend
Source: https://docs.modellix.ai/pixverse/v6-video-extend
/media-model-api/pixverse/pixverse-v2v.json post /pixverse/v6-video-extend
[Core Function] PixVerse v6 Extend continues an existing video, generating additional seconds guided by a text prompt. [Strengths] Seamless continuation of the existing motion and scene. [Best For] Lengthening clips, continuing an action, adding an ending to footage. [Limitations] Do NOT use this to create a video from scratch (use Text-to-Video) or to restyle (use Restyle). It requires an input video; both duration and quality are required, with duration limited to 1-15 seconds added per call and output resolution up to 1080p. [Routing] Use when the user wants to make an existing video longer or continue its action.
# Video Restyle
Source: https://docs.modellix.ai/pixverse/video-restyle
/media-model-api/pixverse/pixverse-v2v.json post /pixverse/video-restyle
[Core Function] PixVerse Restyle re-renders an existing video into a new visual style. [Strengths] Consistent style transfer across all frames. [Best For] Turning footage into anime/3D/painterly looks, stylized remixes. [Limitations] Do NOT use this to change content, motion, or add new scenes; it only re-renders the visual style of an existing video. It requires an input video, and you must provide EITHER restyle_id (a preset style code from the PixVerse restyle list) OR restyle_prompt (free-text style, max 2048 chars), not both. [Routing] Use when the user wants to change the look of an existing video. Use restyle_id for an official preset, restyle_prompt for a custom style.
# Segmented Camera Motion
Source: https://docs.modellix.ai/skyreels/segmented-camera-motion
/media-model-api/skywork/skyreels-i2v.json post /skywork/segmented-camera-motion
**[Core Function]** SkyReels Segmented Camera Motion (audio-to-video) generates a talking-avatar video with directed camera movement across time segments. **[Strengths]** Combines an audio-driven avatar with per-segment camera trajectories such as push, pan, crane and rotation. **[Best For]** Dynamic presenter clips and cinematic avatar shots with controlled camera motion. **[Limitations]** Do NOT use this when you need a completely static camera (use single-actor-avatar). It requires first_frame_image and one audio segment; use camera_control_pro for multi-segment or compound moves. mode=std outputs 720p, mode=pro outputs 1080p. **[Routing]** Set a single traj_type plus camera_control_strength for a simple move, or supply camera_control_pro (a list of per-segment {start_time, end_time, traj_type, ...}) for compound motion; choose mode=pro for 1080p.
# Single Actor Avatar
Source: https://docs.modellix.ai/skyreels/single-actor-avatar
/media-model-api/skywork/skyreels-i2v.json post /skywork/single-actor-avatar
**[Core Function]** SkyReels Single-Actor Avatar (audio-to-video) drives a talking-avatar video from a single portrait image and one audio track. **[Strengths]** Lip-synced single-speaker talking-head video generated from an image plus audio. **[Best For]** Virtual presenters, single-speaker dubbing, and talking avatars. **[Limitations]** Do NOT use this for multi-speaker scenes (use the multi-actor flow) or when you only have text. It requires first_frame_image and exactly one audio segment (<=200s). mode=std outputs 720p, mode=pro outputs 1080p. **[Routing]** Provide a portrait first_frame_image and one audio URL in audios; choose mode=pro for 1080p output.
# Sky Lipsync
Source: https://docs.modellix.ai/skyreels/sky-lipsync
/media-model-api/skywork/skyreels-v2v.json post /skywork/sky-lipsync
**[Core Function]** SkyReels Lip Sync (retalking) re-drives a talking video so the subject's lips match a given audio track. **[Strengths]** Accurate lip re-synchronization on an existing talking-head video. **[Best For]** Dubbing, re-voicing talking-head video, and localizing spoken video. **[Limitations]** Do NOT use this to generate new motion or content from scratch; it only re-times lips on an existing video. Requires video_url and audio_url. Output resolution is fixed at 720p. **[Routing]** Provide the source video_url and the target audio_url; optionally provide reference_char_url to guide the driven face.
# SkyReels I2V
Source: https://docs.modellix.ai/skyreels/skyreels-i2v
/media-model-api/skywork/skyreels-i2v.json post /skywork/skyreels-i2v
**[Core Function]** SkyReels Image-to-Video animates one or more keyframe images into a video guided by a text prompt. **[Strengths]** Supports a start frame, an end frame, and tagged mid-frames for keyframe control; optional audio, 480p/720p/1080p output, and fast/std modes. **[Best For]** Bringing a photo to life, first-last-frame transitions, and keyframe-driven storyboards. **[Limitations]** Do NOT use this for pure text-to-video (use skyreels-t2v) or for editing an existing video (use the Omni / video-to-video models). It requires at least one of first_frame_image, end_frame_image, or mid_frame_images; output is capped at 1080p and 15s, and fast mode supports only sound=false. **[Routing]** Provide first_frame_image to animate from a start image, add end_frame_image for a transition, or supply mid_frame_images (each tag must appear in the prompt as @tag) for keyframe guidance.
# SkyReels Omni
Source: https://docs.modellix.ai/skyreels/skyreels-omni
/media-model-api/skywork/skyreels-v2v.json post /skywork/skyreels-omni
**[Core Function]** SkyReels Omni is a reference-driven video model that generates or edits video using reference images (@image) and/or a reference video (@video), bound by tags in the prompt. **[Strengths]** A single endpoint covers motion reference, subject/background replacement, object insertion/removal, local editing, and video extension. **[Best For]** Video subject or background swap, motion transfer onto an image, object add/remove, local video edits, and extending a reference video. **[Limitations]** Do NOT use this for pure text-to-video (use skyreels-t2v) or simple single-image animation (use skyreels-i2v). Each ref tag must appear in the prompt as @tag; ref_videos supports only one video (<=15s); a reference-type video may combine only with image-type ref_images, while an extend-type video cannot combine with ref_images. When ref_videos is provided, aspect_ratio is ignored (output matches the video). **[Routing]** Provide ref_images (type grid or image) for image references and/or a single ref_videos entry (type reference for motion/edit, type extend for continuation); the reference tags must be used in the prompt.
# SkyReels R2V
Source: https://docs.modellix.ai/skyreels/skyreels-r2v
/media-model-api/skywork/skyreels-i2v.json post /skywork/skyreels-r2v
**[Core Function]** SkyReels Reference-to-Video (multiobject) generates a video from a prompt while preserving the subjects from 1-4 reference images. **[Strengths]** Multi-subject identity preservation, placing specific characters or objects into a newly generated scene. **[Best For]** Putting given characters/products into a new scene, multi-subject composition from reference photos. **[Limitations]** Do NOT use this to animate a single fixed frame (use skyreels-i2v) or for text-only generation (use skyreels-t2v). It requires 1-4 reference images and produces clips up to 5s. **[Routing]** Provide 1-4 subject reference images in ref_images; set aspect_ratio and duration (1-5s) as needed.
# SkyReels T2V
Source: https://docs.modellix.ai/skyreels/skyreels-t2v
/media-model-api/skywork/skyreels-t2v.json post /skywork/skyreels-t2v
**[Core Function]** SkyReels Text-to-Video generates a video purely from a text prompt, with no input media. **[Strengths]** Strong prompt adherence and smooth motion; supports optional audio, 480p/720p/1080p output, and a fast/std quality-speed trade-off. **[Best For]** Turning an idea or script into video, concept visualization, story beats, and social clips generated from text. **[Limitations]** Do NOT use this when the user provides an image, video, or audio input; route to Image-to-Video (skyreels-i2v), Reference-to-Video (skyreels-r2v), or the video-to-video / Omni models instead. Output is capped at 1080p and 15s per clip; fast mode currently supports only sound=false (no audio). **[Routing]** Use for text-only generation. Choose mode=fast for quicker results or mode=std for balanced quality; set resolution and aspect_ratio as needed.
# Video Extension Shot Switching
Source: https://docs.modellix.ai/skyreels/video-extension-shot-switching
/media-model-api/skywork/skyreels-v2v.json post /skywork/video-extension-shot-switching
**[Core Function]** SkyReels Shot-Switching Extension continues a video while transitioning to a new shot or camera angle. **[Strengths]** Cinematic shot transitions (cut-in, cut-out, reverse-shot, multi-angle, cut-away) when extending footage. **[Best For]** Adding a new shot after existing footage and cinematic transitions. **[Limitations]** Do NOT use this for a plain single-shot continuation (use video-extension-single-shot) or for generation from scratch. Requires a prefix_video (mp4 URL); adds 2-5s. **[Routing]** Provide prompt and prefix_video; choose cut_type for the transition style, or Auto to let the model decide.
# Video Extension Single Shot
Source: https://docs.modellix.ai/skyreels/video-extension-single-shot
/media-model-api/skywork/skyreels-v2v.json post /skywork/video-extension-single-shot
**[Core Function]** SkyReels Single-Shot Extension continues an existing single-shot video, generating additional seconds guided by a text prompt. **[Strengths]** Seamless single-shot continuation of the existing motion and scene. **[Best For]** Lengthening clips and continuing an action within one continuous shot. **[Limitations]** Do NOT use this to create a video from scratch (use skyreels-t2v) or to switch shots / add transitions (use video-extension-shot-switching). Requires a prefix_video (mp4 URL); adds 5-30s. **[Routing]** Provide prompt and prefix_video; set duration for how many seconds (5-30) to append.
# Video Restyling
Source: https://docs.modellix.ai/skyreels/video-restyling
/media-model-api/skywork/skyreels-v2v.json post /skywork/video-restyling
**[Core Function]** SkyReels Restyle re-renders an existing video into a preset visual style. **[Strengths]** Consistent style transfer across all frames into a chosen named art style. **[Best For]** Turning footage into simpsons, lego, paper-cutting, amigurumi, animal-crossing, van-gogh, or pixel-art looks. **[Limitations]** Do NOT use this to change content, motion, or add new scenes; it only restyles an existing video (input <=30s). Output resolution is fixed at 720p. **[Routing]** Provide the source video_url and a style_name from the supported list.
# Install the Web Tools MCP
Source: https://docs.modellix.ai/tools/mcp
Connect Cursor, VS Code, Claude Code, Codex, and other MCP clients to Modellix Web Search and Web Fetch at https://tool.modellix.ai/mcp.
Modellix Web Tools MCP is a remote Streamable HTTP server at `https://tool.modellix.ai/mcp`. It exposes Web Search and Web Fetch to MCP clients. Send `POST` requests only; the server does not use GET SSE and does not require `Mcp-Session-Id`.
This is not the [Docs](/ways-to-use/docs-search-mcp) MCP. The **Connect to
Cursor** / **Connect to VS Code** items in the page contextual menu install
`https://docs.modellix.ai/mcp`. Use the buttons and configs on this page for
Tools MCP.
## What You Get
After the client lists tools, you should see:
Tool name `modellix-ai/web-search`. Search the public web and return ranked
results. Request fields and pricing live on the API page.
Tool name `modellix-ai/web-fetch`. Extract readable content from public HTTP
or HTTPS URLs. Request fields and pricing live on the API page.
See [Tools Overview](/tools/overview) for SKUs and rates. Query history with [Get Tool Logs](/api/get-tool-logs).
## Prerequisites
Create a key in the [Modellix console](https://modellix.ai/console/api-key).
Keep it in the MCP client or a secrets store. Do not put it in frontend
code, public repos, or logs.
Every call requires `X-Mdlx-User-Id`: 8–128 characters of letters, digits,
`-`, or `_`. Use a stable business ID such as `user_12345678`. Do not use
nicknames, email, phone numbers, or a new random value per request.
Required headers:
| Header | Required | Value |
| ---------------- | -------- | ---------------------------------------------------------------- |
| `Authorization` | Yes | `Bearer ` |
| `X-Mdlx-User-Id` | Yes | Stable end-user ID |
| `Content-Type` | Yes | `application/json` (clients usually set this) |
| `Accept` | Yes | `application/json, text/event-stream` (clients usually set this) |
## Install
Replace `YOUR_MODELLIX_API_KEY` and `your-end-user-id` before you use the server.
One-click install embeds placeholders. Open the client MCP settings afterward
and replace the API Key and `X-Mdlx-User-Id`.
Install in Cursor
Or add the server by hand:
Press Command + Shift + P (
Ctrl + Shift + P on Windows), search
for **Open MCP settings**, then click **Add custom MCP**.
```json theme={null}
{
"mcpServers": {
"modellix-tools": {
"url": "https://tool.modellix.ai/mcp",
"headers": {
"Authorization": "Bearer YOUR_MODELLIX_API_KEY",
"X-Mdlx-User-Id": "your-end-user-id"
}
}
}
}
```
Install in VS Code
Or create `.vscode/mcp.json`:
```json theme={null}
{
"servers": {
"modellix-tools": {
"type": "http",
"url": "https://tool.modellix.ai/mcp",
"headers": {
"Authorization": "Bearer YOUR_MODELLIX_API_KEY",
"X-Mdlx-User-Id": "your-end-user-id"
}
}
}
}
```
```bash theme={null}
claude mcp add --transport http modellix-tools https://tool.modellix.ai/mcp \
--header "Authorization: Bearer YOUR_MODELLIX_API_KEY" \
--header "X-Mdlx-User-Id: your-end-user-id"
```
Confirm with `claude mcp list`. The entry needs `"type": "http"` if you edit `.mcp.json` by hand.
In `~/.codex/config.toml`:
```toml theme={null}
[mcp_servers.modellix-tools]
url = "https://tool.modellix.ai/mcp"
http_headers = { "Authorization" = "Bearer YOUR_MODELLIX_API_KEY", "X-Mdlx-User-Id" = "your-end-user-id" }
```
Run `codex mcp list` to confirm the server is registered.
Any MCP client that supports remote Streamable HTTP and custom headers can use this config. Field names vary by host; the URL and headers stay the same.
```json theme={null}
{
"mcpServers": {
"modellix-tools": {
"type": "http",
"url": "https://tool.modellix.ai/mcp",
"headers": {
"Authorization": "Bearer YOUR_MODELLIX_API_KEY",
"X-Mdlx-User-Id": "your-end-user-id"
}
}
}
}
```
## Verify
Reload the client, then ask which tools are available. You should see `modellix-ai/web-search` and `modellix-ai/web-fetch`.
If the client returns `401`, check the API Key. If it returns `400`, check `X-Mdlx-User-Id`. Keep `request_id` from the tool result when you contact support.
# Modellix Tools Overview
Source: https://docs.modellix.ai/tools/overview
Get started with Modellix Tools—Web Search and Web Fetch on https://tool.modellix.ai, including per-request and per-URL pricing.
Modellix Tools run at `https://tool.modellix.ai`. Use one Modellix API Key to search the public web or extract readable content from public URLs. Calls are synchronous.
Tools run on `https://tool.modellix.ai`. Image, video, and speech generation
use `https://api.modellix.ai`. LLM chat uses `https://llm.modellix.ai`.
## Available Tools
Search the public web and return ranked results. Choose depth (`lite`,
`standard`, or `rich`) and optionally filter by domain, time, topic, or
country.
Extract readable title and content from public HTTP or HTTPS URLs (up to 20
per request). Failed URLs are returned separately.
Query request history with [Get Tool Logs](/api/get-tool-logs).
## Pricing
Prices are in **USD**. Web Search is billed **per request**. Web Fetch is billed **per successful URL**.
### Web Search
**[Web Search](/api/web-search)** (`POST /v1/web-search`) is charged once per request according to `depth`. It is not billed per result or per URL.
| Depth | SKU | Price (USD / request) |
| ---------- | --------------------- | --------------------: |
| `lite` | `web-search.lite` | \$0.01 |
| `standard` | `web-search.standard` | \$0.015 |
| `rich` | `web-search.rich` | \$0.02 |
Default `depth` is `standard`.
### Web Fetch
**[Web Fetch](/api/web-fetch)** (`POST /v1/web-fetch`) is charged **\$0.002 per successful URL** (`web-fetch`). Failed URLs are not billed.
Availability and prices can change. Confirm live rates in the
[Modellix console](https://modellix.ai/console).
## Quick Start
Create a key in the [Modellix console](https://modellix.ai/console/api-key)
and send it as `Authorization: Bearer`.
POST to `https://tool.modellix.ai/v1/web-search` or
`https://tool.modellix.ai/v1/web-fetch`. Include the required
`X-Mdlx-User-Id` header (8–128 characters; letters, digits, `-`, `_`).
Use [Get Tool Logs](/api/get-tool-logs) to list Web Search and Web Fetch
requests for your team.
## API Reference
| Endpoint | Docs |
| --------------------- | ----------------------------------- |
| `POST /v1/web-search` | [Web Search](/api/web-search) |
| `POST /v1/web-fetch` | [Web Fetch](/api/web-fetch) |
| `GET /v1/logs` | [Get Tool Logs](/api/get-tool-logs) |
# Lip Sync
Source: https://docs.modellix.ai/vidu/lip-sync
/media-model-api/vidu/vidu-v2v.json post /vidu/lip-sync
[Core Function] Vidu Lip Sync is a video-to-video audio synchronization model. [Strengths] It excels at reanimating lip movements in an existing video to precisely match a new replacement audio track, while preserving the original face identity. [Best For] Highly recommended for: dubbing videos into different languages, correcting spoken dialogue post-production, and creating realistic digital avatars. [Limitations] Do NOT use this model if you need to change body movements or generate a new video from scratch. It requires a pre-existing video and a clear audio track. [Routing] Use this model specifically when the user wants to change what a person in a video is saying. If the user wants to transfer body movements, use Vidu Motion Sync instead.
# Motion Sync
Source: https://docs.modellix.ai/vidu/motion-sync
/media-model-api/vidu/vidu-v2v.json post /vidu/motion-sync
[Core Function] Vidu Motion Sync is a video-to-video motion transfer model. [Strengths] It excels at accurately extracting physical motion from a source video (e.g., a dancing person) and applying it to a target character image, preserving the target's identity. [Best For] Highly recommended for: creating dance videos with custom characters, transferring complex choreography, and replicating specific physical actions onto avatars. [Limitations] Do NOT use this model if you want to change what a character is saying (use Lip Sync). It requires both a reference video for motion and a target image for appearance. [Routing] Use this model specifically when the user wants to copy the body movements or actions from one video onto a different character.
# One Click AD Film
Source: https://docs.modellix.ai/vidu/one-click-ad-film
/media-model-api/vidu/vidu-i2v.json post /vidu/one-click-ad-film
[Core Function] Vidu One-Click AD-Film is an automated marketing video generation model. [Strengths] It excels at transforming 1 to 7 product or scene images into a polished, commercial-style advertisement video (10-60s) automatically. [Best For] Highly recommended for: e-commerce product showcases, social media ads, promotional reels, and quick marketing campaigns. [Limitations] Do NOT use this model for narrative storytelling or cinematic films; it is optimized specifically for commercial pacing and product emphasis. [Routing] Use this specifically when the user wants to generate an 'ad', 'commercial', or 'promotional video' from product photos.
# One Click General Film
Source: https://docs.modellix.ai/vidu/one-click-general-film
/media-model-api/vidu/vidu-i2v.json post /vidu/one-click-general-film
[Core Function] Vidu One-Click General Film is an automated cinematic film generation model. [Strengths] It excels at automatically stringing together 1 to 7 user-provided images into a cohesive, cinematic film (up to 180s) with appropriate transitions and pacing. [Best For] Highly recommended for: instant music videos, cinematic montages, automated travel vlogs, and turning photo albums into compelling short films. [Limitations] Do NOT use this model if the user needs precise, frame-by-frame control over camera movements or specific character actions in each shot. [Routing] Use this when the user wants an automated 'done-for-you' long video from a batch of images without manually prompting every single shot.
# One Click Trending Replicate
Source: https://docs.modellix.ai/vidu/one-click-trending-replicate
/media-model-api/vidu/vidu-v2v.json post /vidu/one-click-trending-replicate
[Core Function] Vidu One-Click Trending Replicate is a viral video style cloning model. [Strengths] It excels at analyzing a trending or viral reference video and recreating its specific visual style, transitions, and pacing using the user's own provided subject images. [Best For] Highly recommended for: participating in social media video trends, quickly cloning viral visual effects, and applying popular editing styles to personal photos. [Limitations] Do NOT use this model for original cinematic storytelling. It is strictly meant to mimic the style of a provided reference video. [Routing] Use this when the user explicitly provides a 'trend' or 'viral' video and wants to recreate that exact vibe or transition style with their own images.
# Template Story
Source: https://docs.modellix.ai/vidu/template-story
/media-model-api/vidu/vidu-i2v.json post /vidu/template-story
[Core Function] Vidu Template Story is a narrative video generation model. [Strengths] It excels at placing user-provided character images into predefined, structured narrative templates (like 'love_story' or 'monkey_king') to automatically generate a cohesive short film. [Best For] Highly recommended for: creating instant themed short stories, personalized entertainment videos, and viral social media narrative trends. [Limitations] Do NOT use this model if the user wants completely custom motion or a unique storyline not covered by the available templates. It restricts creativity to the predefined story structure. [Routing] Use this only when the user explicitly requests to put their characters into a specific story template.
# Vidu Q2 Pro Digital Human
Source: https://docs.modellix.ai/vidu/viduq2-pro-digital-human
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq2-pro-digital-human
[Core Function] Vidu Q2 Pro Digital Human is a premium portrait animation model. [Strengths] It excels at generating highly realistic, expressive digital humans from a single portrait image, featuring precise lip-sync to audio and natural facial micro-expressions. [Best For] Highly recommended for: professional virtual spokespersons, high-end educational videos, news anchoring, and realistic character animation. [Limitations] Do NOT use this model for complex full-body physical interactions or videos longer than 10 seconds. [Routing] Use this by default for 'talking head' or 'digital human' requests prioritizing realism over speed.
# Vidu Q2 Pro Multi Frame
Source: https://docs.modellix.ai/vidu/viduq2-pro-multi-frame
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq2-pro-multi-frame
[Core Function] Vidu Q2 Pro Multi-Frame is a sequence-based animation model. [Strengths] It excels at creating continuous, high-quality animation by interpolating through a provided sequence of keyframes (up to 9 images). [Best For] Highly recommended for: complex motion control, precise character posing sequences, and animating detailed storyboards where intermediate states must be strictly followed. [Limitations] Do NOT use this model if you only have one or two images (use standard I2V or FL2V instead). It requires a start image and a list of key images. [Routing] Use this exclusively when the user provides a sequence of 3 to 9 specific frames and wants them animated in order.
# Vidu Q2 Turbo Digital Human
Source: https://docs.modellix.ai/vidu/viduq2-turbo-digital-human
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq2-turbo-digital-human
[Core Function] Vidu Q2 Turbo Digital Human is a fast portrait animation model. [Strengths] It excels at quickly animating a static portrait image into a speaking or moving digital human, syncing lip movements to provided audio with low latency. [Best For] Highly recommended for: rapid generation of talking head videos, quick virtual presenters, and responsive interactive avatars. [Limitations] Do NOT use this model for complex full-body motion, multi-character interactions, or videos longer than 10 seconds. [Routing] Use this model when the user wants to make a portrait 'talk' quickly. For higher realism and better quality, route to Q2 Pro Digital Human.
# Vidu Q2 Turbo Multi Frame
Source: https://docs.modellix.ai/vidu/viduq2-turbo-multi-frame
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq2-turbo-multi-frame
Vidu Q2 Turbo multi-frame animation model. Animates through a sequence of up to 9 frames (1 start + up to 8 key images). Both `start_image` and `key_images` are required. Supports `resolution`.
# Vidu Q3 AD
Source: https://docs.modellix.ai/vidu/viduq3-ad
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq3-ad
[Core Function] Vidu Q3 AD is a keyframe-driven short-play (short drama) Image-to-Video model that turns a film-style script plus character, scene, and prop reference images into a complete multi-shot short video, automatically planning shots and compositing them in one pass. [Strengths] It excels at multi-shot narrative coherence and keeping referenced characters, scenes, and props consistent across shots, driven directly from keyframe reference images without manual shot-by-shot prompting. [Best For] Highly recommended for: short ad films and brand stories, scripted short dramas, multi-scene narrative clips generated from a script, and character-driven reels built from a small cast of reference assets. [Limitations] Do NOT use this model to simply animate a single image (use a standard image-to-video model such as Vidu Q3 Pro/Turbo I2V instead); it requires script_content in traditional screenplay format (scene + characters + dialogue) plus 1-14 reference assets each with an image_uri, outputs 1080p in 16:9 or 9:16, and does not expose per-shot duration or style controls (duration is auto-planned). Asset type must be one of character/scene/tool. [Routing] Choose Vidu Q3 AD when the user provides a script and reference images and wants an automatically directed multi-shot short drama from keyframes. For animating a single image into one continuous clip, route to viduq3-pro-i2v / viduq3-turbo-i2v; for the standard short-play flow with placement/quality/duration/style controls, route to Vidu Q3 Drama.
# Vidu Q3 Drama
Source: https://docs.modellix.ai/vidu/viduq3-drama
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq3-drama
[Core Function] Vidu Q3 Drama (Short Play) is a script-to-video model that turns a written script plus character, scene, and prop reference images into a complete multi-shot short drama, automatically planning the shots, transitions, and camera work in a single pass. [Strengths] It excels at multi-shot narrative coherence, automatic storyboarding and cinematography, and keeping the identity of referenced characters, scenes, and props consistent across every shot. [Best For] Highly recommended for: scripted short dramas and web-series episodes, narrative short-form ads and brand stories, rapid storyboard and pre-visualization, and character-driven reels built from a cast of reference assets. [Limitations] Do NOT use this model to simply animate a single image as-is (use a standard image-to-video model such as Vidu Q3 Pro instead); it requires a script and 1-14 reference assets, supports only 8-12 second clips at 1080p in 16:9 or 9:16, and is not intended for pixel-perfect single-product shots or complex multi-object physics. [Routing] Choose Vidu Q3 Drama when the user provides a script or narrative beats plus character/scene/prop references and wants an automatically directed multi-shot short play. If the user only wants to animate a single image or needs one continuous clip without scripted scene changes, route to a standard image-to-video model such as Vidu Q3 Pro.
# Vidu Q3 Mix R2V
Source: https://docs.modellix.ai/vidu/viduq3-mix-r2v
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq3-mix-r2v
[Core Function] Vidu Q3 Mix R2V is a mixed-style reference-to-video generation model. [Strengths] It excels at generating highly consistent character videos by synthesizing and blending features from multiple reference images (up to 7) based on a text prompt. [Best For] Highly recommended for: maintaining strict character consistency across different styles, generating videos of a specific subject in entirely new environments, and blending concepts from multiple reference images. [Limitations] Do NOT use this model if you just want to animate a single image as is (use standard I2V). This model focuses on extracting character/style features and generating new content. [Routing] Use this when the user provides reference images of a character/subject and wants a video of them doing a specific new action from a text prompt, prioritizing mixed-style consistency.
# Vidu Q3 Pro Fast I2V
Source: https://docs.modellix.ai/vidu/viduq3-pro-fast-i2v
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq3-pro-fast-i2v
[Core Function] Vidu Q3 Pro Fast I2V is a high-speed Image-to-Video generation model. [Strengths] It excels at generating smooth, physically accurate continuous motion from a single starting frame with extremely low latency. [Best For] Highly recommended for: fast prototyping, short dynamic product showcases, quick cinematic transitions, and scenarios where generation speed is prioritized over maximum detail. [Limitations] Do NOT use this model if the user requires 4K resolution, complex multi-character interactions, or highly stylized 2D anime deformations. It only supports 720p/1080p resolutions. [Routing] Choose this 'Fast' model when the user emphasizes 'quick', 'fast', or needs immediate results. If the user demands ultimate cinematic quality, choose the standard Q3 Pro I2V model instead.
# Vidu Q3 Pro FL2V
Source: https://docs.modellix.ai/vidu/viduq3-pro-fl2v
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq3-pro-fl2v
[Core Function] Vidu Q3 Pro FL2V is a premium First-Last frame transition video model. [Strengths] It excels at generating highly detailed, cinematic, and logically consistent video transitions between a starting image and an ending image. [Best For] Highly recommended for: high-end commercial transitions, complex subject morphing, professional time-lapse effects, and cinematic storyboard completion. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. [Routing] Use this by default when the user provides exactly two images (start and end) and wants a video bridging them. For faster but lower-quality results, use Q3 Turbo FL2V.
# Vidu Q3 Pro I2V
Source: https://docs.modellix.ai/vidu/viduq3-pro-i2v
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq3-pro-i2v
[Core Function] Vidu Q3 Pro I2V is a premium Image-to-Video generation model. [Strengths] It excels at transforming a single starting image into high-fidelity, cinematic video with stable character consistency, complex motion, and synchronized audio-visual capabilities. [Best For] Highly recommended for: bringing concept art to life, professional film production, high-end commercial showcases, and creating immersive environments from still images. [Limitations] Do NOT use this model if you need instant/real-time generation, as rendering takes longer. It does not support 4K resolution. [Routing] Use this model by default for high-quality image-to-video requests. If the user requires faster generation, route to Q3 Pro Fast or Q3 Turbo.
# Vidu Q3 Pro T2V
Source: https://docs.modellix.ai/vidu/viduq3-pro-t2v
/media-model-api/vidu/vidu-t2v.json post /vidu/viduq3-pro-t2v
[Core Function] Vidu Q3 Pro T2V is a premium cinematic text-to-video generation model. [Strengths] It excels at generating top-tier, realistic videos from text with support for advanced multi-shot 'smart cuts', complex physics, and simultaneous audio-visual generation. [Best For] Highly recommended for: cinematic storytelling, professional advertising, short films, and high-fidelity concept visualizations. [Limitations] Do NOT use this model if the user is looking for an instant, low-latency preview, as generation takes longer. It does not support automatic BGM addition. [Routing] Use this model by default for high-quality text-to-video requests. If the user specifically asks for 'fast' or 'quick' generation, switch to the Q3 Turbo T2V model.
# Vidu Q3 R2V
Source: https://docs.modellix.ai/vidu/viduq3-r2v
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq3-r2v
[Core Function] Vidu Q3 R2V is a high-quality reference-to-video generation model. [Strengths] It excels at generating detailed, cinematic videos that precisely follow a text prompt while highly preserving the character identity from provided reference images. [Best For] Highly recommended for: professional character-driven storytelling, high-fidelity avatar generation in new scenes, and cinematic films requiring consistent actors. [Limitations] Do NOT use this model if you just want to add motion to an existing image (use I2V). This model creates new scenes based on the prompt while keeping the character. [Routing] Use this by default when the user wants to generate a video of a specific character (provided via image) doing something new (provided via text prompt).
# Vidu Q3 Turbo FL2V
Source: https://docs.modellix.ai/vidu/viduq3-turbo-fl2v
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq3-turbo-fl2v
[Core Function] Vidu Q3 Turbo FL2V is a fast First-Last frame transition video model. [Strengths] It excels at rapidly generating a smooth video transition bridging a specific starting image and an ending image. [Best For] Highly recommended for: quick visual morphs, before-and-after transitions, time-lapse simulations, and rapid storyboard filling. [Limitations] Do NOT use this model with only a single image; both a start and end frame are strictly required. Do NOT use if you need the highest possible cinematic detail. [Routing] Use this when the user provides exactly two images and wants a fast transition between them. For higher quality transitions, use Q3 Pro FL2V.
# Vidu Q3 Turbo I2V
Source: https://docs.modellix.ai/vidu/viduq3-turbo-i2v
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq3-turbo-i2v
[Core Function] Vidu Q3 Turbo I2V is a fast Image-to-Video generation model. [Strengths] It excels at quickly animating a starting frame into a video sequence with strong motion dynamics and low latency. [Best For] Highly recommended for: rapid social media content creation, quick animatics, and fast visual iterations from reference images. [Limitations] Do NOT use this model if you require the absolute highest cinematic fidelity or complex audio-visual synchronization. [Routing] Choose this model for general quick image-to-video tasks. For the highest quality, route to Q3 Pro I2V.
# Vidu Q3 Turbo R2V
Source: https://docs.modellix.ai/vidu/viduq3-turbo-r2v
/media-model-api/vidu/vidu-i2v.json post /vidu/viduq3-turbo-r2v
[Core Function] Vidu Q3 Turbo R2V is a fast reference-to-video generation model. [Strengths] It excels at quickly generating dynamic videos based on a text prompt while preserving the identity of the subjects from provided reference images. [Best For] Highly recommended for: rapid character animation, quick conceptual mockups with specific subjects, and fast social media content featuring consistent characters. [Limitations] Do NOT use this model if you need the absolute highest cinematic quality or if you just want to animate an image directly without a text prompt. [Routing] Choose this model for fast generation of a character performing actions based on a text prompt. For better quality, use Q3 R2V or Q3 Mix R2V.
# Vidu Q3 Turbo T2V
Source: https://docs.modellix.ai/vidu/viduq3-turbo-t2v
/media-model-api/vidu/vidu-t2v.json post /vidu/viduq3-turbo-t2v
[Core Function] Vidu Q3 Turbo T2V is a fast text-to-video generation model. [Strengths] It excels at rapidly generating smooth, dynamic videos from text descriptions with very low latency. [Best For] Highly recommended for: fast prototyping, quick visual brainstorming, generating background b-roll, and scenarios where generation speed is prioritized. [Limitations] Do NOT use this model if you need ultimate cinematic quality, complex audio-visual synchronization, or multi-shot 'smart cuts'. It does not support automatic BGM addition. [Routing] Choose this 'Turbo' model when the user emphasizes 'quick', 'fast', or needs immediate results. If the user demands the highest cinematic quality or advanced audio-visual features, choose the Q3 Pro T2V model instead.
# Use Modellix Agent Canvas in Coding Agents
Source: https://docs.modellix.ai/ways-to-use/agent-canvas
Install the Modellix Agent Canvas local stdio MCP plugin in Codex, Cursor, Claude Code, or OpenCode for visual image work, HTML drafts, and presentations.
[Modellix Agent Canvas](https://github.com/Modellix/modellix-agent-canvas) is a local, workspace-bound `stdio` MCP plugin for visual AI work. It brings an Excalidraw infinite canvas, Modellix image generation and editing, paid-operation confirmation, durable task recovery, HTML drafts, presentations, and project-local persistence into one workspace.
Canvas runs on your machine. It does **not** require a deployed Canvas service. Hosts with [MCP Apps](https://modelcontextprotocol.io/) can embed the full canvas; other compatible hosts open the same server through a short-lived loopback page.
Agent Canvas is separate from the [Modellix Plugin](/ways-to-use/plugin), which teaches agents to run media tasks through `modellix-cli` or the REST API. Use Agent Canvas when you want a visual workspace; use the Plugin or [Skill](/ways-to-use/skill) when you want CLI-first image, video, and speech generation.
## What You Get
* **Infinite canvas** — text, shapes, lines, arrows, freehand drawing, frames, images, grouping, locking, layers, alignment, and undo/redo
* **Multi-page projects** — create, rename, duplicate, reorder, and delete pages with independent viewports
* **Image placeholders** — reserve a target area before generation; results replace the placeholder and stay undoable
* **Image generation and editing** — text-to-image, single-image edit, ordered multi-image references, transparent backgrounds, and 1–4 outputs
* **Paid-operation safety** — prepare is free; submit needs explicit one-time confirmation; unknown submissions are never retried automatically
* **Durable tasks** — task IDs and local results persist so work can continue after the host or browser closes
* **HTML drafts and presentations** — sandboxed HTML preview, slide layouts, presentation mode, and PNG-sequence export
* **Local credentials** — API keys use the `modellix-cli` system credential store and are never written to chat, MCP arguments, URLs, or project files
Source and release notes: [GitHub](https://github.com/Modellix/modellix-agent-canvas) · npm [`@modellix/agent-canvas`](https://www.npmjs.com/package/@modellix/agent-canvas)
## Requirements
* Node.js `^20.19.0` or `>=22.12.0`, with npm on `PATH`
* Network access to `https://api.modellix.ai` and `https://registry.npmjs.org`
* A Modellix API key from the [Modellix Console](https://www.modellix.ai/console/api-key)
Install from the host once. Codex, Cursor, and Claude Code load the plugin from Git or Marketplace and resolve the pinned npm runtime in the background. OpenCode and generic MCP hosts add the npm-backed MCP once. You do **not** run a second global CLI install.
Verify the runtime on any host:
```bash theme={null}
npx -y --package @modellix/agent-canvas modellix-agent-canvas --doctor
```
## Install
Choose one path for your host.
```bash theme={null}
codex plugin marketplace add Modellix/modellix-agent-canvas
codex plugin add modellix-agent-canvas@modellix
```
You can also install **Modellix Agent Canvas** from `/plugins` or the desktop Plugins directory. Start a new task after installation so the session loads the skills and MCP tools.
On first use, the Codex adapter caches the pinned npm runtime in a user-local directory. Warm starts reuse that cache without a persistent `npx` wrapper.
In Cursor 2.6+, run:
```text theme={null}
/add-plugin modellix-agent-canvas
```
For a GitHub or local checkout, open **Customize → Plugins → + Add** and select the repository root. Cursor reads `.cursor-plugin/marketplace.json` and offers **Modellix Agent Canvas** from the `modellix` personal marketplace.
For a direct MCP setup, add the repository `mcp.json` (or an equivalent `stdio` entry that runs `modellix-agent-canvas` with `--host cursor --supports-mcp-apps true`). Cursor supplies the active workspace through MCP Roots—do **not** pass a literal `${workspaceFolder}` argument.
Reload Cursor and confirm `modellix-agent-canvas` is connected in MCP settings.
```bash theme={null}
claude plugin marketplace add Modellix/modellix-agent-canvas
claude plugin install modellix-agent-canvas@modellix
```
After enabling or upgrading, run `/reload-plugins`, then `/mcp` to verify the connection.
Merge the server from [`adapters/opencode/opencode.json`](https://github.com/Modellix/modellix-agent-canvas/blob/main/adapters/opencode/opencode.json) into your project `opencode.json`. OpenCode V2 beta users should merge [`adapters/opencode/opencode-v2.json`](https://github.com/Modellix/modellix-agent-canvas/blob/main/adapters/opencode/opencode-v2.json) instead.
Example for the stable OpenCode shape:
```json theme={null}
{
"$schema": "https://opencode.ai/config.json",
"mcp": {
"modellix-agent-canvas": {
"type": "local",
"command": [
"npx",
"-y",
"--package",
"@modellix/agent-canvas",
"modellix-agent-canvas",
"--host",
"opencode",
"--supports-mcp-apps",
"false"
],
"cwd": ".",
"enabled": true
}
}
}
```
Restart OpenCode after saving. Prefer pinning the exact npm version published in the [repository README](https://github.com/Modellix/modellix-agent-canvas) when you need a fixed runtime.
Configure a local `stdio` MCP server:
```json theme={null}
{
"command": "npx",
"args": [
"-y",
"--package",
"@modellix/agent-canvas",
"modellix-agent-canvas",
"--host",
"generic",
"--supports-mcp-apps",
"false",
"--project-dir",
"/absolute/path/to/project"
]
}
```
`--project-dir` must be an existing real absolute directory (not a symlink). One MCP process binds to one workspace.
## Host Compatibility
| Host | Local MCP | Canvas Surface |
| --------------------- | ----------------- | ------------------------------------ |
| Codex | `stdio` | MCP Apps widget, with local fallback |
| Cursor 2.6+ | `stdio` | MCP Apps |
| Claude Code | `stdio` | Short-lived local page |
| OpenCode | Local MCP command | Short-lived local page |
| Other stdio MCP hosts | Local MCP command | Short-lived local page |
Full protocol details live in the repository [host compatibility](https://github.com/Modellix/modellix-agent-canvas/blob/main/docs/host-compatibility.md) guide.
## First Use and API Key
After install, ask your agent to open Canvas for the active project:
```text theme={null}
get_modellix_canvas_status with refresh true and the absolute workspace path
open_modellix_canvas with the same absolute workspace path
```
Codex skills often supply `workspacePath` automatically. The path must be a real absolute directory, and one MCP session binds to one workspace.
If status is `missing` or `invalid`, Canvas shows a password field in its credential card. The form is an isolated, one-time loopback page that expires after five minutes. Submitting the key validates it through the bundled CLI, stores it in the system credential store, and refreshes status. The key never enters Canvas state or MCP tool arguments.
Credential resolution order:
1. Reuse a valid credential already saved by `modellix-cli`
2. Otherwise enter a key in the Canvas credential card
Never put an API key in chat, MCP config, command-line arguments, repository files, screenshots, project backups, or task reports. Canvas does not store keys in `localStorage`, `sessionStorage`, or IndexedDB.
## Typical Image Workflow
Paid image work follows a prepare → confirm → submit → finalize path:
1. **Prepare** — `prepare_modellix_image_task` resolves references, selects a model, and returns routing reason, effective specification, limitations, and estimated total cost. Prepare is free and does not create a paid task.
2. **Confirm** — review the model, quantity, warnings, and cost in the UI or agent reply.
3. **Submit** — after explicit approval, `submit_modellix_image_task` submits the unchanged intent with the same confirmation fingerprint.
4. **Poll** — `get_modellix_image_task` tracks registered tasks only. If status is `SUBMISSION_UNKNOWN`, query only—do not automatically resubmit.
5. **Finalize** — `finalize_modellix_image_task` downloads assets into the project and places them on the canvas.
Prompts and input images are sent to Modellix only after you confirm the paid task. Completed outputs are downloaded into project assets so you are not dependent on expiring remote URLs.
## Useful MCP Tools
Agent-facing tools include:
| Tool | Purpose |
| ------------------------------------------------------------------ | --------------------------------------------------------- |
| `get_modellix_canvas_status` | Check runtime and credential status |
| `start_modellix_api_key_setup` | Open the short-lived local key form |
| `open_modellix_canvas` | Open the canvas for a workspace |
| `get_canvas_context` | Read canvas context for the agent |
| `create_canvas_page` / `rename_canvas_page` / `delete_canvas_page` | Manage pages |
| `prepare_modellix_image_task` | Free prepare step with cost and routing |
| `submit_modellix_image_task` | Paid submit after confirmation |
| `get_modellix_image_task` | Poll a registered task |
| `list_modellix_canvas_tasks` | List local canvas tasks |
| `finalize_modellix_image_task` | Download and place results |
| `cleanup_modellix_canvas_uploads` | Clean terminal temporary uploads (`confirmCleanup: true`) |
## Project Data
Each bound workspace stores Canvas data under:
```text theme={null}
.modellix/canvas/
├── project.json
├── pages/
├── assets/
├── tasks/
├── recovery/
└── locks/
```
Uninstalling the plugin does **not** delete this directory or shared system credentials. Back up project data before you remove it. To remove a stored CLI profile, check `modellix-cli auth status --json`, then run `modellix-cli auth logout --profile ` only when you intend to drop that credential.
## Upgrade and Uninstall
| Host | Upgrade |
| ---------------------- | ------------------------------------------------------------------------------------------ |
| Codex | `codex plugin marketplace upgrade modellix`, then update `modellix-agent-canvas@modellix` |
| Claude Code | `claude plugin marketplace update modellix`, update from `/plugin`, then `/reload-plugins` |
| Cursor | Update from the plugin page, or bump the exact npm version in direct MCP config |
| OpenCode / generic MCP | Bump the exact npm package version and restart the host |
## Troubleshooting
| Problem | What to Do |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| MCP server does not connect | Run `npx -y --package @modellix/agent-canvas modellix-agent-canvas --doctor`, confirm Node.js version, then reload the host. |
| Canvas asks for an API key | Enter the key in the Canvas credential card, or reuse a valid `modellix-cli` profile. Do not put the key in MCP config. |
| `${workspaceFolder}` errors in Cursor | Remove that argument. Cursor binds the workspace through MCP Roots. |
| Paid task status is unknown | Query with `get_modellix_image_task` or check account activity. Do not automatically resubmit. |
| Wrong or empty workspace | Pass a real absolute project path as `workspacePath` / `--project-dir`. One MCP process binds to one workspace. |
## Next Steps
CLI-first media generation for coding agents.
Install the Modellix Agent Skill on its own.
Call Modellix media endpoints directly.
Source, host adapters, and release details.
# Use the Modellix REST API for Media Generation
Source: https://docs.modellix.ai/ways-to-use/api
Use the Modellix REST API to authenticate, upload media inputs, submit an asynchronous image, video, or audio task, poll its status, and retrieve output assets.
## Steps
Log in to the [Modellix console](https://modellix.ai/console).
In the Modellix console, go to "[API Key](https://modellix.ai/console/api-key)" and create an API Key.
The API Key is only displayed once after creation, so be sure to save it first.
Find the model you want to use in **Model API**, or [Model Index](https://docs.modellix.ai/llms.txt), and call the API using your API Key. For example:
```bash theme={null}
curl --request POST \
--url https://api.modellix.ai/api/v1/alibaba/qwen-image-plus/async \
--header 'Authorization: Bearer ' \
--header 'Content-Type: application/json' \
--data '
{
"prompt": "A cute cat playing in a garden on a sunny day"
}
'
```
Since model calls are asynchronous tasks, after a successful call, you will first receive a `task_id` as shown below:
```json theme={null}
{
"code": 0,
"message": "success",
"data": {
"status": "pending",
"task_id": "task-abc123",
"model_id": "qwen-image-plus",
"get_result": {
"method": "GET",
"url": "https://api.modellix.ai/api/v1/tasks/task-abc123"
}
}
}
```
You can query the task result later using the `task_id`. For example:
```bash theme={null}
curl --request GET \
--url https://api.modellix.ai/api/v1/tasks/{task_id} \
--header 'Authorization: Bearer '
```
After a successful query, you will get the generated result. For example:
```json theme={null}
{
"code": 0,
"message": "success",
"data": {
"status": "success",
"task_id": "task-abc123",
"model_id": "qwen-image-plus",
"duration": 3500,
"result": {
"resources": [
{
"url": "https://cdn.example.com/images/abc123.png",
"type": "image",
"width": 1024,
"height": 1024,
"format": "png",
"role": "primary"
}
],
"metadata": {
"image_count": 1,
"request_id": "req-123456"
},
"extensions": {
"submit_time": "2024-01-01T10:00:00Z",
"end_time": "2024-01-01T10:00:03Z"
}
}
}
}
```
The model's generated results will all be placed in the `result` object.
You can find the generated image URL in the `result`. For example:
```json theme={null}
{
"url": "https://cdn.example.com/images/abc123.png"
}
```
You have now successfully used the Modellix model API and obtained the generated result.
Generated results are only saved for 7 days, so please make sure to save them promptly.
## Request logs
List your team's media request history with:
```bash theme={null}
curl -sS "https://api.modellix.ai/api/v1/logs?start_time=1700000000&end_time=1700086400&page=1&page_size=20" \
-H "Authorization: Bearer ${API_KEY}"
```
`start_time` / `end_time` are UNIX seconds (span ≤ 30 days). Optional `mdlx_user_id` filters by the end-user id sent as `X-Mdlx-User-Id` on async inference. Full reference: [List media request logs](/api/get-logs).
For LLM request logs on `https://llm.modellix.ai`, see [LLM request logs](/llm/api/api#request-logs).
## User ID
Optional. Tag media async inference requests with your own end-user identifier so you can filter [request logs](#request-logs) later. Invalid values return `400`.
| Header | Rules |
| ---------------- | ---------------------------------------------------------------- |
| `X-Mdlx-User-Id` | Optional; length 8–128; ASCII letters, digits, `-`, and `_` only |
```bash theme={null}
curl --request POST \
--url https://api.modellix.ai/api/v1/alibaba/qwen-image-plus/async \
--header 'Authorization: Bearer ' \
--header 'Content-Type: application/json' \
--header 'X-Mdlx-User-Id: end_user_01' \
--data '
{
"prompt": "A cute cat playing in a garden on a sunny day"
}
'
```
The list response does **not** include a `mdlx_user_id` field; filter with the query parameter instead.
## Webhooks
Modellix supports webhooks to notify your application automatically when a media generation task is completed. Instead of polling the task status, you can configure a Webhook URL to receive the task results asynchronously.
### Triggering Webhooks
To enable webhooks for a task, include the `X-Webhook-URL` header when calling any prediction creation API:
```http theme={null}
X-Webhook-URL: https://example.com/webhook
```
Once the task reaches a terminal state, the system asynchronously sends a `POST` request to this URL. Terminal states include:
* `success`
* `failed`
* `canceled`
### Webhook URL Requirements
Your webhook endpoint must meet the following requirements:
* Must be a publicly accessible **HTTPS** address.
* Cannot be `localhost` or `127.0.0.1`.
* Cannot be a private IP address (e.g., `10.x.x.x`, `172.16.x.x` to `172.31.x.x`, `192.168.x.x`).
* Cannot contain username and password credentials in the URL.
For local development and testing, you can use tools like [ngrok](https://ngrok.com/) or [cloudflared](https://github.com/cloudflare/cloudflared) to expose your local server via a public HTTPS URL.
### Callback Request
The system will initiate a `POST` request to your configured `X-Webhook-URL` with the following headers:
| Header | Description |
| :----------------------- | :------------------------------------------------------ |
| `Content-Type` | Always `application/json` |
| `User-Agent` | Always `modellix-webhook/1.0` |
| `X-Modellix-Event` | The event type that triggered the webhook |
| `X-Modellix-Task-ID` | The unique ID of the generation task |
| `X-Modellix-Delivery-ID` | The unique ID of this webhook delivery attempt |
| `X-Modellix-Retry-Count` | The number of retries for this delivery (starts at `0`) |
#### Event Types
The `X-Modellix-Event` header can have one of the following values:
* `prediction.task.succeeded` — Sent when the task completes successfully.
* `prediction.task.failed` — Sent when the task fails.
* `prediction.task.canceled` — Sent when the task is canceled.
### Callback Payload
The webhook request body structure is identical to the response of the [Query Task Result](/api/get-task-result) API.
```json Success Example theme={null}
{
"code": 0,
"message": "success",
"data": {
"status": "success",
"task_id": "task_id",
"model_id": "provider/model",
"duration": 12345,
"result": {},
"billing": {
"status": "succeeded",
"amount": "12.3456"
}
}
}
```
```json Failure Example theme={null}
{
"code": 0,
"message": "success",
"data": {
"status": "failed",
"task_id": "task_id",
"model_id": "provider/model",
"error": "provider timeout"
}
}
```
### Receiver Response Requirements
To acknowledge receipt of the webhook, your server must return an HTTP `2xx` status code.
* **Recommended response:** `HTTP/1.1 200 OK` with a plain text body of `ok`.
* **Alternative response:** `HTTP/1.1 204 No Content` with an empty body.
The Modellix webhook system only checks the HTTP status code and does not parse or validate the response body.
### Retry Rules
If delivery fails, the system will attempt to redeliver the webhook under specific conditions.
#### Retried Errors
The system will automatically retry delivery for the following errors:
* HTTP status `429` (Too Many Requests)
* HTTP status `5xx` (Server Errors)
* Network timeouts
* Temporary network errors
#### Non-Retried Errors
The system will **not** retry delivery for the following errors:
* HTTP status `3xx` (Redirection)
* HTTP status `4xx` (Client Errors, except `429`)
* Invalid Webhook URLs
* URLs pointing to private networks or `localhost`
* Permanent connection errors (e.g., connection refused, unresolved DNS)
### Best Practices
To ensure reliable and secure webhook processing, we recommend following these guidelines:
1. **Idempotency:** Use the `X-Modellix-Delivery-ID` header to deduplicate incoming webhooks and prevent processing the same event multiple times.
2. **Asynchronous Processing:** To avoid timeouts, quickly persist the incoming payload or push it to a message queue, and immediately return a `200` or `204` response. Perform any heavy business logic asynchronously.
3. **Security:** Verify the request source or signature (if signature verification is supported in a future update).
## Error Handling
All error responses follow a unified JSON format:
```json theme={null}
{
"code": 400,
"message": "Invalid parameters: parameter 'prompt' is required"
}
```
* `code` (integer) — equals the HTTP status code (`0` on success)
* `message` (string) — formatted as `": "`
### Error Codes
| HTTP Status | Description | Common Scenarios | Retryable |
| ----------- | --------------------- | ----------------------------------------------------------- | ----------------------------- |
| 400 | Bad Request | Missing required parameters, invalid format, invalid values | No — fix parameters first |
| 401 | Unauthorized | Invalid API key, missing API key, expired API key | No — provide a valid key |
| 402 | Payment Required | Insufficient balance, account in arrears | No — recharge your account |
| 404 | Not Found | Task ID not found, model not found, provider not found | No — check resource ID |
| 429 | Too Many Requests | Rate limit exceeded, concurrent limit exceeded | Yes — use exponential backoff |
| 500 | Internal Server Error | Internal processing error, unexpected error | Yes — retry up to 3 times |
| 503 | Service Unavailable | Service temporarily unavailable, circuit breaker open | Yes — retry with backoff |
For **429** responses, check the `X-RateLimit-Reset` header to know when you can retry. Use exponential backoff (1 s → 2 s → 4 s) for **500** and **503** errors.
## Upload Media Files
Use the File API to upload images, videos, and audio that you can pass into Modellix prediction APIs. Uploads are authenticated with your API Key and are not billed. Files are retained for a limited time (default **7 days**).
Upload a single file via `multipart/form-data`.
List non-expired files for your team.
Delete a file and free upload quota immediately.
### Typical Workflow
Send a `POST` request to `/api/v1/media/files` with `multipart/form-data`. The form field name must be `file`.
The file must use a supported extension and contain valid content for that type.
```bash cURL theme={null}
curl -X POST 'https://api.modellix.ai/api/v1/media/files' \
-H 'Authorization: Bearer $API_KEY' \
-F 'file=@./image.png'
```
```python Python theme={null}
import os
import requests
resp = requests.post(
"https://api.modellix.ai/api/v1/media/files",
headers={"Authorization": f"Bearer {os.environ['API_KEY']}"},
files={"file": open("image.png", "rb")},
)
print(resp.json())
```
On success, the response includes a `file_id`, media `type` (`image`, `video`, or `audio`), and a `url`:
```json theme={null}
{
"code": 0,
"message": "success",
"data": {
"file_id": "550e8400-e29b-41d4-a716-446655440000",
"type": "image",
"url": "https://file.modellix.ai/example/550e8400-e29b-41d4-a716-446655440000.png",
"filename": "image.png",
"size": 102400,
"created_at": 1784084400000
}
}
```
Pass `data.url` into the image, video, or audio input fields of a prediction API (for example `image_url`).
```bash theme={null}
curl --request POST \
--url https://api.modellix.ai/api/v1/alibaba/qwen-image-edit/async \
--header 'Authorization: Bearer $API_KEY' \
--header 'Content-Type: application/json' \
--data '{
"prompt": "Change the background to a sunny garden",
"image_url": "https://file.modellix.ai/example/550e8400-e29b-41d4-a716-446655440000.png"
}'
```
Prefer the returned `url` over hosting the asset yourself when you only need a temporary public URL for model input.
List media files that belong to your team and have not yet expired. Default page size is `100`; `limit` is capped at `100`.
```bash theme={null}
curl -X GET 'https://api.modellix.ai/api/v1/media/files?limit=10&offset=0' \
-H 'Authorization: Bearer $API_KEY'
```
```json theme={null}
{
"code": 0,
"message": "success",
"data": {
"items": [
{
"file_id": "550e8400-e29b-41d4-a716-446655440000",
"type": "image",
"url": "https://file.modellix.ai/example/550e8400-e29b-41d4-a716-446655440000.png",
"filename": "image.png",
"size": 102400,
"created_at": 1784084400000
}
],
"total": 1,
"limit": 100,
"offset": 0
}
}
```
Delete a file you no longer need. After deletion, it no longer counts toward your upload limit.
```bash theme={null}
curl -X DELETE 'https://api.modellix.ai/api/v1/media/files/550e8400-e29b-41d4-a716-446655440000' \
-H 'Authorization: Bearer $API_KEY'
```
```json theme={null}
{
"code": 0,
"message": "success",
"data": {}
}
```
Returns `404` if the file is not found.
### Limits and Supported Formats
#### Default Limits
| Limit | Default |
| --------------------------- | ------------ |
| Max file size | 16 MB |
| Files per team | 10 |
| Concurrent uploads per team | 2 |
| Retention | About 7 days |
Files expire after the retention window. Delete unused files early if you are approaching the per-team file count limit.
#### Allowed File Extensions
| Category | Extensions |
| -------- | ------------------------------------------------------------------------- |
| image | `jpg`, `jpeg`, `png`, `webp`, `gif`, `bmp`, `tiff`, `tif`, `heic`, `heif` |
| video | `mp4`, `m4v`, `webm`, `mov`, `mkv`, `avi` |
| audio | `mp3`, `wav`, `m4a`, `aac`, `ogg`, `flac`, `opus`, `weba` |
The response `type` field is derived from the extension and is one of `image`, `video`, or `audio`.
### File API Errors
Error responses use the same JSON shape as other Modellix APIs:
```json theme={null}
{
"code": 400,
"message": "Invalid format: unsupported file format"
}
```
| HTTP status | Common causes |
| ----------- | ---------------------------------------------------------------------------------------- |
| 400 | Missing `file`, unsupported format, content mismatch, invalid image, or rejected content |
| 401 | Missing or invalid API Key |
| 404 | File not found (delete only) |
| 413 | File exceeds the maximum allowed size |
| 429 | Upload quota or concurrency limit exceeded |
| 500 | Internal server error |
Full request and response schemas: [Upload Media File](/api/upload-media-file), [List Media Files](/api/list-media-files), [Delete Media File](/api/delete-media-file).
# Modellix CLI for Image, Video, and Audio Generation
Source: https://docs.modellix.ai/ways-to-use/cli
Use modellix-cli to authenticate, discover models, inspect request schemas, submit async image, video, and audio tasks, wait for results, and download assets from the terminal or CI.
`modellix-cli` is the official command-line client for [Modellix](https://modellix.ai).
Use it to manage authentication profiles, discover and run models, inspect request schemas, wait for asynchronous tasks, download results, and produce stable output for scripts and agents.
For full API behavior and response fields, see the [REST API](/ways-to-use/api) guide.
Package details: [npm](https://www.npmjs.com/package/modellix-cli) · [GitHub](https://github.com/Modellix/modellix-cli)
## Requirements
* Node.js 18.17 or later
* A Modellix API key from the [Modellix Console](https://www.modellix.ai/console/api-key) (`model get-schema` is public and does not need one)
## Install
```bash theme={null}
npm install --global modellix-cli
modellix-cli --version
```
Running `modellix-cli` with no arguments prints a local Quickstart. It does not call the API or start an interactive wizard.
```bash theme={null}
modellix-cli
modellix-cli quickstart
modellix-cli --help
```
## Quickstart
```bash theme={null}
modellix-cli init
```
Or set `MODELLIX_API_KEY` for temporary and CI usage (see [Authentication](#authentication-and-profiles)).
```bash theme={null}
modellix-cli doctor
```
```bash theme={null}
modellix-cli model list
```
```bash theme={null}
modellix-cli model get-schema bytedance/seedream-4.5-t2i
```
```bash theme={null}
modellix-cli model run \
--model-slug bytedance/seedream-4.5-t2i \
--body '{"prompt":"A cute cat playing in a sunny garden"}'
```
```bash theme={null}
modellix-cli task get task-abc123
```
## Authentication and Profiles
Authenticated commands select a profile in this order:
1. `--profile`
2. `MODELLIX_PROFILE`
3. the saved `currentProfile`
4. `default`
They resolve the API key independently in this order:
1. `--api-key`
2. `MODELLIX_API_KEY`
3. the selected saved profile
### Recommended Setup
```bash theme={null}
modellix-cli auth login
modellix-cli auth login --profile work
modellix-cli auth status --profile work
```
`modellix-cli init` is the short setup command and supports the same profile selection.
Interactive prompts hide the key, and the CLI validates it before writing.
Non-interactive setup:
```bash theme={null}
modellix-cli init --api-key "$MODELLIX_API_KEY" --yes
modellix-cli init --api-key "$MODELLIX_API_KEY" --yes --json
modellix-cli init --api-key "$MODELLIX_API_KEY" --check
```
### Environment Variable (CI and Temporary Use)
```bash theme={null}
# macOS / Linux
export MODELLIX_API_KEY="your_api_key"
```
```powershell theme={null}
# Windows PowerShell
$env:MODELLIX_API_KEY = "your_api_key"
```
Prefer the hidden prompt or an environment variable over `--api-key` on the command line, which can remain in shell history.
`auth status`, `auth whoami`, `config show`, and JSON status output never print the credential value.
### Auth Commands
```bash theme={null}
modellix-cli auth login [--profile NAME]
modellix-cli auth status [--profile NAME] [--json]
modellix-cli auth whoami [--profile NAME] [--json]
modellix-cli auth logout [--profile NAME] [--yes]
```
### Local Configuration
```bash theme={null}
modellix-cli config path
modellix-cli config show
modellix-cli config show --json
modellix-cli config clear --profile work --yes
```
## Diagnose the Environment
```bash theme={null}
modellix-cli doctor
modellix-cli doctor --json
```
`doctor` checks the Node.js version, reports the API-key source without printing the key, validates API connectivity, and reads the team balance when authentication succeeds.
A failed required check returns a non-zero exit status.
## Discover Models
```bash theme={null}
modellix-cli model list
modellix-cli model list --type text-to-image --output slugs
modellix-cli model list --provider google --limit 20
modellix-cli model list --search banana
```
Use `--quiet` or `--output slugs` to print one slug per line.
Inspect one model:
```bash theme={null}
modellix-cli model describe google/nano-banana-2
modellix-cli model describe google/nano-banana-2 --json
```
## Get a Model API Schema
Use the exact `provider/model` value returned by `model list`.
This command calls the public [Get Schema](/api/get-schema) endpoint on `https://www.modellix.ai` and does not require an API key.
```bash theme={null}
modellix-cli model list --output slugs
modellix-cli model get-schema alibaba/qwen-image-3.0-pro
modellix-cli model get-schema alibaba/qwen-image-3.0-pro --output human
modellix-cli model get-schema alibaba/qwen-image-3.0-pro --quiet
```
JSON is the default. It preserves the complete `servers` and `post` fields returned by Modellix.
Human output summarizes the inference endpoint, summary, description, request body, and responses.
Quiet output prints only `servers[0].url` for scripts.
The printed inference URL is the async generate path on `https://api.modellix.ai`. Submitting a task with that URL, or with `model run`, still requires a Bearer API key.
The default schema host is `https://www.modellix.ai`. Pass `--base-url` only when testing a compatible local schema endpoint, for example `http://127.0.0.1:3000`. Setting `MODELLIX_BASE_URL` to `https://api.modellix.ai` overrides that default and will miss the public schema endpoint.
## Run a Model
### Inline JSON Body
```bash theme={null}
modellix-cli model run \
--model-slug bytedance/seedream-4.5-t2i \
--body '{"prompt":"A cute cat"}'
```
### JSON File Body
```bash theme={null}
modellix-cli model run \
--model-slug alibaba/qwen-image-edit \
--body-file ./payload.json
```
### Pipe JSON from Stdin
```bash theme={null}
printf '%s' '{"prompt":"A cute cat"}' | \
modellix-cli model run --model-slug google/nano-banana-2 --body-file -
```
### Common Flags
* `--model-slug` (required): exact `provider/model` value from `model list`
* `--body`: request JSON string
* `--body-file`: path to a JSON file, or `-` for stdin
* `--api-key`: API key (overrides environment and saved profile)
* `--wait`: poll until the task reaches a terminal state
* `--timeout`: wait deadline (for example `5m`, `30s`, or bare seconds)
* `--output task-id`: print only the new task ID for shell pipelines
Use either `--body` or `--body-file`, not both.
The request body must be a JSON object and is capped at 64 MiB.
### Submit and Wait in One Command
```bash theme={null}
modellix-cli model run \
--model-slug google/nano-banana-2 \
--body '{"prompt":"A cute cat"}' \
--wait --timeout 5m --quiet
```
The default remains asynchronous.
If waiting times out after a task ID is received, the error prints that ID and a safe `task wait` recovery command.
Paid POST submissions are never automatically retried.
If a network or protocol failure leaves the submission outcome unknown, do not immediately repeat the same request—check `task history` and account activity first.
### Compatibility Alias
`modellix-cli model invoke` remains an alias of `model run`.
New scripts should use `model run`.
Example success response (async submit):
```json theme={null}
{
"code": 0,
"message": "success",
"data": {
"status": "pending",
"task_id": "task-abc123",
"model_id": "alibaba/qwen-image-plus",
"get_result": {
"method": "GET",
"url": "https://api.modellix.ai/api/v1/tasks/task-abc123"
}
}
}
```
## Batch Model Tasks
`model batch` accepts one JSON object per line (JSONL):
```json theme={null}
{"modelSlug":"google/nano-banana-2","body":{"prompt":"First image"}}
{"modelSlug":"google/nano-banana-2","body":{"prompt":"Second image"}}
```
```bash theme={null}
modellix-cli model batch tasks.jsonl --max-tasks 10 --concurrency 3
cat tasks.jsonl | modellix-cli model batch - --yes --wait --quiet
```
Batch submission requires either `--max-tasks` or explicit `--yes`, because every line can create a paid task.
All slugs and JSON bodies are validated before the first POST.
Concurrency is limited to 1–10, and the absolute local limit is 1000 tasks.
## Query, Wait, and Download Tasks
### Get a Task Result
```bash theme={null}
modellix-cli task get task-abc123
modellix-cli task get task-abc123 --output human
modellix-cli task get task-abc123 --quiet
```
Example success response:
```json theme={null}
{
"code": 0,
"message": "success",
"data": {
"status": "success",
"task_id": "task-abc123",
"result": {
"resources": [
{
"url": "https://cdn.example.com/images/abc123.png",
"type": "image"
}
]
}
}
}
```
### Wait for One or More Tasks
```bash theme={null}
modellix-cli task wait task-abc123
modellix-cli task wait task-a task-b --interval 5s --timeout 10m --concurrency 8
```
Bare time values remain seconds.
One invocation accepts at most 1000 unique IDs.
Any failed terminal task exits `1`.
An overall timeout exits `124`; JSON mode includes completed responses and `unfinishedTaskIds`.
### Download Task Resources
```bash theme={null}
modellix-cli task download task-abc123 --output-dir ./results
modellix-cli task download task-abc123 --output-dir ./results --json
```
Downloads require a successful task.
Existing files are preserved by default; use `--overwrite` deliberately.
JSON output contains local paths and byte counts, not signed source URLs.
### Local Task History
Successful submissions are recorded locally so task IDs remain recoverable:
```bash theme={null}
modellix-cli task history
modellix-cli task history --limit 50 --json
modellix-cli task history --profile work --json
modellix-cli task history --clear --yes
```
History stores Task ID, profile, API origin, model slug, status, and timestamps—never an API key or request body.
## Output, Automation, and CI
Built-in Modellix business commands accept these common output controls:
| Flag | Behavior |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------- |
| `--json` / `--output json` | One machine-readable JSON document. Failures use `{ "ok": false, "error": { "exitCode", "message" } }`. |
| `--quiet` / `-q` / `--output quiet` | Only the primary value (slugs, task IDs, resource URLs, inference URLs, or local paths). |
| `--output human` | Concise readable output. |
| `--output slugs` | Compatible format for `model list` (one slug per line). |
| `--output task-id` | Compatible format for `model run` (task ID only). |
When flags overlap, quiet wins over JSON, then JSON wins over the command default.
Business output uses stdout; warnings, debug details, and recovery instructions use stderr.
For CI:
```bash theme={null}
modellix-cli init --api-key "$MODELLIX_API_KEY" --yes --json
modellix-cli doctor --json --no-color --no-progress
```
`CI` automatically disables colors, progress-capable output, and the background update check.
Set `MODELLIX_CLI_SKIP_NEW_VERSION_CHECK=true` to disable the update check explicitly.
### Exit Codes
| Code | Meaning |
| ----- | ------------------------------------------------------------------------------------ |
| `0` | Success |
| `1` | API, task, operation, or command validation failure |
| `2` | Argument-parser rejection or an explicit safety guard (for example batch cost limit) |
| `124` | Local task wait timeout; the remote task may still be running |
| `127` | Unknown command; suggestions are never executed automatically |
## Networking and Diagnostics
Override the API origin for a trusted gateway or local development server:
```bash theme={null}
modellix-cli doctor --base-url https://gateway.example.com
MODELLIX_BASE_URL=https://gateway.example.com modellix-cli model list
```
HTTPS is required; HTTP is accepted only for `localhost`, `127.0.0.1`, or `::1`.
Sanitized request diagnostics go to stderr:
```bash theme={null}
modellix-cli model list --verbose
modellix-cli model list --debug
```
Diagnostics include method, endpoint path, retry attempt, response status, and elapsed time.
They exclude API keys, request bodies, and response bodies.
## Shell Completion
```bash theme={null}
modellix-cli autocomplete
modellix-cli autocomplete bash
modellix-cli autocomplete zsh
modellix-cli autocomplete powershell
```
## Troubleshooting
| Problem | What to do |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Missing API key | Run `modellix-cli init`, set `MODELLIX_API_KEY`, or pass `--api-key`. |
| `401 Unauthorized` | Repair auth with `init` or `auth login`, then verify with `doctor`. |
| `402 Payment Required` | Recharge in the Modellix Console and retry. |
| `429 Too Many Requests` | Read-only commands already retry within their deadline. Do not blindly retry paid submissions. |
| Paid submission outcome unknown | Check `task history` and account activity before submitting again. |
| Schema slug not found | Confirm the slug with `model list --output slugs`, then retry `model get-schema`. |
| Download blocked | Downloads default to HTTPS and public network destinations. Use `--allow-insecure-http` or `--allow-private-network` only in explicitly trusted environments. |
API error codes the CLI surfaces:
* `400`: fix parameters or request body format before retrying
* `401`: API key missing, invalid, or expired
* `402`: insufficient balance
* `404`: verify `task_id`, `--model-slug`, or the `model get-schema` slug
* `429`: rate or concurrency limits
* `500` / `503`: temporary server-side issue
## Help
```bash theme={null}
modellix-cli --help
modellix-cli help --nested-commands
modellix-cli model run --help
modellix-cli model get-schema --help
modellix-cli task wait --help
```
Misspelled commands receive a nearest-command suggestion.
Suggestions are never executed automatically.
# Use Modellix with the DeepSeek Harness Plugin
Source: https://docs.modellix.ai/ways-to-use/deepseek-harness
Install the dsh-modellix plugin in DeepSeek Harness to use Modellix LLM models, media generation, and Web Search and Web Fetch with one API key.
[dsh-modellix](https://github.com/Modellix/dsh-modellix) is the official [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) plugin for Modellix. One [API key](https://www.modellix.ai/console/api-key) enables a chat-first workflow: ask in the session, keep the context in the conversation, and inspect completed work without leaving Harness.
The plugin is also listed in the [awesome-dsh-plugin](https://awesome-dsh-plugin.com/) community catalog.

The plugin registers its own Harness tools and provider. It does **not** install or invoke [`modellix-cli`](/ways-to-use/cli) at runtime.
This page covers the DeepSeek Harness plugin. For the Open Plugins package used in Claude Code, Codex, and Cursor, see [Plugin](/ways-to-use/plugin). To add only the LLM gateway by hand—without media or Web tools—see [DeepSeek Harness](/llm/agent/deepseek-harness).
Harness and this plugin currently use prerelease interfaces. Check the plugin [CHANGELOG](https://github.com/Modellix/dsh-modellix/blob/main/CHANGELOG.md) and peer dependencies before upgrading Harness.
## What the Plugin Provides
| Area | What you can do |
| --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **LLM** | Choose live Modellix models from the Harness model selector. The plugin loads the current catalog; it does not invent fallback entries when the catalog is unavailable. |
| **Media** | Ask the Agent to create, edit, animate, or narrate media. Results appear as live chat cards and in the right-side **Modellix Design** panel. |
| **Web** | Ask a current, external, or source-verification question in normal language. The Agent calls [Web Search](/api/web-search) and [Web Fetch](/api/web-fetch) when it needs the public web. |
You can turn Design, LLM, and Web on or off independently in Modellix settings.
## Requirements
* [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) `0.1.1-rc.2`
* Node.js `^22.19.0` or `>=24.0.0` for the published package
* A Modellix [API key](https://www.modellix.ai/console/api-key)
## Install
Install the package in the Harness Web profile, inspect the merged configuration, then start or restart that profile:
```bash theme={null}
dsh plugin --profile web add dsh-modellix
dsh --profile web --dump-config
dsh --profile web
```
`--dump-config` should contain the `dsh-modellix` bundle layer and a plugin row with id `modellix`. Replace `web` if you use another profile. Restart the running profile after you install or update the plugin.
To install a trusted local build from a clone of [dsh-modellix](https://github.com/Modellix/dsh-modellix):
```bash theme={null}
pnpm install --frozen-lockfile
pnpm run verify:release:static
pnpm pack
dsh plugin --profile web add ./dsh-modellix-0.2.0.tgz
```
## Configure the API Key
Create or copy a key in the [Modellix Console](https://www.modellix.ai/console/api-key) before you connect the plugin.
Start the Harness Web UI. On first use, open **Connect Modellix**.
Enter a valid Modellix [API key](https://www.modellix.ai/console/api-key). Keep **Design**, **LLM**, and **Web** enabled unless you intend to turn one capability off.
Choose **Save and enable**. Then open **Settings → Modellix** and confirm that the Credential status and LLM catalog are healthy.
The stored key is write-only. After save, the UI shows status and source, never the key itself.

Alternatively, provide `MODELLIX_API_KEY` in the Harness launch environment. Environment credentials are read-only in the UI and require a Harness restart after you replace the value.
Do not put a real key in a repository, command argument, URL, browser storage, log, screenshot, or test snapshot.
Select **Configure later** only when you want to postpone setup. The plugin stays unavailable and asks again the next time you use an enabled Modellix capability.
## Use Modellix LLM Models
When LLM is enabled, the plugin reads the live catalog and adds those models to the Harness model selector.
In **Settings → Modellix**, leave LLM on. Check catalog status and count, or choose **Refresh** if the list looks stale.
Open the Harness model selector and choose a model in the Modellix group. Send the next Agent turn.
The provider is OpenAI-compatible and uses the Modellix LLM gateway. Provider retries are `0`. If the catalog cannot load, the plugin does not fabricate model entries.
You can still add Modellix as a custom `llm-pi-ai` provider by hand. That path covers LLM only. See [DeepSeek Harness](/llm/agent/deepseek-harness) for the form and `settings.yaml` values.
Full Model IDs and rates: [Models & Pricing](/llm/overview#models-and-pricing).
## Create Media in Chat
Describe the outcome in the conversation. You do not need to name tools or walk through a catalog first.
```text theme={null}
Create a polished 16:9 architectural hero image of a glass botanical research pavilion floating above a dawn cloud sea, with restrained lapis-blue and warm-gold tones, realistic premium materials, no people, no text, and no watermark.
```
The Agent can then:
1. Search the live media catalog when a compatible model is not already known.
2. Read that model's live API schema and use only published fields and values.
3. Reuse the latest relevant result URL for edit, image-to-video, or video-to-video follow-ups instead of starting a new text-to-media task.
4. Upload a session attachment or a workspace file when the schema requires a public media URL.
5. Submit the generation once. An unknown submission outcome is never replayed automatically.
6. Check the task once in the same turn. A background watcher updates the existing result card when the job finishes.
Continue from the previous result in the same session:
```text theme={null}
Turn the image just completed in this conversation into a five-second cinematic video. Slowly push the camera forward and preserve the composition and palette.
```
Generate speech the same way, and keep voice names and audio parameters to values published in the model schema:
```text theme={null}
Generate this English voiceover with a professional narrator, calm emotion, MP3 at 44.1 kHz: “From one idea to images, video, and sound, Modellix Design keeps creation flowing naturally in the conversation.”
```
### Result Cards and Modellix Design
* While a task is running or has failed, the chat card shows a concise header and status. Preview and JSON appear only after success.
* A successful card updates in place with Preview and JSON tabs, image enlargement, and native video or audio players.
* One task maps to one card. A later result lookup does not create a duplicate.
* The right-side **Modellix Design** panel lists only tasks from the current Harness conversation.
* **Add URL to chat** appends the selected resource URL to the composer so you can edit or transform it. **Download** opens the upstream file.
* Result URLs follow the upstream expiry. If the API does not provide one, the plugin applies a seven-day local display limit. It does not keep a permanent media copy.
Open **Modellix Design** from the far right of the conversation header, next to **Session log**. On large screens it is a split panel; on narrow screens it is a full-width overlay.

Design must stay enabled for media tools, chat result cards, and the result panel. Turning Design off removes those tools from the Agent.
## Search and Fetch the Web
Ask the question normally. You do not need to say “use search” or “use fetch.”
```text theme={null}
Verify the official Modellix page for alibaba/wan2.7-videoedit. Give its title, one required parameter and what it means, with a source. Do not answer from memory.
```
For current, changing, external, or source-verification questions, the Agent calls `modellix_web_search`. When you provide a public URL, or a search result needs the full page, it calls `modellix_web_fetch`. Failed or unknown Web requests are not repeated automatically.

If you explicitly ask the Agent not to browse, it does not call these tools. Web must stay enabled in Modellix settings.
API reference: [Web Search](/api/web-search) · [Web Fetch](/api/web-fetch).
## Settings and Recovery
The Modellix settings section shows:
* Credential configured or verification status, and whether the source is a saved key or `MODELLIX_API_KEY`
* Replace and remove actions for a writable local credential
* Independent **Design**, **LLM**, and **Web** switches
* Live LLM catalog health, model count, refresh time, and a manual refresh
Only HTTP `401` marks a credential invalid. Other failures keep their own recovery states:
| Status | What to do |
| ------------------------------------ | ------------------------------------------------------------------------------------------ |
| `402` | Check account status and balance in the [console](https://www.modellix.ai/console/api-key) |
| `429` | Wait for the rate-limit window, then retry |
| Offline or timeout | Restore connectivity. Do not assume the key is invalid |
| `5xx` | Treat it as a service error. Retry only when the operation is safe to repeat |
| Unknown generation or upload outcome | Inspect the task, transcript, or Modellix record before submitting again |
## Troubleshooting
| Problem | What to do |
| ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Modellix Design** is missing | Confirm Design is enabled, inspect `dsh --profile web --dump-config` for the `dsh-modellix` layer and plugin id `modellix`, then fully restart the profile |
| The Agent does not use Web automatically | Confirm Web is enabled and start a new session so routing instructions load. The tools should appear as `modellix_web_search` and `modellix_web_fetch` |
| A model schema is reported as unavailable | Refresh the live media catalog, use the exact slug from the catalog, and read the schema again. Unsupported contracts block submission instead of guessing |
| A task never updates | Keep the conversation open for the client watcher, confirm the credential has not changed, and check the right-side panel. Do not resubmit because an older assistant sentence still says the job is running |
| Media cannot play | Confirm the result succeeded, the upstream URL has not expired, and the browser can reach the file origin. Running and failed tasks have no player |
## Uninstall
Remove a writable local key in Modellix settings first, or revoke an environment credential in your secret manager. Then remove the plugin and restart the profile:
```bash theme={null}
dsh plugin --profile web remove dsh-modellix
dsh --profile web --dump-config
dsh --profile web
```
Uninstalling does not delete upstream Modellix tasks, external environment variables, or all Harness profile data.
## Next Steps
Add Modellix as a custom provider without the plugin.
Review Model IDs and rates for the live catalog.
Call media models directly with submit-and-poll.
Use the same Web tools outside Harness.
# Docs
Source: https://docs.modellix.ai/ways-to-use/docs-search-mcp
Connect AI clients like Cursor and Claude Desktop to the Docs MCP to search the official documentation without leaving your editor.
This MCP searches Modellix documentation. It does not generate images, videos, or speech. For media generation through MCP, see [Modellix Media Generation](/ways-to-use/media-generation-mcp).
MCP (Model Context Protocol) is an open-source standard for connecting AI applications to external systems. For more information, please refer to the [MCP documentation](https://modelcontextprotocol.io/).
Compatible with both [Cursor](https://cursor.sh/) and [Claude Desktop](https://claude.ai/download)!
Modellix Docs MCP is also compatible with any MCP client.
Modellix Docs MCP provides a search tool for your AI application to initiate search requests within the Modellix documentation.
## Remote Server
The easiest way to take advantage of Modellix Docs MCP is by using the remote URL. This provides a seamless experience without requiring local installation or configuration.
Simply use the remote MCP server URL:
```plaintext theme={null}
https://docs.modellix.ai/mcp
```
### Connect to Clients

You can also connect Modellix Docs MCP to popular coding AI tools like Cursor and VS Code with just one click.
Find the "Copy Page" menu on the right side of the title on each documentation page, where you'll find buttons to connect to Cursor and VS Code with one click.
### With Smithery
You can also connect to Modellix Docs MCP on [Smithery](https://smithery.ai/server/modellix/modellix-docs).
### With MCP.so
You can also connect to Modellix Docs MCP on [MCP.so](https://mcp.so/server/modellix-docs/Modellix).
### OpenAI
Allow models to use remote MCP servers to perform tasks.
```python theme={null}
from openai import OpenAI
client = OpenAI()
resp = client.responses.create(
model="gpt-4.1",
tools=[
{
"type": "mcp",
"server_label": "modellix-docs",
"server_url": "https://docs.modellix.ai/mcp",
"require_approval": "never",
},
],
input="Do you have access to the modellix docs mcp server?",
)
print(resp.output_text)
```
## How It Works
Once your AI application integrates with Modellix Docs MCP, it can directly search the Modellix documentation instead of performing a generic web search when responding to user prompts. Modellix Docs MCP provides access to all indexed content on the documentation site.
* AI applications can proactively search Modellix documentation while generating responses, not just when explicitly requested.
* AI applications decide when to use the search tool based on the conversation context and the relevance of Modellix documentation to the current topic.
* Each search (tool call) occurs during the generation process, so AI applications retrieve the latest information from Modellix documentation to generate responses.
## Search Filter Parameters
The MCP search tool supports optional filter parameters that AI applications can use to narrow down search results.
| Field Name | Description |
| ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `version` | Filter results to a specific documentation version. For example, `v0.7`. Only returns content with the specified version tag, or content that is common across all versions. |
| `language` | Filter results to a specific language code. For example, `en`, `zh`, or `es`. Only returns content in the specified language, or content that is common across all languages. |
| `apiReferenceOnly` | When set to true, only returns API reference documentation pages. |
| `codeOnly` | When set to true, only returns code snippets and examples. |
AI applications decide when to apply these filters based on the context of the user's query. For example, if a user asks about a specific API version or requests code examples, the AI application may automatically apply the appropriate filters to provide more relevant results.
# Modellix Media Generation
Source: https://docs.modellix.ai/ways-to-use/media-generation-mcp
Modellix Media Generation is an MCP server for image, video, and speech generation. It is in development.
**Modellix Media Generation** is in development. Stay tuned — install instructions and client configs will land on this page when the server is ready.
**Modellix Media Generation** is an MCP server for generating images, videos, and speech through Modellix from MCP clients such as Cursor and Claude Desktop.
This is not the [Docs](/ways-to-use/docs-search-mcp) MCP, which only searches these docs.
## Until It Ships
Use these supported paths for media generation today:
Submit async image, video, and speech tasks over HTTP.
Teach coding agents how to discover models and run generation.
Authenticate, submit tasks, wait, and download results from the terminal.
Generate and preview media in the browser with no code.
# Generate Images, Videos, and Audio in the Playground
Source: https://docs.modellix.ai/ways-to-use/playground
Generate images, videos, and speech audio in the Modellix Playground with no code required. Discover models, tune prompts, and preview outputs in your browser.
For creators, designers, and prompt engineers, you don't need to write code to experience the power of Modellix models. We provide a fully functional **Playground** directly on the Modellix website.
## Discover Models
Before generating content, you can easily find the model that best suits your needs:
* **[Modellix Models](https://www.modellix.ai/models)**: Browse our curated list of featured models, latest releases, and model families.
* **[Modellix Explore](https://www.modellix.ai/explore)**: Use our comprehensive search page to filter and find any model by provider, category (e.g., text-to-image, text-to-video, text-to-speech), or specific features.
## Create in the Playground
Once you find a model you want to try, simply click on it to enter its detail page.
Every model's detail page features a built-in **Playground**. Here, you can:
* Input your text prompts (You can optimize your prompts with AI).
* Upload reference images, videos, or audio when the model supports them.
* Adjust model-specific parameters like resolution, aspect ratio, duration, voice, and more.
* Generate and preview the results instantly.
### Playground Interface

### Benefits for Creators
* **Zero Setup**: No need to configure API keys or development environments. Just log in and start creating.
* **Visual Controls**: Intuitively adjust parameters with sliders and dropdowns instead of JSON payloads.
* **Instant Feedback**: Iterate rapidly on your prompts and see the results immediately.
* **Inspect JSON Payload**: Switch to the **JSON** tab in the result panel to inspect the exact API response, seamlessly bridging the gap between prototyping and development.
* **Transparent Pricing**: The detailed pricing for every parameter combination is clearly displayed, so you always know exactly what you're spending.
* **One-Click Delivery**: Send your generated media (including images, videos, and audio) directly to external storage or workflows with a single click. We currently support exporting to:
* Designated Webhook URLs
* Google Drive
* Dropbox
# Modellix Plugin for Coding Agents
Source: https://docs.modellix.ai/ways-to-use/plugin
Install the official Modellix Plugin in Claude Code, Codex, Cursor, and other agent hosts to generate images, videos, and speech via CLI or REST API.
The [Modellix Plugin](https://github.com/Modellix/modellix-plugin) packages everything your coding agent needs to generate images, videos, and speech audio with [Modellix](https://modellix.ai): the official Modellix skill, model discovery guidance, default model choices, credential handling, and retry rules that match the CLI exit codes.
For a visual Excalidraw workspace with image prepare/confirm, HTML drafts, and presentations, see [Agent Canvas](/ways-to-use/agent-canvas) instead.
The repository follows the [Open Plugins](https://open-plugins.com/plugin-builders/specification) specification, so the same layout installs into Claude Code, Codex, Cursor, OpenClaw, OpenCode, Pi, Hermes, and any Agent Skills host.
## Plugin or Skill
The repository ships in two shapes. Pick the one your host supports.
| Install shape | What you get | When to use it |
| ------------- | -------------------------------------------------------- | ----------------------------------------------------------------------------- |
| **Plugin** | Repository root: plugin manifests plus `skills/modellix` | Your host has a plugin marketplace (Claude Code, Codex, Cursor, OpenClaw, Pi) |
| **Skill** | Only the `skills/modellix` Agent Skill | Your host has no plugin marketplace, or you only want the skill |
If you only need the Agent Skill, see the [Skill](/ways-to-use/skill) guide, which covers the standalone `modellix-skill` repository as well.
## What the Plugin Provides
* A CLI-first workflow: `modellix-cli doctor` → `model run --wait` → `task download`
* Automatic REST fallback when the CLI is not installed
* Sensible default models when you do not name one
* Model discovery through `modellix-cli model list`, `model describe`, and `model get-schema`, plus the live model index at [llms.txt](https://docs.modellix.ai/llms.txt)
* Retry and error guidance aligned with CLI exit codes, including paid-submission safety rules
* Credential handling for `MODELLIX_API_KEY` and saved CLI auth profiles
## Requirements
* A Modellix API key from the [Modellix Console](https://www.modellix.ai/console/api-key)
* Recommended: [`modellix-cli`](https://www.npmjs.com/package/modellix-cli) on Node.js 18.17 or later
```bash theme={null}
npm install --global modellix-cli@latest
modellix-cli doctor --json
```
The plugin works without the CLI by calling the [REST API](/ways-to-use/api) directly, but the CLI gives your agent faster task submission, built-in waiting, and safer downloads.
## Install as a Plugin
Install:
```text theme={null}
/plugin marketplace add Modellix/modellix-plugin
/plugin install modellix@modellix
```
Update:
```text theme={null}
/plugin marketplace update modellix
/plugin update modellix@modellix
```
You can also use the `/plugin` UI, then run `/reload-plugins` if the plugin does not appear.
For local development against a clone:
```bash theme={null}
claude --plugin-dir /path/to/modellix-plugin
claude plugin validate /path/to/modellix-plugin
```
Add the marketplace, then install `modellix` from the `/plugins` list:
```bash theme={null}
codex plugin marketplace add Modellix/modellix-plugin
```
To update, reopen `/plugins` and update from there, or re-add the marketplace and reinstall.
Browse or submit the plugin through [cursor.directory](https://cursor.directory/plugins/new) using the repository URL `https://github.com/Modellix/modellix-plugin`.
To install from a local clone:
```bash theme={null}
git clone https://github.com/Modellix/modellix-plugin.git
ln -sfn "$PWD/modellix-plugin" ~/.cursor/plugins/local/modellix
```
Run **Developer: Reload Window**, then confirm the plugin under **Customize**.
Update a symlinked clone:
```bash theme={null}
git -C ~/.cursor/plugins/local/modellix pull
```
Reload the window again after pulling.
The ClawHub package `@modellix/modellix-plugin` ships the Open Plugins layout as a content and skill bundle, not a TypeScript runtime plugin.
```bash theme={null}
openclaw plugins install clawhub:@modellix/modellix-plugin
```
Install from a local checkout or Git instead:
```bash theme={null}
openclaw plugins install .
openclaw plugins install git:github.com/Modellix/modellix-plugin
```
To update, reinstall from the same ClawHub, Git, or path source. `openclaw plugins update` only refreshes npm-tracked installs.
[Pi](https://github.com/badlogic/pi-mono) loads the repository as a [Pi package](https://docs.pi.dev/packages) that contributes skills only.
```bash theme={null}
pi install git:github.com/Modellix/modellix-plugin
```
Local checkouts work too:
```bash theme={null}
pi install /path/to/modellix-plugin
```
Update:
```bash theme={null}
pi update --extensions
```
## Install as a Skill
When your host has no plugin marketplace, or you only want `skills/modellix`, install the Agent Skill instead:
```bash theme={null}
npx skills add https://github.com/Modellix/modellix-plugin --skill modellix
```
The [Skill](/ways-to-use/skill) guide covers skill installs for skills.sh, ClawHub, OpenCode, Cursor, Pi, Hermes, and Smithery.
## Configure Your API Key
Every install shape reads the same credential.
```bash theme={null}
export MODELLIX_API_KEY="your_api_key"
```
The CLI resolves the key in this order: `--api-key`, then `MODELLIX_API_KEY`, then the selected saved profile.
* REST calls always require `MODELLIX_API_KEY`.
* The CLI can use the environment variable or a saved profile created by `modellix-cli auth login` or `modellix-cli init`.
* In Cursor, you can set the key as the `MODELLIX_API_KEY` plugin variable.
* In Hermes, prefer `~/.hermes/.env` or the secure prompt shown when the skill loads.
Prefer session-only keys and persist them only when you intend to. Never commit an API key to a repository or print it in logs.
## Verify the Install
Confirm the environment first, then generate one asset end to end.
```bash theme={null}
modellix-cli doctor --json
```
`doctor` reports the Node.js version, the API-key source, API connectivity, and your team balance without printing the key.
```text theme={null}
Generate an image of a cinematic sunset over a futuristic city skyline.
```
The plugin routes the request through the CLI, picks a default model, and waits for the task to finish.
The agent downloads results into your working directory. To reproduce the same flow manually:
```bash theme={null}
modellix-cli model run \
--model-slug google/nano-banana-2-lite \
--body '{"prompt":"A cinematic sunset over a futuristic city skyline"}' \
--wait --timeout 5m --json
modellix-cli task download --output-dir ./outputs --json
```
## Supported Task Types
| Type | Description |
| ---------------- | ----------------------------------------------- |
| `text-to-image` | Generate images from text prompts |
| `image-to-image` | Edit or transform images with text instructions |
| `text-to-video` | Create videos from text descriptions |
| `image-to-video` | Convert static images into video sequences |
| `video-to-video` | Transform existing videos |
## Default Models
When you do not name a model, the plugin uses these defaults:
| Task type | Default model slug |
| ------------------------------ | --------------------------------- |
| Text-to-image | `google/nano-banana-2-lite` |
| Image editing / image-to-image | `google/nano-banana-2-lite-edit` |
| Text-to-video | `bytedance/seedance-2.0-mini-t2v` |
| Image-to-video | `bytedance/seedance-2.0-fast-i2v` |
| Video-to-video | `bytedance/seedance-2.0-fast-v2v` |
Name any other model in your prompt to override the default. To discover and inspect models:
```bash theme={null}
modellix-cli model list --type text-to-image --output slugs
modellix-cli model describe google/nano-banana-2 --json
modellix-cli model get-schema google/nano-banana-2
```
Request-body schemas come from `modellix-cli model get-schema` (JSON is the default; the endpoint is public and does not need an API key). See the [CLI schema command](/ways-to-use/cli#get-a-model-api-schema).
If the CLI is unavailable, the plugin uses Docs MCP when connected, then `docs_url` from `model describe`, or the model page in [llms.txt](https://docs.modellix.ai/llms.txt). The [Skill](/ways-to-use/skill#inspect-a-request-schema) guide covers when the agent fetches the schema versus when it skips the call.
## How the Plugin Runs Tasks
Understanding these rules helps you predict what your agent will do:
1. It prefers the CLI when installed, and falls back to the [REST API](/ways-to-use/api) otherwise.
2. It reads request schemas with `model get-schema` before non-trivial submits, and skips that call for documented default examples or a complete body you already supplied.
3. It uses `model run --wait` or `task wait` instead of hand-written polling loops.
4. It never blindly retries a paid `model run` after an unknown submission outcome. It checks `modellix-cli task history` first.
5. It may use the helper scripts bundled in `skills/modellix/scripts/`, and falls back to plain CLI commands if a helper fails.
6. It treats `modellix-cli --help` and the [npm package](https://www.npmjs.com/package/modellix-cli) as the source of truth for CLI behavior.
## Troubleshooting
| Problem | What to do |
| -------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| The agent does not use Modellix | Confirm the plugin or skill is installed and enabled, then start a new session. In Claude Code, run `/reload-plugins`. |
| A model schema is unavailable | Confirm the slug with `model list --output slugs`, then retry `model get-schema`. If the CLI is missing, use Docs MCP or [llms.txt](https://docs.modellix.ai/llms.txt). |
| `401 Unauthorized` | Set `MODELLIX_API_KEY`, or repair the saved profile with `modellix-cli auth login`, then verify with `doctor`. |
| `402 Payment Required` | Recharge in the [Modellix Console](https://www.modellix.ai/console/api-key) and retry. |
| A paid submission outcome is unknown | Run `modellix-cli task history` and check account activity before submitting again. |
| `task download` fails with a private or reserved network error | Local proxies sometimes map CDN hosts into `198.18.0.0/15`. Retry with `--allow-private-network` for trusted Modellix CDN hosts, or download the resource URL with `curl`. |
| `openclaw plugins update` does nothing | That command only refreshes npm-tracked installs. Reinstall from the original ClawHub, Git, or path source. |
## Next Steps
Install the Agent Skill on its own.
Learn the `modellix-cli` commands the plugin runs.
Call Modellix directly without the CLI.
Review model pricing before running paid tasks.
# Install the Modellix Agent Skill
Source: https://docs.modellix.ai/ways-to-use/skill
Install the Modellix Agent Skill so coding agents in Cursor, Claude Code, and other hosts can discover models, inspect request schemas, and generate images, videos, and speech.
Agent Skills are folders of instructions, scripts, and resources that agents discover and load on demand to work more accurately. For background, see [Agent Skills](https://agentskills.io/home).
The official Modellix Skill teaches your coding agent how to generate images, videos, and speech audio with [Modellix](https://modellix.ai). It ships as `skills/modellix` inside the [modellix-plugin](https://github.com/Modellix/modellix-plugin) repository and installs into any Agent Skills host.
The skill previously lived in a separate `modellix-skill` repository. That repository is now `modellix-plugin`, and the skill is maintained at `skills/modellix`. Old URLs still redirect, but update your install commands to the new repository.
## Skill or Plugin
| Install shape | What you get | When to use it |
| ------------- | ------------------------------------------------ | ----------------------------------------------------------------------------- |
| **Skill** | Only `skills/modellix` | Any Agent Skills host, or when you do not want a full plugin install |
| **Plugin** | Repository root: plugin manifests plus the skill | Your host has a plugin marketplace (Claude Code, Codex, Cursor, OpenClaw, Pi) |
See the [Plugin](/ways-to-use/plugin) guide for marketplace installs. Both shapes deliver the same skill and behave the same at runtime.
## What the Skill Gives Your Agent
* A CLI-first workflow: `modellix-cli doctor` → `model get-schema` when the body is non-trivial → `model run --wait` → `task download`
* REST fallback with submit-and-poll logic when the CLI is unavailable
* Default models for each task type, so simple prompts do not require a catalog scan
* Model discovery through `model list` and `model describe`, plus live request schemas from `model get-schema`
* An API key lifecycle policy: discover an existing key before asking, prefer session-only keys, and persist only on request
* Retry rules mapped to HTTP status codes and CLI exit codes, including a hard rule against blindly re-running paid submissions
## Requirements
* A Modellix API key from the [Modellix Console](https://www.modellix.ai/console/api-key)
* Recommended: [`modellix-cli`](https://www.npmjs.com/package/modellix-cli) on Node.js 18.17 or later
```bash theme={null}
npm install --global modellix-cli@latest
modellix-cli doctor --json
```
The skill works without the CLI by calling the [REST API](/ways-to-use/api), but the CLI gives your agent single-command waiting, safer downloads, and local task history.
## Install
Works with any Agent Skills host:
```bash theme={null}
npx skills add https://github.com/Modellix/modellix-plugin --skill modellix
```
Target one agent:
```bash theme={null}
npx skills add https://github.com/Modellix/modellix-plugin --skill modellix --agent cursor
```
Update:
```bash theme={null}
npx skills update
```
Install from [ClawHub](https://clawhub.ai/Modellix/modellix) with the slug `modellix`:
```bash theme={null}
clawhub install modellix
```
Or through OpenClaw:
```bash theme={null}
openclaw skills install modellix
```
Update:
```bash theme={null}
clawhub update modellix
clawhub update --all
```
The skill slug `modellix` is separate from the OpenClaw bundle package `@modellix/modellix-plugin`, which is covered in the [Plugin](/ways-to-use/plugin) guide.
```bash theme={null}
npx skills add https://github.com/Modellix/modellix-plugin --skill modellix --agent cursor
```
Update:
```bash theme={null}
npx skills update
```
You can also set the key as the `MODELLIX_API_KEY` plugin variable if you install the full [plugin](/ways-to-use/plugin) instead.
OpenCode [plugins](https://opencode.ai/docs/plugins/) are JavaScript event hooks, which Modellix does not use. Install the [Agent Skill](https://opencode.ai/docs/skills/) instead:
```bash theme={null}
npx skills add https://github.com/Modellix/modellix-plugin --skill modellix
```
Or symlink the skill from a clone:
```bash theme={null}
# Global
mkdir -p ~/.config/opencode/skills
ln -sfn /path/to/modellix-plugin/skills/modellix ~/.config/opencode/skills/modellix
# Project-local
mkdir -p .opencode/skills
ln -sfn /path/to/modellix-plugin/skills/modellix .opencode/skills/modellix
```
Load it in a session with `skill({ name: "modellix" })`.
[Pi](https://github.com/badlogic/pi-mono) can load the repository as a package, which is the preferred path documented in the [Plugin](/ways-to-use/plugin) guide. For a skill-only install:
```bash theme={null}
npx skills add https://github.com/Modellix/modellix-plugin --skill modellix
```
Pi also scans `~/.agents/skills/`. To symlink the skill tree directly:
```bash theme={null}
mkdir -p ~/.pi/agent/skills
ln -sfn /path/to/modellix-plugin/skills/modellix ~/.pi/agent/skills/modellix
```
[Hermes](https://hermes-agent.nousresearch.com/) uses Agent Skills rather than plugins:
```bash theme={null}
hermes skills install Modellix/modellix-plugin/skills/modellix
```
Or symlink into the Hermes skills tree:
```bash theme={null}
mkdir -p ~/.hermes/skills
ln -sfn /path/to/modellix-plugin/skills/modellix ~/.hermes/skills/modellix
```
To reuse a shared Agent Skills directory, add it to `~/.hermes/config.yaml`:
```yaml theme={null}
skills:
external_dirs:
- ~/.agents/skills
```
Start a new session after installing, then invoke `/modellix`. Hermes may prompt securely for `MODELLIX_API_KEY` on first load because the skill declares it as a required environment variable.
```bash theme={null}
npx @smithery/cli@latest skill add modellix/modellix-skill
npx @smithery/cli@latest skill add modellix/modellix-skill --agent cursor
```
To update, run the same `skill add` command again.
## Configure Your API Key
```bash theme={null}
export MODELLIX_API_KEY="your_api_key"
```
The skill follows a `discover → request → use for the session → optionally persist` policy:
1. It first checks the session environment for `MODELLIX_API_KEY`.
2. It then checks for a saved CLI profile through `modellix-cli auth status` or `doctor`.
3. Only when neither exists does it ask you for a key.
The CLI resolves keys in this order: `--api-key`, then `MODELLIX_API_KEY`, then the selected saved profile.
The skill does not persist your key automatically. Ask for persistence explicitly, and prefer `modellix-cli auth login` or `modellix-cli init` so the CLI validates and stores the profile. Never commit an API key or print it in logs.
## Verify the Install
```bash theme={null}
modellix-cli doctor --json
```
`doctor` reports the Node.js version, the API-key source, API connectivity, and your team balance without printing the key.
```text theme={null}
Generate an image of a cinematic sunset over a futuristic city skyline.
```
The skill picks a default model, submits the task, waits for it, and downloads the result.
```bash theme={null}
modellix-cli model run \
--model-slug google/nano-banana-2-lite \
--body '{"prompt":"A cinematic sunset over a futuristic city skyline"}' \
--wait --timeout 5m --json
modellix-cli task download --output-dir ./outputs --json
```
## Supported Task Types
| Type | Required body fields |
| ---------------- | ------------------------------------------------------------------------------ |
| `text-to-image` | `prompt` |
| `image-to-image` | `prompt` and an `image` array |
| `text-to-video` | `prompt` |
| `image-to-video` | At least one of `first_frame_image`, `last_frame_image`, or `reference_images` |
| `video-to-video` | A `video_urls` array |
Exact schemas come from `modellix-cli model get-schema`. See [Inspect a Request Schema](#inspect-a-request-schema).
## Default Models
When you do not name a model, the skill uses these defaults immediately instead of scanning the catalog:
| Task type | Default model slug |
| ------------------------------ | --------------------------------- |
| Text-to-image | `google/nano-banana-2-lite` |
| Image editing / image-to-image | `google/nano-banana-2-lite-edit` |
| Text-to-video | `bytedance/seedance-2.0-mini-t2v` |
| Image-to-video | `bytedance/seedance-2.0-fast-i2v` |
| Video-to-video | `bytedance/seedance-2.0-fast-v2v` |
Name any other model in your prompt to override the default:
```bash theme={null}
modellix-cli model list --type text-to-image --output slugs
modellix-cli model describe google/nano-banana-2 --json
modellix-cli model get-schema google/nano-banana-2
```
Model slugs are exact, and decimals matter. Use `bytedance/seedance-2.0-mini-t2v`, not a slug guessed from a documentation filename.
## Inspect a Request Schema
The skill reads the live OpenAPI-style request and response contract with `modellix-cli model get-schema `. The endpoint is public and does not need an API key. JSON is the default.
```bash theme={null}
modellix-cli model get-schema alibaba/qwen-image-3.0-pro
modellix-cli model get-schema alibaba/qwen-audio-3.0-tts-flash --output human
```
Command details: [CLI schema command](/ways-to-use/cli#get-a-model-api-schema). REST contract: [Get Schema](/api/get-schema).
The agent calls `get-schema` when:
* you named a model that is not a skill default
* the body needs fields beyond the documented examples
* a previous submit returned HTTP `400`
* it needs to report required fields, such as TTS voices
It skips `get-schema` when the skill default plus the example already lists the required fields, or you supplied a complete body.
If the CLI is unavailable, the skill falls back to Docs MCP when connected, then `docs_url` from `model describe`, or the model page in [llms.txt](https://docs.modellix.ai/llms.txt).
## What Is Inside the Skill
| Path | Purpose |
| --------------------------------------- | ----------------------------------------------------------------------------- |
| `SKILL.md` | Execution policy, default models, credential lifecycle, error and retry rules |
| `references/cli-playbook.md` | Install, auth, get-schema, run, wait, download, batch, and recovery commands |
| `references/rest-playbook.md` | REST submit-and-poll flow when the CLI is unavailable |
| `references/capability-matrix.md` | CLI to REST mapping and fallback rules |
| `scripts/preflight.py` | Optional environment check that wraps `doctor` and recommends CLI or REST |
| `scripts/invoke_and_poll.py` | Optional wrapper around submit, wait, and download |
| `assets/output/task-result.schema.json` | Schema for task result payloads |
Your agent loads references progressively, reading only the files a task needs.
## Error Handling the Skill Follows
| Situation | Skill behavior |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| `400` | Does not retry. Fixes parameters or the request body first, using `model get-schema` when the contract is unclear. |
| `401` | Does not retry. Repairs auth with `doctor` or `auth login`. |
| `402` | Does not retry. Reports insufficient balance. |
| `404` | Does not retry. Verifies the task ID or model slug. |
| `429` or read-only `5xx` | Relies on CLI retries for safe GET requests, and never re-POSTs a paid submission blindly. |
| Unknown paid submission outcome | Checks `modellix-cli task history` and console activity before submitting again. |
| Exit code `124` | Treats it as a local wait timeout and recovers with `task wait` or `task get`. |
| Exit code `2` | Treats it as an argument or safety-guard rejection and fixes the flags. |
## Troubleshooting
| Problem | What to do |
| -------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| The agent ignores the skill | Confirm the skill is installed and enabled, then start a new session. Skills load on demand, so mention Modellix or image, video, or speech generation in your prompt. |
| A model schema is unavailable | Confirm the slug with `model list --output slugs`, then retry `model get-schema`. If the CLI is missing, use Docs MCP or [llms.txt](https://docs.modellix.ai/llms.txt). |
| The agent asks for a key you already set | Export `MODELLIX_API_KEY` in the shell that launched the agent, or save a profile with `modellix-cli auth login`. |
| `task download` fails with a private or reserved network error | Local proxies sometimes map CDN hosts such as `file.modellix.ai` into `198.18.0.0/15`. Retry with `--allow-private-network` for trusted Modellix CDN hosts, or download the resource URL with `curl`. |
| A result URL no longer works | Resource URLs expire after roughly 24 hours. Download results promptly. |
| A bundled helper script fails | Fall back to the direct CLI commands. The scripts are optional. |
## Next Steps
Install the full plugin from a host marketplace.
Learn the `modellix-cli` commands the skill runs.
Call Modellix directly without the CLI.
Review model pricing before running paid tasks.
# Grok Imagine Image
Source: https://docs.modellix.ai/xai/grok-imagine-image
/media-model-api/xai/xai-t2i.json post /xai/grok-imagine-image
[Core Function] Grok Imagine Image is xAI's standard text-to-image generation model. [Strengths] It excels at quickly generating solid, visually appealing images from a text prompt across a wide range of aspect ratios. [Best For] Highly recommended for: rapid prototyping, social media content, and general-purpose image generation. [Limitations] Do NOT use this model when maximum detail or fidelity is required; the Quality variant produces richer detail. [Routing] Choose this model for fast, general image generation. When the user demands maximum fidelity, route to Grok Imagine Image (Quality).
# Grok Imagine Image 2.0
Source: https://docs.modellix.ai/xai/grok-imagine-image-2-0
/media-model-api/xai/xai-t2i.json post /xai/grok-imagine-image-2.0
[Core Function] Grok Imagine Image 2.0 generates images from a text prompt. [Strengths] It supports quality (low or medium), 1k or 2k resolution, a wide range of aspect ratios including 21:9 and 5:2, and up to 10 images per request. [Best For] Highly recommended for: concept art, marketing visuals, cinematic banners, and batch generation. [Limitations] This model does not edit existing images. [Routing] For image editing, use Grok Imagine Image 2.0 Edit.
# Grok Imagine Image 2.0 Edit
Source: https://docs.modellix.ai/xai/grok-imagine-image-2-0-edit
/media-model-api/xai/xai-i2i.json post /xai/grok-imagine-image-2.0-edit
[Core Function] Grok Imagine Image 2.0 Edit edits one to three source images from a text prompt. [Strengths] It supports quality (low or medium), 1k or 2k resolution, and the same aspect ratios as Grok Imagine Image 2.0. [Best For] Highly recommended for: restyling, combining up to 3 references, and iterative refinement. [Limitations] Requires at least one source image; at most 3 images per request. This model does not generate from text alone. [Routing] For text-to-image generation, use Grok Imagine Image 2.0.
# Grok Imagine Image Edit
Source: https://docs.modellix.ai/xai/grok-imagine-image-edit
/media-model-api/xai/xai-i2i.json post /xai/grok-imagine-image-edit
[Core Function] Grok Imagine Image Edit is xAI's standard image editing model. [Strengths] It excels at quickly applying prompt-guided edits and style changes to one or more source images. [Best For] Highly recommended for: fast restyling, quick variations, and lightweight image edits. [Limitations] Do NOT use this model when maximum edit fidelity is required; the Quality variant preserves more detail. A maximum of 3 source images is supported. [Routing] Choose this model for fast edits. When the user demands maximum fidelity, route to Grok Imagine Image Edit (Quality).
# Grok Imagine Image Quality
Source: https://docs.modellix.ai/xai/grok-imagine-image-quality
/media-model-api/xai/xai-t2i.json post /xai/grok-imagine-image-quality
[Core Function] Grok Imagine Image (Quality) is xAI's high-fidelity text-to-image generation model. [Strengths] It excels at producing richly detailed, high-quality images from a text prompt, with flexible aspect ratios and an optional 2K resolution. [Best For] Highly recommended for: detailed concept art, marketing visuals, and any scenario where image quality is prioritized over generation speed. [Limitations] Do NOT use this model when latency is critical, as generation is slower than the standard model. [Routing] Use this model by default when the user emphasizes quality or detail. For faster, lighter generation use Grok Imagine Image (standard).
# Grok Imagine Image Quality Edit
Source: https://docs.modellix.ai/xai/grok-imagine-image-quality-edit
/media-model-api/xai/xai-i2i.json post /xai/grok-imagine-image-quality-edit
[Core Function] Grok Imagine Image Edit (Quality) is xAI's high-fidelity image editing model. [Strengths] It excels at applying detailed, prompt-guided edits and style transformations to one or more source images while preserving fine detail. [Best For] Highly recommended for: high-quality restyling, detailed inpainting-style edits, and combining up to 3 source images. [Limitations] Do NOT use this model when latency is critical, as it is slower than the standard edit variant; a maximum of 3 source images is supported. [Routing] Use this model by default for quality-sensitive edits. For faster edits, route to Grok Imagine Image Edit (standard).
# Grok Imagine Video
Source: https://docs.modellix.ai/xai/grok-imagine-video
/media-model-api/xai/xai-t2v.json post /xai/grok-imagine-video
[Core Function] Grok Imagine Video is xAI's text-to-video generation model. [Strengths] It excels at generating short, dynamic video clips directly from a text prompt, with controllable duration, aspect ratio, and resolution. [Best For] Highly recommended for: short social clips, animated concepts, and dynamic scene generation from a description. [Limitations] Do NOT use this model when you have a starting image or reference subjects, or when you need resolutions above 720p or clips longer than 15 seconds; it is limited to 480p/720p and 15s. [Routing] Use this model when the user wants a video from text only. If a starting image is provided, route to the Image-to-Video model; for reference-driven character video, use Reference-to-Video.
# Grok Imagine Video 1.5 I2V
Source: https://docs.modellix.ai/xai/grok-imagine-video-1-5-i2v
/media-model-api/xai/xai-i2v.json post /xai/grok-imagine-video-1.5-i2v
[Core Function] Grok Imagine Video 1.5 I2V animates a single starting image into a video using the Grok Imagine 1.5 generation backbone. [Strengths] It excels at producing motion from one starting frame with the improved 1.5 model. [Best For] Highly recommended for: animating a photo or illustration when the 1.5 generation backbone is preferred. [Limitations] Do NOT use this model for text-only generation, for resolutions above 1080p, or for clips longer than 15 seconds; it requires a starting image and is limited to 480p/720p/1080p and 15s. [Routing] Use this model when the user provides one starting image and prefers the 1.5 backbone.
# Grok Imagine Video Edit
Source: https://docs.modellix.ai/xai/grok-imagine-video-edit
/media-model-api/xai/xai-v2v.json post /xai/grok-imagine-video-edit
[Core Function] Grok Imagine Video Edit applies a prompt-guided transformation to an input video. [Strengths] It excels at restyling and modifying an existing video while keeping its original duration and aspect ratio. [Best For] Highly recommended for: restyling clips, applying visual effects, and prompt-driven video edits. [Limitations] Do NOT use this model to change the duration, aspect ratio, or resolution; the output preserves the input video's duration and aspect ratio, and those parameters are not configurable. Input video constraints (e.g. format/length) are enforced by the upstream provider. [Routing] Use this model when the user provides a video and wants it edited/restyled. To make a video longer, use Video Extend.
# Grok Imagine Video Extend
Source: https://docs.modellix.ai/xai/grok-imagine-video-extend
/media-model-api/xai/xai-v2v.json post /xai/grok-imagine-video-extend
[Core Function] Grok Imagine Video Extend continues an existing video, generating additional footage beyond its end. [Strengths] It excels at seamlessly extending a clip with new prompt-guided motion. [Best For] Highly recommended for: lengthening short clips, continuing a scene, and adding follow-on action. [Limitations] Do NOT use the `duration` parameter expecting it to set the total video length; it only controls the length of the appended segment (2-10 seconds). Input video constraints are enforced by the upstream provider. [Routing] Use this model when the user wants to make a video longer. To restyle or modify an existing video, use Video Edit.
# Grok Imagine Video I2V
Source: https://docs.modellix.ai/xai/grok-imagine-video-i2v
/media-model-api/xai/xai-i2v.json post /xai/grok-imagine-video-i2v
[Core Function] Grok Imagine Video I2V animates a single starting image into a video. [Strengths] It excels at producing smooth motion from one starting frame, guided by a text prompt for the desired movement. [Best For] Highly recommended for: bringing a photo or illustration to life, dynamic product showcases, and quick cinematic motion from a still. [Limitations] Do NOT use this model for text-only generation, for resolutions above 720p, or for clips longer than 15 seconds; it requires a starting image and is limited to 480p/720p and 15s. [Routing] Use this model when the user provides exactly one starting image. For reference-driven character video, use Reference-to-Video; for text-only video, use Text-to-Video.
# Grok Imagine Video R2V
Source: https://docs.modellix.ai/xai/grok-imagine-video-r2v
/media-model-api/xai/xai-i2v.json post /xai/grok-imagine-video-r2v
[Core Function] Grok Imagine Video R2V generates a video from a text prompt while preserving the subjects shown in up to 7 reference images. [Strengths] It excels at keeping character/subject identity consistent across a newly generated scene driven by the prompt. [Best For] Highly recommended for: character-driven video, placing a specific subject into a new scene, and blending features from multiple references. [Limitations] Do NOT use this model to simply animate a single image as-is (use Image-to-Video), or for clips longer than 10 seconds or resolutions above 720p; duration is capped at 10s and resolution at 480p/720p. [Routing] Use this model when the user provides reference images of a subject and wants a new action/scene described by a prompt. To simply animate a single image as-is, use Image-to-Video.
# Grok Voice ASR
Source: https://docs.modellix.ai/xai/grok-voice-asr
/media-model-api/xai/xai-s2t.json post /xai/grok-voice-asr
[Core Function] Grok Voice ASR transcribes a single public audio URL into text via an async task. [Strengths] Word-level timestamps, optional speaker diarization, multichannel transcription, Inverse Text Normalization (format + language), keyterm biasing, and filler-word control. [Best For] Meeting notes, call-center recordings, captions, and batch audio-to-text pipelines. [Limitations] Do NOT use file upload; URL-only input. Do NOT use for live or real-time streaming transcription. Audio must be publicly reachable (max 500 MB). [Routing] Use this model for xAI Grok Voice ASR quality with URL-based audio.
# Grok Voice TTS
Source: https://docs.modellix.ai/xai/grok-voice-tts
/media-model-api/xai/xai-tts.json post /xai/grok-voice-tts
[Core Function] Grok Voice TTS converts text into natural spoken audio with expressive voices and optional speech tags embedded in the text. [Strengths] Supports 20+ languages (plus auto-detect), 26 built-in voices, multiple codecs (mp3/wav/pcm/mulaw/alaw), and custom voice IDs. [Best For] Product voiceovers, IVR/telephony (mulaw/alaw), and multilingual narration. [Limitations] Do NOT exceed 15,000 characters per request. Audio is returned on the completed task after polling. [Routing] Use this model for xAI Grok Voice quality or custom cloned voices. [Built-in voices] Original: eve (default), ara, leo, rex, sal. Flagship: altair, atlas, carina, castor, celeste, cosmo, helios, helix, iris, kepler, lumen, luna, lux, naksh, orion, perseus, rigel, sirius, ursa, zagan, zenith — case-insensitive; custom voice IDs are also accepted as voice_id.