AMT Interpolation
Video editing
Interpolate between video frames
Find a model. See what it can do.
122 models in this collection.
Video editing
Interpolate between video frames
Text to video · Video editing
Animate your ideas!
Text to video · Video editing
Animate your ideas in lightning speed!
Video editing
Video background removal version of bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)
Video editing
Remove backgrounds from any video with Bria's VRMBG 3.0. Fast, accurate background removal across talking heads, podcasts, product videos, commercials, and cinematic footage.
Image to video
Omnihuman v1.5 is a new and improved version of Omnihuman. It generates video using an image of a human figure paired with an audio file. It produces vivid, high-quality videos where the character’s emotions and movements maintain a strong correlation with the audio.
Text to video · Image to video
Text to Video endpoint for Seedance 1.0 Pro Fast, a next-generation video model designed to deliver maximum performance at minimal cost
Text to video · Image to video
Generate videos with audio with Seedance 1.5
Text to video · Image to video · Video editing
Generate video from text using NVIDIA's 2B Cosmos Post-Trained Model
Text to video
Generate video from text and videos using NVIDIA's 2B Cosmos Distilled Model
audio to video
Generate dubbed videos or audios using ElevenLabs Dubbing feature!
Image to video
FLUX 3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.
Image to video
VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video
Image to video
Animates a still image into video with audio. Extends a single frame into coherent motion, grounded in Gemini's physical understanding of how scenes and subjects behave.
Text to video · Image to video
Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint generates video with synchronized native audio from a text prompt, grounded in Gemini's real-world knowledge and physics understanding, with cinematic camera control expressed in natural language.
VideoText to video
Google Veo 3 for text to video.
VideoText to video
Google Veo 3 Fast for text to video.
VideoImage to video
Google Veo 3 Fast Image to Video for image to video.
VideoImage to video
Google Veo 3 Image to Video for image to video.
Text to video · Image to video
Generate videos with audio from text using Grok Imagine Video.
Text to video · Image to video
Generate videos from prompts with audio using xAI's Grok Imagine 1.5 Video model.
Text to video
Generates 768p video with audio in a 16-bit pixel-art style from text prompts or an optional first-frame image. Supports durations of 5–15 seconds.
Text to video · Image to video
H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
Text to video · Image to video
Generate 1080p video with synchronized native audio from a text prompt. Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4. Duration: 3–15s.
Take it into the studio. See the price before you create.
Browse models here without an account. Sign in when you are ready to create.