glossary

What is image-to-video (i2v)?

Image-to-video (i2v) is AI video generation that starts from a still picture: the model treats your image as the first frame and animates what happens next. Text-to-video (t2v) starts from a written prompt alone, so the model invents the subject, the setting and the framing itself.

The practical difference is control. With text-to-video you describe a scene and accept the face, product and composition the model picks. With image-to-video those are already decided by your picture, and the prompt only has to describe motion: what moves, how the camera travels, what happens. That is why i2v is the usual choice when a specific person, product or brand look has to survive into the clip.

On AIGE, every video model runs as text-to-video until you give it a start picture, and then the same model runs as image-to-video. Kling V3, Veo 3.1, Seedance and H3 Max all accept a start picture. Most also accept an end picture, so the model animates the move from one frame to the other; Kling 3 Turbo is the exception and takes a start picture only. On several models, Kling included, the clip takes its shape from the start picture, so crop the image to the aspect ratio you want before you generate.

Two related modes build on the same idea. Extend takes the last frame of a clip and generates what happens next, and a bridge takes the last frame of one shot and the first frame of the next and generates the shot between them. Reference-to-video is different: on Seedance and Seedance 2.5 you supply images as references for a character, a set or a palette, and the model composes new frames from them without starting on any one of them.

A saved character can stand in for the start picture. When you ask for a video of a saved character through AIGE's MCP server and give no frame, the server uses the character's saved picture as the starting frame, which is how the person in your clips matches the person in your images.

Frequently asked questions

What is the difference between image-to-video and text-to-video?

Image-to-video animates a picture you supply, so the subject and framing are fixed and the prompt describes motion. Text-to-video generates everything from the prompt, so the model chooses the subject and framing.

Which is better, image-to-video or text-to-video?

Use image-to-video when a specific face, product or look must appear in the clip. Use text-to-video for exploring ideas, or when no particular subject has to match.

Can I set both the first and the last frame?

On AIGE, yes for most video models: give a start picture and an end picture and the model generates the motion between them. Kling 3 Turbo takes a start picture only.

Does image-to-video cost more than text-to-video?

On AIGE the price depends on the model, the length and the resolution of the clip. The studio shows the exact credit cost before anything runs.